For decades, the central frustration of genomics was not sequencing DNA — we got fast at that. The frustration was interpretation. A genome is three billion base pairs long, and for most of it, we had almost no idea what any given stretch was doing, when it was doing it, or why mutations in one region would cause disease in another region entirely. The genome looked like a wiring diagram written in a language nobody had ever formally been taught.
That is starting to change in a way that feels genuinely discontinuous from everything before it.
The models now being deployed in regulatory genomics — the subfield concerned with which genes get switched on, in which cells, at which times — are doing something that resembles reading that wiring diagram for real. Deepmind’s Enformer, and the wave of architectures that followed it, demonstrated that a transformer trained on genome sequences could predict gene expression levels across hundreds of cell types directly from raw DNA. Not from curated features a biologist had pre-selected. From the sequence itself. The model had to learn, internally, what a transcription factor binding site looks like, how enhancers communicate across hundreds of thousands of base pairs, how chromatin accessibility shapes which signals even reach a promoter. It learned the grammar of gene regulation by reading enough examples of the text.
What makes the current moment sharper than even that is the shift toward models that can reason about genetic variants. The critical clinical question is not just “what does this genome do” but “what does this genome do differently from that one, and why does the difference matter?” Sequence-to-function models can now be queried: insert this single nucleotide change, re-run the prediction, compare the output. The result is a quantitative score for how much that variant perturbs regulatory activity in each cell type. For common disease variants identified in genome-wide association studies, this is a substantial acceleration. Researchers can computationally prioritize thousands of candidate variants down to a handful worth pursuing experimentally, in weeks rather than years.
The spatial dimension is coming online too. Single-cell sequencing tells you which genes are active in which cells; spatial transcriptomics tells you where those cells sit in a tissue. Models trained jointly on both modalities are beginning to construct something like a functional map of how a tissue is organized at the molecular level — not a static snapshot but a predictive model of how perturbations propagate. Disrupt this regulatory element, and the model can sketch out downstream consequences in cell types several steps removed from the original target. That kind of systems-level prediction has always been the goal of molecular biology. We are getting close enough to it to be useful.
Perhaps the most striking recent direction is applying these approaches to non-coding RNA. For a long time, the vast stretches of the genome that do not code for protein were treated as background noise, or at best vaguely labeled “regulatory.” Long non-coding RNAs, of which there are tens of thousands, play roles in gene silencing, chromatin organization, and developmental timing that are still being worked out. Sequence-based models are turning out to have real predictive power over their expression patterns and secondary structures, which means for the first time there is a systematic way to generate hypotheses about what they are doing in disease contexts where their levels are anomalous.
The genome has always been a complete specification for a living organism. We just did not have the tools to parse the specification. What is emerging now is not a single breakthrough but an entire class of model that brings the language of DNA into reach — readable, queryable, and increasingly predictive. The implications for understanding disease, designing therapeutics, and grasping how evolution actually works at the molecular level are hard to overstate. The wiring diagram is finally starting to make sense.