top of page

DeepMind Precomputes the Whole Human Genome's Mutations, Which Turns Genetic Search Into a Lookup

1 day ago
5 min read
DeepMind Precomputes the Whole Human Genome's Mutations, Which Turns Genetic Search Into a Lookup
DeepMind Precomputes the Whole Human Genome's Mutations, Which Turns Genetic Search Into a Lookup

The human genome holds roughly three billion base pairs. At every one of those positions, three single-letter substitutions are possible, which puts the total number of ways a single DNA letter can change at nine billion. On September 8, Google DeepMind published predicted molecular consequences for all nine billion of them in a resource called the AlphaGenome Atlas, built on top of the AlphaGenome model the company released in January.


The scale is the headline. The dataset runs to about one petabyte, which DeepMind describes as roughly 30 times larger than its AlphaFold protein-structure database, itself among the most heavily used scientific resources the company has produced. Where AlphaFold answered a structural question, how a protein folds, Atlas answers a functional one: what a given DNA change is likely to do to the machinery that turns genes on and off.


How the Atlas Was Built


AlphaGenome, the model underlying the Atlas, takes a DNA sequence of up to one million base pairs and predicts thousands of molecular measurements from it, including gene expression, chromatin accessibility, and RNA splicing. When the Nature paper describing it appeared in January, AlphaGenome had achieved state-of-the-art results on 25 of 26 variant-effect prediction benchmarks tested.


Turning that model into a genome-wide catalogue meant running it against every possible substitution at every position in a reference genome, then comparing each result to the unmutated sequence. Žiga Avsec, DeepMind's genomics initiative lead, told IEEE Spectrum the team's early estimate was that they needed to raise their computation speed by a factor of 80 to finish the project in a reasonable amount of time. He described the initial scope as one that "seemed impossible to do that computationally."


Each of the nine billion variants in the finished Atlas carries an average of about 27,000 individual predictions, spanning hundreds of human and mouse cell and tissue types. The dataset also includes a compendium of 2,601 recurring DNA sequence motifs, the short regulatory patterns that control when and where genes activate, mapped across those same cell types, along with predicted effects for more than 100 million short insertions and deletions observed in human genomes.


A Single Score for Nine Billion Variants


Raw prediction volume on that scale creates its own problem. Researchers querying the AlphaGenome API had repeatedly asked for a simpler way to know which variants were worth a closer look, according to reporting from Nature's news team. DeepMind's answer is the AlphaGenome Variant Impact score, a single number per variant built by combining AlphaGenome's regulatory predictions with AlphaMissense's protein-impact predictions, along with conservation statistics and predicted protein loss-of-function effects. One technical account puts the underlying feature count at 18. Each AVI score ships with attributions that break the number down into the specific molecular process driving it, such as splicing disruption or altered expression.


The split reflects a division of labor DeepMind has built across two separate tools. Avsec described it directly: AlphaMissense evaluates protein-coding changes, while AlphaGenome covers the regulatory genome, the roughly 98 percent of DNA that does not code for protein but governs how the coding 2 percent gets used. According to coverage of the accompanying technical paper, the AVI score reached state-of-the-art performance on variant pathogenicity and rare-disease benchmarks, with its largest gains concentrated in that non-coding majority of the genome, the part existing tools have historically handled worst.


A preprint from Cheng and colleagues, published alongside the Atlas launch, reports that the AVI score reliably separated disease-causing mutations from benign ones in a clinical genomics database. That preprint has not yet completed peer review.


Early Use in Unsolved Cases


The clearest test of a resource like this is whether it changes an actual diagnosis. One example comes from the GREGoR Consortium and the Broad Institute, where researchers including Laura Covill and Anne O'Donnell-Luria used AVI scores to triage variants in a rare-disease case that had gone unsolved through standard sequencing. The score flagged a heterozygous non-coding change in intron 10 of the gene DNM1, with the attribution data pointing specifically at splicing disruption as the mechanism.


A separate case, reported in a medRxiv preprint and using the underlying AlphaGenome model rather than the Atlas itself, illustrates the same approach applied earlier. Researchers scored 91 non-coding variants near the PLA2G6 gene in a family affected by PLA2G6-associated neurodegeneration, after conventional exome sequencing, genome sequencing, and other standard tests had found only a single coding variant. The model identified an intronic change predicted to create a cryptic splice site, a prediction later confirmed by RNA sequencing showing the abnormal splice junction in an unrelated carrier. The finding completed a diagnosis nearly two decades after the family's symptoms began.


What the Predictions Do Not Settle


DeepMind's own materials are direct about where the Atlas stops. The company's disclaimer states that AlphaGenome "has not been validated for, and is not approved for, any clinical use" and is "not intended to be a substitute for professional medical advice." The technical paper describing Atlas and the AVI score goes further, stating that both are research tools whose predictions belong to the evidence chain leading toward a clinical diagnosis without ever completing it alone.


That caveat matches how the tool performed in the cases described above. AVI scores narrowed a search space from dozens or thousands of candidate variants down to one worth testing, and laboratory confirmation did the remaining work: RNA sequencing that showed the abnormal splice junction directly, in the PLA2G6 case. The Atlas removes repeated model inference from that initial search. The experiment at the end of it stays exactly where it was.


That distinction matters most for the roughly 98 percent of the genome that does not code for protein. Standard clinical genetic testing has historically focused on the coding 2 percent, in large part because tools for interpreting non-coding variants lagged far behind. A change buried in an intron, like the DNM1 and PLA2G6 variants above, produces no altered protein sequence for a standard test to flag. It can still break splicing, silence a nearby gene, or shift when and where that gene turns on, and those effects are exactly what AlphaGenome was built to predict. Widening the searchable fraction of the genome from 2 percent to effectively all of it is the underlying reason DeepMind built a database this large rather than a smaller, curated one.


Access and What Comes Next


The AlphaGenome Atlas is available now through a browser-based portal that requires no coding, alongside the existing AlphaGenome API and a new skill built for Google Antigravity, DeepMind's agentic development platform. Access is free for non-commercial research as of the September 8 release. DeepMind says commercial access will follow through a Google Cloud service, though it has not given a date. Separately, the technical paper states that a static download of AVI scores carries a permissive license covering both commercial and non-commercial use.


For researchers working on rare and undiagnosed diseases in particular, the practical shift is that a genome-wide search that once required writing code and running a computationally demanding model one variant at a time is now a lookup. The Atlas cannot tell a researcher which variant caused a disease. It can tell them, out of nine billion candidates, which few hundred are worth the months of laboratory work it takes to find out. What happens next depends on how the broader research community uses that shortlist, and on how the preprints now circulating, including the AVI validation study, hold up once they clear peer review.

About the author: David Borish writes about frontier AI, enterprise deployment, and the gap between demonstrated capability and real-world adoption. He is the author of The Tony Hawk Paradox: When Video Games Predict Reality and publishes analysis at davidborish.com.

 
 

JOIN THE AI SPECTATOR MAILING LIST

CONTACT

Contacting You About:

Thanks for submitting!

New York, NY           

Db @DavidBorish.com           

  • LinkedIn
  • Instagram
  • Facebook
  • X
Back to top

© 2026 by David Borish IP, LLC, All Rights Reserved

bottom of page