News · 2026-09-08
DeepMind precomputed every single-letter change in human DNA
Google DeepMind released AlphaGenome Atlas on September 8, 2026: a free, browsable database containing a predicted molecular effect for every one of the roughly 9 billion possible single-letter changes to human DNA. The dataset runs to about one petabyte, roughly thirty times the size of the AlphaFold protein database, and each variant carries an average of about 27,000 individual predictions condensed into a single impact score. DeepMind states that it has not been validated for clinical use.
Key facts
- 9 billion single-nucleotide variants, each with an average of about 27,000 predictions.
- The dataset is roughly 1 petabyte, about 30 times the size of the AlphaFold Database.
- Free for academic and non-commercial research through a web portal and API; commercial access via Google Cloud.
- Primary source: DeepMind's announcement, September 8, 2026; covered by Nature and IEEE Spectrum.
The problem this addresses is one of the least glamorous and most consequential in modern medicine. Sequencing a patient's genome is now cheap and routine. Interpreting it is neither. A sequencing run turns up thousands of places where a patient's DNA differs from the reference, and for most of them nobody can say whether the difference matters. These get filed as "variants of uncertain significance" — a phrase that means the test worked and the answer is still a shrug. Families with an undiagnosed genetic disease frequently end up there.
The old approach was to compute an answer when someone asked. DeepMind's approach was to compute all of them in advance. Rather than requiring a researcher to install a genome language model, find a GPU and run inference, the Atlas ships as a lookup: type in a position, get the prediction. That is a meaningful accessibility change for the large fraction of working biologists and clinical geneticists who do not write code.
The scientifically interesting move is the compression. AlphaGenome produces roughly 27,000 numbers about a single variant — how it changes gene expression, how it affects the way DNA is transcribed, what happens across many different cell types. The Atlas adds the AlphaGenome Variant Impact score, or AVI, which collapses all of that into one number meant to tell you at a glance whether a variant is worth attention. Crucially it works across both the 2 percent of the genome that codes for proteins and the 98 percent that does not — the regulatory bulk, where most trait-associated variants actually live and where older prediction tools were weakest.
That compression is also the honest place to point a reader's attention. Going from 27,000 numbers to one is enormously useful and necessarily lossy. A single score cannot preserve which cell type an effect appears in, which direction it runs, or whether the mechanism is regulatory or structural. A variant that badly disrupts liver tissue and does nothing in neurons, and a variant that mildly perturbs everything, can arrive at the same headline number by completely different routes. The score is a triage signal, and triage signals are dangerous exactly when they are convenient.
The analogy is a national weather service that publishes a forecast for every square metre of the country rather than every city. The resolution is genuinely astonishing and genuinely useful. It is still a forecast — a model's opinion about what would happen, not a record of what did.
DeepMind's own validation language is worth reading precisely. The company says "our testing shows that the AVI score provides best-in-class performance across many variant pathogenicity and rare disease benchmarks," and points to collaborator work — rare disease identification with the GREGoR Consortium, studies using the UK Biobank — rather than to a named benchmark table on the announcement page. That is a real claim backed by real users, but it is not the same as an independently audited accuracy figure, and the distinction matters when the output is a number that looks like a diagnosis.
DeepMind says so itself, and this is the most important sentence on the page: "AlphaGenome has not been validated for, and is not approved for, any clinical use." The Atlas is not a diagnostic tool, and the company is explicit that its output is not medical advice.
Why it matters: precomputing an entire hypothesis space is a pattern that will spread. Once a model is good enough and inference is cheap enough, the move from "run the model when asked" to "run it on everything once and publish the answers" changes who can use it — from labs with GPUs to anyone with a browser. That is how AlphaFold reshaped structural biology, and the Atlas is thirty times larger. The risk that comes with it is equally predictable: a precomputed score in a searchable database acquires an authority that a model output does not, and the caveat travels less well than the number. The question worth asking over the next year is not whether the predictions are impressive — they are — but what happens the first time an AVI score shows up in a clinical workflow it was explicitly not approved for.
Key questions
What is AlphaGenome Atlas?
Can doctors use the Atlas to diagnose patients?
What is the AVI score?
Cite this
APA
Ground Truth. (2026, September 8). DeepMind precomputed every single-letter change in human DNA. Ground Truth. https://groundtruth.day/news/deepmind-precomputed-every-single-letter-change-in-human-dna.html
BibTeX
@misc{groundtruth:deepmind-precomputed-every-single-letter-change-in-human-dna,
title = {DeepMind precomputed every single-letter change in human DNA},
author = {{Ground Truth}},
year = {2026},
month = {sep},
url = {https://groundtruth.day/news/deepmind-precomputed-every-single-letter-change-in-human-dna.html}
}
Comments are replies to this story on Bluesky — reply with any Bluesky account to join in.