Field Notes

A Score for Every Letter

2026-09-105 min readAIScienceHealth

AlphaGenome Atlas has a prediction for every possible single-letter DNA change. Exhaustive coverage can make a useful estimate feel more complete than the knowledge beneath it.

A researcher arrives with one genetic variant from an unsolved rare-disease case. The search space contains millions of plausible distractions. Before another sample is prepared or another paper is opened, a number is already waiting.

Google DeepMind's new AlphaGenome Atlas contains predictions for roughly nine billion single-nucleotide variants: each of the three possible substitutions at nearly every letter in the human genome. The dataset is one petabyte. The model does not have to wake up and think about a variant when a researcher enters it. The prediction has been computed in advance.

This is an extraordinary change in the texture of a research question. A blank space has become a lookup.

AlphaGenome takes up to one million letters of DNA as input and predicts thousands of molecular signals, including gene expression, splicing, chromatin accessibility, transcription-factor binding, and the physical contacts made by folded DNA. In its original Nature paper, the model matched or exceeded the strongest external systems in 25 of 26 variant-effect evaluations.

The Atlas puts those predictions behind a portal and adds the AlphaGenome Variant Impact score. The AVI combines signals from AlphaGenome and AlphaMissense into one number so researchers can rank possible variants, then inspect feature attributions showing which predicted biological processes pushed the score up or down.

The score answers a practical question: where should scarce attention go first?

That is not a small service. In one case described by DeepMind, researchers investigating a severe epilepsy used the Atlas to prioritize a non-coding change in the DNM1 gene. The predicted mechanism pointed toward an incorrect splice site. Literature review and laboratory experiments followed, eventually producing enough evidence to reclassify the variant. In another analysis, a region containing 526 candidates narrowed to four for closer study.

The number did useful work because it did not have to do all the work.

Still, completeness has a peculiar authority. When every possible single-letter change has a score, it becomes easy to feel that the system knows something complete about every possible change. The database has no empty row to remind us how much remains uncertain. Coverage and certainty begin to share the same visual language.

They are not the same property.

AlphaGenome predicts molecular effects from sequence. It does not observe a person's symptoms, family history, environment, treatment, or the other variants in their genome. A recent technical assessment notes that its predictions of combined, personal-genome effects remain limited and that sequence-to-function models hold much of the surrounding cellular state fixed. The Atlas also starts with single-letter substitutions. Biology has not agreed to become a table merely because the table is very large.

Clinical interpretation already has a language for this distinction. ClinGen's guidance treats calibrated computational predictions as one possible strength of evidence inside a broader classification process. Its recommendations for non-coding variants also ask whether the regulatory element is relevant, whether the proposed mechanism fits the condition, whether population evidence supports it, and whether functional testing confirms the effect. A molecular prediction can help form the case. It is not the case by itself.

The Atlas appears designed for research, not direct diagnosis. That boundary deserves to remain visible as the data becomes easier to reach. DeepMind already offers the Atlas through an AI science skill that can rank variants and generate an explanation in a chat window. The workflow can move from millions of candidates to a plausible biological story in minutes.

Speed makes the seams more important.

A useful result should keep three things visually separate: what the model predicted, what researchers observed, and what clinicians concluded for a particular person. The score should travel with its model version, tissue and cell context, calibration, and known blind spots. Feature attributions should open the prediction rather than decorate it. An explanation should end with the evidence that would test it, not with the smooth confidence of a mystery already solved.

This sits upstream from Health AI Needs A Handoff. That Field Note asked how uncertainty and responsibility move from a capable answer into human care. The Atlas shows that the handoff starts earlier. A ranked variant may move from database to agent, from agent to researcher, from researcher to laboratory, and eventually into a clinical interpretation. Each step changes what the number is allowed to mean.

There is a generous version of this future. Researchers who cannot run a giant model or write an analysis pipeline can explore a vast body of predictions. Rare-disease teams can spend less time sorting noise and more time testing plausible mechanisms. Non-coding regions that were once especially difficult to interpret become more available to inquiry.

The generosity depends on preserving the difference between availability and authority.

An atlas can cover an entire territory without predicting every journey through it. AlphaGenome Atlas gives science a remarkable new map of possible molecular effects. The score makes the search space smaller. It should not make the question look finished.