Tahoe-100M's chemical perturbation atlas in history reads like a cheat sheet. 95.6 million cells. 47 cell lines. Exactly 379 unique molecules.

1,135 counts perturbation arms, each arm being the same molecule at one dose. Somewhere in the retelling, an arm became a molecule, and 379 became "thousands."

Of those 379, 69 percent are approved drugs spread across 25 mechanisms. The atlas is heavily weighted toward known biology and established chemical scaffolds.

Here is the part that matters for anyone building on it: a model trained on 379 molecules can interpolate beautifully among them. Every prediction about molecule activity is extrapolation, and extrapolation is where deep learning fails silently. It does not tell you it is extrapolating. It gives you a number that looks confident.

Data quantity was never the problem in the right dimension. You can sample a hundred million cells from the same 379 molecules and the model still only knows 379 molecules.

When you size a perturbation experiment or buy into an atlas narrative, ask two questions: how many unique compounds, and how much of the space is truly new chemistry? The answers decide whether your model is predicting or guessing.

When did you last check the chemical vocabulary of your training set?

Core Scientific Takeaway

A model trained on 379 molecules can interpolate among known chemical scaffolds, but predicting activity on novel chemotypes is pure extrapolation. Scale narratives that emphasize cell counts over molecular diversity overstate what deep learning can reliably learn.