Seven deep learning foundation models entered a clean perturbation benchmark.

All seven lost to a simple linear addition of single gene effects.

Ahlmann-Eltze, Huber and Anders published the most valuable perturbation prediction result of the cycle: a clean benchmark where seven deep learning models, including scGPT, Geneformer, GEARS, and CPA, all lost to a linear baseline on unseen gene perturbations.

The winning model had no foundation model story. It added known single perturbation effects together and sat down. A linear model pretrained on a different cell line won the whole exam.

Read it as a discipline lesson. "Bigger model" was never the question. The real question is: compared to what? The additive baseline is now the free, public yardstick every perturbation model must clear. It is the reporting discipline a field adopts when it can falsify itself.

For anyone running experiments on a model prediction, this is a way to cut risk. A model that beats the baseline tells you something worth testing. A model that does not tells you to spend the wet lab budget elsewhere.

Five independent groups now converge on the same result, and the paper code is public.

What is the simplest baseline in your pipeline that you are not comparing against?

#foundationmodels #perturbation #computationalbiology

置顶第一条评论(已发布)

Reference & reproducible code for this benchmark:

  1. Nature Methods paper: https://doi.org/10.1038/s41592-025-02772-6

  2. Open-source benchmark pipeline: https://github.com/const-ae/linear_perturbation_prediction-Paper

For more translational data strategy deep-dives and computational biology notes: https://omicswithjohnson.net/insights