How well a model matched a drug's transcriptome told you almost nothing about whether it got the mechanism right: median Spearman 0.052. 12 models plus 3 baselines; a plain linear-regression baseline led on mechanism fidelity.
A single-cell model can predict a transcriptome almost perfectly and still get the drug response backwards. That is a serious problem if we want “virtual cells” to do more than reconstruct gene-expression profiles. Let introduce scDrugPerturb-Bench, a mechanism-aware benchmark for single-cell drug perturbation models. Instead of asking only: Does the predicted transcriptome resemble the measured one? it asks: Did the model recover the biological response that actually matters? The benchmark brings together 2.5 million cells, 181 datasets, 101 drugs and 423 literature-curated drug-response cases, with experimentally supported directional changes for 717 key genes across different cellular contexts. The authors then evaluated 12 perturbation-prediction models, simple baselines and multiple single-cell foundation-model representations. The result is uncomfortable: expression similarity was only weakly aligned with mechanism fidelity. A prediction could look convincing at the whole-transcriptome level while getting key genes, pathway direction or mechanism specificity wrong. One example makes the problem tangible. For a drug perturbation in human iPSC-derived microglia, the predicted profiles achieved very low reconstruction error, yet 0 of 8 annotated key genes changed in the correct direction. The transcriptome looked right. The biology was backwards. To capture this distinction, the authors introduce a Mechanism Fidelity Score that tests whether predictions preserve key-gene direction, effect size, gene-set coherence, mechanism specificity and pathway-level polarity - not just global expression similarity. The hard-negative experiments are perhaps even more revealing. Some models could generate plausible-looking perturbation responses by exploiting generic transcriptional programmes or overall perturbation strength, without recovering the specific drug-response signature. This matters far beyond benchmarking. If virtual-cell models are eventually used to prioritize compounds, infer mechanisms or design experiments, then a biologically wrong prediction does not become useful simply because its transcriptome correlation is high. Prediction quality depends on what we choose to measure. Reconstructing expression is one task. Recovering mechanism is another. And for biological AI, those two should not be treated as the same thing.
1
1
4
1,351
Li et al., A mechanism-annotated benchmark reveals limited fidelity to drug-response signatures in single-cell perturbation models, bioRxiv preprint, August 2026 (not peer reviewed), Harbin Institute of Technology Shenzhen and MindFlow.ai. biorxiv.org/content/10.64898…

Sep 24, 2026 · 12:00 AM UTC

163
Sort replies: Relevant Recent Liked