Anthropic shipped Mythos 5.1 today, leading the ProteinGym benchmark at 49.3% rank correlation against real lab measurements. Two weeks ago they published Claude running autonomous protein binder design campaigns at a 27% hit rate versus the 10-15% that's typical today.
Both results are impressive, but both point to a problem most people aren't (yet) seeing.
Start with the shape of that ProteinGym chart:
Mythos 5.1 - 49.3%
Opus 5 - 47.7%
Mythos 5 - 45.8%
Gemini 3.1 Pro - 37.0%
Sonnet 5 - 36.6%
GPT-5.6 - 35.5%
Six frontier general-purpose models inside a 14-point band. None of them dedicated protein models. All trained on roughly the same public corpus.
That's not anyone achieving a moat; it's a a floor rising.
What ProteinGym actually measures: ~217 deep mutational scanning assays covering millions of variants, where every one of them is an experiment somebody already ran and published.
Predicting results that exist is a different problem from generating results that don't.
The binder campaign has the same shape, and Anthropic says so in its own technical report. Co-folding confidence turned out to be a useful filter but not a guarantee, and "experimental screening remains the only way to learn" which targets a campaign actually succeeded on. Adaptyv, who ran the wet lab, called it an open-loop experiment and said the next step is closing the loop.
You can't close that loop at 30 designs per target with a multi-week CRO turnaround. You get one shot and no statistical power to learn anything from it.
In contrast – last March,
@ManifoldBio and
@NVIDIA tested 1,000,000 designs against 127 targets in a single multiplexed experiment, measuring over 100 million protein-protein interactions. Roughly 750x the design count of the Anthropic campaign, in one run. (disclosure: I'm on Manifold's board, and an investor.)
The model they were validating in that study was Proteina-Complexa, which is one of the ten open-source design models Claude used in the Anthropic campaign.
Design is converging on a shared toolkit, and ... commoditizing. Measurement is not.
Which brings me to the ladder below. Every public benchmark, and both of Anthropic's recent headline results, live near the bottom of it. Expression. Binding. In vitro affinity.
There is no ProteinGym for biodistribution and for what actually translates into human biology (and into therapeutics that work). Not because nobody wants one, but because the data doesn't exist at benchmark scale.
Someone has to generate it, at billion-scale. Whoever does has the data that matters to train the models that will actually make a difference.
My bet is that the first to million-scale will be the first to billion-scale.