A database with no record of binding is not a record of non-binding

An antibody recognizes its target by a patch of its surface — the epitope. In therapeutic discovery this becomes a huge selection problem: there is a library of antibodies, there is a target protein, and the lab can measure only a small fraction of the possible pairs. Structural models therefore promise to screen candidates in advance. A new preprint from Arnav Solanki and colleagues asks whether AlphaFold3 can distinguish a binding antibody from an unsuitable one; a convincing three-dimensional scene is not enough to make that call.

The authors took 3,401 experimentally determined complexes from the SAbDab database and added 23,798 control combinations. With ten independent runs, AlphaFold3 called 1,486 of the known pairs positive. After rerunning the missed pairs with 100 random initializations, the number rose to 1,808: 53% of the known complexes. At the same time, under the chosen rule, 848 control pairs received a positive score. This is not a paper about a finished antibody, nor a clinical result; it is a map of how much of the laboratory queue the model can cut, and where it creates a new queue of errors.

What exactly AlphaFold3 predicts

An antibody sequence contains two variable arms. Each carries the complementarity-determining loops (CDRs): these are what usually fit the contours of the antigen. AlphaFold3 takes the sequences of the antibody and the protein, generates candidate joint structures and assigns internal scores. Two of them were decisive in this benchmark: PAE, the predicted error in the relative positioning of parts of the model, and ipTM, a confidence score for the interface.

The authors called a prediction positive if PAE was below 2.5 or ipTM exceeded 0.7. That is a threshold on the model's confidence in the contact it has drawn, not a measurement of binding strength. Answering the latter question requires concentrations, kinetics and biochemical experiment: surface plasmon resonance, biolayer interferometry or a cell-based assay, for example. A structure predictor arranges atoms; the experiment shows whether the complex exists in solution and whether it works in a biological system.

The distinction looks formal until drug discovery begins. A high score can send a pair into expensive laboratory validation, and a low score can drop an antibody that really does bind the target. The usefulness of the model is therefore set by two error streams at once: false candidates and lost candidates.

How the test works and what the 848 "false" complexes mean

The positive set consists of antibody–antigen structures already deposited in the Protein Data Bank. For the negative set, the researchers combined antibodies with random human proteins, then ran every antibody against six fixed proteins from human, rat, fruit fly, bacteria and plant. SAbDab has no record of a complex for such combinations. This is a practical way to create many hard decoys for the model, but not a series of direct experiments showing that each pair fails to bind. The 848 results therefore show the hit rate within this computational control; they do not amount to 848 proven cases of physical binding or its absence.

Still, the pattern of the errors is telling. AlphaFold3 placed spurious antibodies on the surfaces of the six control proteins, and some regions became recurring "hot spots". The model did not single out the same antibody as universally sticky: most antibodies had zero or one false target, and fewer than ten picked up three. What emerges is a more troubling kind of error: for a given incorrect pair, the model can assemble a locally plausible interface.

This is where the gap between structural plausibility and specificity opens up. In a separate preprint from Markus Hoffmann and colleagues, the authors arrived at a similar question using shuffled nanobody–antigen pairs: the internal structure scores separated genuine combinations from incorrect ones poorly. Both studies set developers the same test: benchmark the model against real and deliberately constructed negative pairs, rather than measuring only accuracy on complexes that are already known.

Two different errors: missing the target and missing the pose

The standard structural metric, DockQ, compares the predicted interface with the experimental one: it accounts for shared contacts and for the geometry of the arrangement. It is useful, but a single number blends two questions. Did the model find the right patch of the protein? And did it give the antibody the right angle, distance and CDR loop shapes?

Solanki and colleagues added two measures. "Epitope shift" measures how far the predicted contact patch has moved from the true one across the antigen surface. "Antibody displacement" measures how far its center has moved after the antigens are aligned. Among 1,915 false negatives, 651 — about 34% — had an epitope shift of less than 10 Å. In other words, the model often brought the antibody to the right neighborhood, but the precise interface geometry remained weak and DockQ was low.

For a pipeline these are two different routes. A pair with the wrong epitope calls for a different antibody, a different target or a different selection approach. A pair with the right epitope and a poor pose can become a candidate for local docking, molecular dynamics simulation and subsequent laboratory testing. The authors name such methods as a possible way to refine the conformation; the benchmark itself did not test how many candidates could be rescued that way.

Why protein size and flexible loops break confidence

The larger the antigen surface, the more places where an antibody pose has to be considered. In this work the fraction of positive predictions fell as target length, surface area and volume increased. Flexible, intrinsically disordered regions are particularly awkward: they have no single stable shape that can be reliably represented as a landscape for binding. In the control protein CD274, spurious antibodies settled on the folded globular portion and avoided its disordered N-terminus.

There is a dataset skew as well. AlphaFold3 recovered complexes solved by X-ray crystallography noticeably more often than complexes from cryo-electron microscopy: recall was 0.52 against 0.30 in the baseline run. That may reflect both the richer representation of X-ray structures in the training data and properties of the molecules themselves: crystallization favors more stable, ordered objects, whereas cryo-EM often captures flexible states. The authors checked for direct overlap with AlphaFold3's training set and found no dominant effect for the specific antibodies it had "seen", but the sample is still built on PDB structures.

The broad FoldBench study produced a similar picture on a different set: across 172 low-homology antibody–antigen complexes, AlphaFold3 led the five all-atom models but reached 47.9% successful predictions by DockQ. The FoldBench authors tie the difficulty to the flexible CDR-H3 loop, correct registration of the paratope with the epitope, and the fact that confidence scores are poor at picking the best variant among the many structures generated. On ordinary protein–protein interfaces AlphaFold3 was substantially stronger; antibodies remain a class of problem apart.

More runs give more variants — and larger compute bills

In the new benchmark, 100 seeds recovered 344 complexes missed with ten. Each initialization takes the model along a different generation trajectory, so a broader search can stumble on a rare successful pose. But the improvement is not free: the number of runs went up tenfold for the false negatives alone, and the final call was still made by the same set of internal scores.

FoldBench observed the same economics. AlphaFold3's success rate rose with the number of samples, yet ranking quality capped the gain: a good variant could sit among those generated without the model's own ranking score putting it first. For a practical program this means the budget should be spread across stages: a broad, cheap filter, deeper generation for a subset of pairs, independent structural analysis and experiment. Sending every pair from a large library through 100 or 1,000 runs turns computational screening into a new bottleneck in short order.

What the next benchmark should be able to do

The present work already moves past the familiar question of whether a predicted structure resembles the PDB. It counts negative pairs, separates the target site from the pose geometry, and shows the dependence on size and on the structural method. The next step will require datasets holding antibody and antigen sequences together with binding measurements, epitopes and genuine negative results. Such data are harder to assemble: publications tend to preserve the successful molecules and lose the long list of failures.

For now AlphaFold3 remains a generator of well-motivated hypotheses. Its strength is narrowing the search space and handing the researcher coordinates for a possible interface. The lab decides whether that hypothesis became a molecule. For antibody development this is a useful division of labor: the model chooses what to test next, and biophysics answers whether binding actually occurred.