We introduce an Antibody Discovery Benchmark, a benchmark for testing whether AI agents can make scientific decisions across the stages of therapeutic antibody discovery.
100 evaluations span ten areas from concrete drug programs: from target and modality selection through assay design, binder discovery, binding characterization, cellular pharmacology, antibody engineering, and preclinical candidate de-risking.
Therapeutic antibody discovery is a multiparameter, context-dependent process. Affinity must be considered alongside specificity, stability, solubility, expression, and biological activity.
Every measurement must be interpreted in the context of the experimental system that produced it. Display enrichment can reflect amplification bias rather than binding. Strong binding to purified antigen may not translate to recognition in its native cell-surface context. Apparent affinity can arise from avidity, while changes in valency or molecular geometry can alter function without changing the underlying binding domains.
Across 20 model–harness configurations, even the strongest systems passed only about half of all attempts. Opus 5 with Claude Code led at 53%, with models from Google and xAI following closely. GPT-5.6 Sol with PI reached only 33.8%. Performance also varied substantially by competency: Opus was strongest on target opportunity and cellular pharmacology, Gemini on epitope, escape, and structural mechanism, and GPT-5.6 Sol on sequence, enrichment, and next-cycle engineering decisions.
Manuscript:
latch.bio/txbench-ab
Leaderboard:
benchmarks.bio/txbench-ab
Sample evaluations and trajectories:
github.com/latchbio/txbench-…