Claude Opus 5's performance on MedAgentBench drops substantially compared to its predecessors, Opus 4.7 and 4.8: roughly half their success rate. π₯
We benchmarked Opus 5 as an autonomous clinical agent across 10 EHR task types, and it landed at ~40% pass@1 (avg of 3Γ 300-task runs), well below both earlier Opus models on the same benchmark.
We dug into why π§΅
1
1
17
2,901
The cause isn't clinical knowledge, it's agentic behavior. In ~60% of turns, Opus 5 hallucinates the API's response inside its own message (emitting a tool call and a fabricated server reply together), instead of issuing one call and waiting. Opus 4.8 did this ~7% of the time.
To isolate the effect, we ran a diagnostic that ignores the fabricated text and executes only the real call. Score rises to ~68% but still short of 4.8. Retrieval tasks recover; multi-step order-writing tasks stay low (potassium replacement 7%, HbA1c ordering 23%), and it sometimes reports real patients as "not found." A capable model held back by not respecting the tool-use loop.
Note: the ~40% is the leaderboard figure, same harness as every other model. The ~68% is a diagnostic only, with a temporary parser tweak, and isn't directly comparable.
π Leaderboards: medicalsphere.ai/benchmarks
π§ͺ Full pipeline: github.com/medicalsphere/Medβ¦
Jul 28, 2026 Β· 9:30 PM UTC
5
279

