We received early access to
@claudeai Opus 5.5 and evaluated its defensive security capabilities on
@depthfirstlabs dfbench.
At high reasoning effort, the model achieved 54.9% recall and 48.5% precision, up from 48.1% and 44.2% at medium. Detection cost averaged $8.42 per task, roughly one-third the cost of GPT 5.6 Sol at high and half the cost of Grok 4.7 at high, though with lower recall than both.
Surprisingly, increasing reasoning effort improved both recall and precision. We’ve typically seen a tradeoff in our evaluations, with models finding more vulnerabilities at the expense of more false positives. Opus 5.5 improved on both.
It also reached 79.5% macro recall on differential analysis, up from 75.3% at medium. At $1.78 per task, it showed strong performance in tracking vulnerabilities across code changes.