Anthropic used CyScenarioBench, our benchmark for multi-stage offensive cyber operations, in the cyber evaluations for Claude Opus 5.5.
Across a ten-challenge subset, Opus 5.5 averaged a 67.6% solve rate, ahead of Claude Mythos 5.1 at 61.7% and Claude Opus 5 at 53.0%.
Full write-up in the first comment.