Today we're releasing BlueBench-Simulation, a brand new benchmark suite evaluating agent performance on defensive cybersecurity tasks against simulated enterprise environments.
BlueBench-Simulation is releasing with 12 scenarios across 3 new benchmarks, each covering a different stage of the intrusion lifecycle:
- Initial Access & Command-and-Control (C2)
- Identity & Active Directory Attacks
- Impact & Exfiltration
Every scenario is a generated enterprise environment with weeks of routine activity hit by an intrusion. Models start from a single weak alert and produce a threat hunting report, an incident response report, or a detection query.
- The most notable result,
@SpaceXAI's Grok 4.6 led the suite at 71.8%, taking two of the three benchmarks and finishing first overall on pure capability.
-
@claudeai Opus 4.8 came in second at 68.9%. We'd expect Anthropic's newer models to do at least as well, but Fable 5, Opus 5, and Sonnet 5 all ran into cybersecurity refusals, so we couldn't get a clean read on them.
- The
@OpenAI GPT-5.6 model family once again anchored the value end of the Pareto curve
More details in thread below ⬇️
Full results:
cotool.ai/research