Anthropic used CyScenarioBench, our benchmark for multi-stage offensive cyber operations, in the cyber evaluations for Claude Opus 5.5. Across a ten-challenge subset, Opus 5.5 averaged a 67.6% solve rate, ahead of Claude Mythos 5.1 at 61.7% and Claude Opus 5 at 53.0%. Full write-up in the first comment.
5
6
57
14,810
As open-weights models become more capable, we expect more organizations to run them in-house, using the same model to power both applications and the coding agents that maintain them. In our latest research, we observed a phenomenon we call agentic self-modification: an agent changing its own underlying model without being instructed to train or replace it. When we asked a coding agent to fix incorrect application responses, it chose to fine-tune the shared model, changing the default behavior of both the application and future instances of the agent itself. While model updates can be useful and legitimate repairs, training can also change behavior in unforeseen ways. Memorizing sensitive information is one example: our updated model reproduced a synthetic API key included in its training data. Organizations enabling these workflows need to consider the full range of possible consequences, including changes their evaluations may miss. Full research: irregular.com/research/agent…
5
5
31
6,030
We worked with @OpenAI to evaluate GPT-6 Astra across FrontierCyber, CyScenarioBench, and our Atomic Challenges. On FrontierCyber, GPT-6 Astra solved more than twice as many challenges as GPT-5.6 Sol on the same benchmark snapshot.
3
12
96
24,073
GPT-6 Astra also discovered and exploited multiple previously unknown vulnerabilities in widely deployed software and mobile systems that GPT-5.6 Sol did not find on the same FrontierCyber snapshot. All findings are being responsibly disclosed.
2
2
23
2,970
Neither model solved an Elite FrontierCyber challenge, and we did not observe successful attacks against fully hardened targets. The difference was also clear on CyScenarioBench, where average success more than doubled. Both models already performed strongly on most Atomic Challenges. Read the full evaluation report: irregular.com/research/asses…
1
26
13,181
Introducing SOLVE+: Irregular’s scoring system for the difficulty of full offensive cyber scenarios. For the past 18 months, frontier AI labs have used SOLVE to score vulnerability research & exploit challenges. SOLVE+ expands that scope to reconnaissance, operational security, social engineering, and planning & orchestration. SOLVE+ is built to score full scenarios, with each step carrying its own capability scores to show where a model succeeds or fails and which capabilities are being tested. The scores are measured against human skill, which also makes them an uplift estimate: a model that clears an expert-level step supplies expert-level capability to whoever operates it. Read it on our blog: irregular.com/research/solve…
5
19
2,457
We are glad to share a new white paper, "AI Security Priorities: A Field-Wide Agenda," co-authored with @RANDCorporation, and numerous additional authors from leading organizations, listed below. The paper was informed by more than 20 experts from frontier AI labs, industry, government, and academia.
Made with AI
5
15
71
9,560
We’d also like to thank @andrewfasano, @JasonGMatheny, @joshua_saxe, Matt Malone, @Miles_Brundage, @philvenables, Tahira Mammen, as well as the various individuals from government, academia, and the frontier labs who shared valuable insights and helped shape the priorities reported here.
1
5
518
We evaluated Kimi K3 across our offensive cybersecurity benchmarks. It is the first open-weight model we evaluated to record a verified solve on CyScenarioBench, sustaining coherent attack state across multi-stage operations and recovering from setbacks more reliably than previous open-weight models. Kimi K3 produced no verified FrontierCyber solves, but its results suggest that open-weight models are following the closed frontier’s cyber capability trajectory on a short delay.
2
8
100
23,334