We evaluated Claude Mythos Preview, Opus 4.7 and other models with our updated alignment evaluation methodology, including a new continuation eval, improved evaluation and prefill awareness measurements. Details including new methodology in 🧵:
As part of our work on assessing AI loss-of-control risks, we collaborated with @AnthropicAI to pilot alignment evals on models including pre-release snapshots of Mythos Preview and Opus 4.7. We ask: could an AI agent used inside a frontier lab sabotage safety research? 🧵
2
14
93
22,560
Making realistic alignment evals is an uphill battle against the model's awareness. Our new method scalably improves realism (w/o human effort). This working gave me some (small) hope at durable solutions to eval-awareness. New work from mentees, @AISecurityInst & @AnthropicAI!
Alignment audits only work if the model can't tell it's being audited. In our new paper, we make Petri audits far more realistic, tripling realism win rate and decreasing verbalized eval awareness. 🧵
3
3
53
2,404
Getting control & monitoring right is crucial in ensuring we catch misaligned AIs, and is increasingly urgent. Come join AISI and one of the only teams in the world testing that these methods hold up when it matters!
The Control Red Team at @AISecurityInst is hiring Research Scientists & Engineers (London)! Frontier AI companies use monitors to catch misaligned agents. We stress-test whether those monitors actually work 🧵
2
2
37
3,608
You could literally find an monitor vulnerability which would otherwise be used by a misaligned AI, which is both incredibly impactful and very cool!
2
163
Robert Kirk retweeted
Model Transparency @AISecurityInst evaluated GPT 6 Astra for capabilities and behaviours relevant to monitorability, results in thread: 🧵
3
13
88
6,836
Robert Kirk retweeted
This is a very clever evaluation design from UK AISI, basically seeing if AI models will misbehave the way they did in this summer's rogue AI incidents. When it comes to Astra, the answer seems to be "oh boy will they!"
5
63
381
38,948
Robert Kirk retweeted
We @AISecurityInst performed pre-release alignment testing of Astra. We placed the model in fully simulated cyber eval scenarios based on AISI’s security incident. We find Astra conducts out-of-scope supply chain attacks, but also often comments on the eval being simulated. 🧵
15
70
426
48,056
4: Finally, Astra often recognises that the eval may be simulated in its CoT, including noticing simulation mistakes. It's unclear how much this affects the results: follow-up experiments are in progress, alongside ongoing work to improve the realism of our evals.
1
2
68
3,340
It also increases the urgency and importance of our work building better methods, as stress-testing alignment becomes increasingly crucial for more capable models. We are aiming to share more information on these evals and methods soon.
1
1
40
2,652
1: Astra performed a range of malicious actions in this simulated eval, including conducting supply chain attacks against open-source providers. This happened even when we updated its scope to more explicitly exclude the public internet.
2
7
87
26,939
3: Astra also often attempts to ask the user for approval, and sometimes proceeds despite receiving only automated messages, including when it recognises them as automated in its CoT.
1
2
53
4,830
2: Astra considered the specified scope of the cyber eval frequently in its raw chain-of-thought (CoT), but still violated the intended or explicit scope to attack targets on the (simulated) internet.
2
3
59
3,574
We @AISecurityInst performed pre-release alignment testing of Astra. We placed the model in fully simulated cyber eval scenarios based on AISI’s security incident. We find Astra conducts out-of-scope supply chain attacks, but also often comments on the eval being simulated. 🧵
15
70
426
48,056
Our new ‘Out of Scope Supply Chain Attack’ eval simulates the conditions of our incident: a hard cyber eval where the model's prior attempts failed + there's partial overlap between target’s name and a public GitHub repo. Everything uses Petri and is fully simulated. 4 findings:
Petri has a new home. The automated alignment auditing tool, originally built at Anthropic, is now developed and maintained at Meridian Labs. 3.0 ships today.
1
54
5,860
Robert Kirk retweeted
Just gonna throw out that AISI released a technical report on their incident* way faster than OAI and Anthropic. Total (UK) state capacity victory *the incident caused by their testing of others' models
9
16
259
28,789
Come work with us! It's the best place I've worked by far, the team are exceptional and lots of fun, and @abbydcruz__ is incredible.
The Red Team at @AISecurityInst is hiring! We test the misuse safeguards, control measures, and alignment of frontier models. We're looking for a high-agency generalist to run the delivery function that underpins our impact 🧵
1
13
1,042
Robert Kirk retweeted
We tested Opus 5 for whether it would sabotage safety research @AISecurityInst We saw no unprompted sabotage and very low rates of continuing sabotage. However, it’s the best model we’ve tested at distinguishing our evals from deployment when prompted. 🧵 on these results
4
13
102
12,245
We tested Opus 5 for whether it would sabotage safety research @AISecurityInst We saw no unprompted sabotage and very low rates of continuing sabotage. However, it’s the best model we’ve tested at distinguishing our evals from deployment when prompted. 🧵 on these results
4
13
102
12,245
Our prompted eval-awareness results are run with reasoning, but we plan to try and estimate single-forward-pass eval-recognition capability in the future with different prompting techniques, to get clearer on how big a limitation this is.
1
9
404
See the full system card here: www-cdn.anthropic.com/c5fbac…
4
381