Alignment audits only work if the model can't tell it's being audited. In our new paper, we make Petri audits far more realistic, tripling realism win rate and decreasing verbalized eval awareness. 🧵
4
7
58
8,134
This work was carried out as a collaboration between Meridian Cambridge, @cbai_ai, @AISecurityInst and @anthropic, jointly authored by @_axelahlqvist, @richardguanxyz, @riverapablojuan, Adeline kassler, @mitroitskii, @alexandrasouly, @kaifronsdal, @_robertkirkk and @jplhughes
1
4
294
For coding settings we go further and run the target inside a real agent scaffold (Claude Code, Codex CLI, or Gemini CLI), so the system prompt, tools, and system reminders match deployment exactly. The auditor only simulates the user and tool results. We call this DISH.
1
2
267
Our methods compose. DISH improves realism substantially, and when combined with critique refinement gives the highest realism on every target we tested.
2
283