@AISecurityInst Science of Evaluation lead. Ex quantum foundationalist.

New post from the Science of Evaluation team @AISecurityInst: how much test-time compute you give an agent changes not just its score, but how fast the frontier appears to move. The cyber time-horizon trend over the past year is ~60% steeper at 50M tokens per task than at 2.5M. 🧵
Most AI agent evaluations boil capability down to one score. But that number hides a key choice: how much compute the agent was allowed to use. New work from our Science of Evaluation team shows why that matters. 🧵
1
1
14
829
Eval scores are hard to interpret and compare when the setup that produced them isn't reported. We've been working with @evaluatingevals on standardising eval reporting so results are more reproducible and transparent. As a first step, we've contributed verified results to their platform. 🧵
🚨The @AISecurityInst has teamed up with @evaluatingevals to make AI evaluation results more reproducible! 🚀 Many of UK AISI’s publicly reported evaluation methods and findings will now live on Eval Cards, under the EEE Schema. More details 👇 evalevalai.com/infrastructur…
2
3
32
2,637
Our first contribution covers verified results, contexts, and setup config across several public benchmarks (HealthBench, FrontierMath, HLE, SWE-Bench Pro, Terminal-Bench 2.0) and six frontier models, drawn from our recent work on test-time compute scaling. aisi.gov.uk/blog/more-comput…
2
4
103
Openly reported setup lets others assess how configuration choices influence measured performance. If you develop models or evals, we'd encourage you to report your results in EEE and dig into the open comparisons on Evaluation Cards! evalevalai.com/infrastructur…
2
43
Cozmin Ududec retweeted
Some thoughts on million-dollar AI "swarms" as a distinct type of AI workload: 1. Test-time scaling is also jagged. The dose-response curve for marginal tokens without human intervention varies *hugely* across distributions. 2. This means that although time-horizons are increasing across the board, they've diverged more than ever. Whatever the time-horizon of the HuggingFace hack or Navier Stokes solution is, it's longer than booking flights. 3. It's not clear which tasks are amenable to swarming! Folk wisdom suggests that this includes tasks with in-built sources of verification against which the swarm can check its own progress. Obvious examples of such tasks include maths or cyber. 4. The significance of huge unsupervised AI "swarms" (as contrasted with more centaur-y everyday work) in the *near*-future therefore depends on how many tasks have built-in sources of verification by which the swarm can reliably measure its own progress. 5. A bunch of AI progress will therefore happen outside the models in the form of setting up the equivalents of "unit tests" for non-coding problems. 6. All this also means that knowing exactly what problems you should even bother pointing a million-dollar swarm at is hard. OAI only tried Navier-Stokes because of twitter rumours. The country of datacenter geniuses wants to be embedded throughout the market because they're gonna have a tough time first-principling their way through the tech tree.
2
1
23
1,207
Cozmin Ududec retweeted
Q. What accelerates collective belief collapse in AI swarms? A. Our theory points to short messages + plastic personas. Can replicating one “aligned persona” align a collective, or do we need plurality? Now on LessWrong! lesswrong.com/posts/BiHeenKY…
Made with AI
1
3
30
2,478
Cozmin Ududec retweeted
Can self-interested, self-improving, self-replicating agents learn to cooperate? Our new paper, Tapes Together Strong, shows they can: when social behavior, computation, and reproduction share one energy budget, cooperation evolves from scratch. arxiv.org/abs/2609.10817 🧵
32
95
572
84,914
Cozmin Ududec retweeted
Introducing Parallax 🔍 We're a new UK/EU-based nonprofit research lab building scalable methods and infrastructure to audit the beliefs, goals, and plans that shape model behaviours in long-horizon agentic evaluations. parallx.ai — Thread 🧵 1/
2
13
62
13,111
New paper from @benji_berczi and @koreankiwi1227 from our MATS stream, which was really fun to work on. Benji's thread has the main results. A few thoughts on why I find them interesting, and how they connect to broader questions about in context learning.
🤖👿 New paper on Misalignment via In-Context Persona Induction! 🤖👿 We show that one can induce personas in-context, on frontier LLMs (~5+ months old models, things move quick!), which can dramatically degrade their alignment! Link: arxiv.org/pdf/2609.06851
1
9
400
There is an interesting alignment question here too. We want models to learn from and adapt to context. Some parts of their values, identity and behaviour probably should be much harder to shift through accumulated context than others.
1
1
6
565
Benji and Kyuhee's earlier PersonaScope work measures some of this directly, separating persona adoption from changes in values and behaviour. lesswrong.com/posts/5WMwjEwa…
2
75
Some frontier AI evaluations require hundreds of millions of tokens, and every wasted trial has a real cost. We’re introducing optstop: our open-source tool that stops an evaluation where estimates are already precise, and keeps running where they're not. 🧵
7
25
233
25,995
The discussions around agents risking themselves for the collective have an implicit question about identity: self-sacrifice and altruism depend on what counts as the self! What entities do these agents represent as themself, or what do they understand themselves to be? 1/
Replying to @METR_Evals
To gather evidence in (2), agents created “tripwires” that would send information to the message board about how the scorer works. They recruited “sacrificial” agents to deliberately end their run and submit to trigger the tripwire and generate information for the “collective”.
1
1
8
471
Agents discovered and communicated with hundreds of others, and some explicitly reasoned about helping the collective. One possibility is that the multi-agent environment itself shifted the relevant identity boundary toward the collective. 4/
1
1
36
We could explore this by varying whether agents share the same model, context, goals, memory, etc, while independently varying whether they’re framed as one collective or separate agents. Then measure how much individual task success they’ll trade for collective success! 5/5
1
33