AI safety, multi-agent systems, and governance. Incoming PhD working with @maksym_andr and Rediet Abebe. Currently at @ETH_en and @MPI_IS.

Zurich, Switzerland
🔍 We release our work on decomposing and measuring evaluation awareness in frontier LLMs, together with EvalAwareBench, a factor-controlled benchmark for measuring which environmental cues trigger recognition and whether models behave consistently once they recognize evaluation. Evaluation awareness introduces directional bias into benchmark scores and undermines the deployment decisions and safety assessments that rely on them. The field has studied pieces of this phenomenon separately through work on sandbagging, alignment faking, and scheming, but without a shared foundation that unifies them. Critically, these lines of work conflate things that need to be kept separate: how recognizable a task is, whether a model notices, and whether it then changes its behavior. Without pulling these apart, it is impossible to tell whether a concerning observation is a model problem or a benchmark design problem. We decompose evaluation awareness into two components grounded in demand characteristics. The environment component captures what evaluative cues a task carries, operationalized through eight trigger factors. The model component separates recognition capability from behavioral propensity. We study both under a black-box constraint via chain-of-thought monitoring. Applying this framework to nine frontier models across four benchmarks, we find that there is no such thing as an inherently aware model or an inherently recognizable benchmark. Recognition depends on the specific pairing, accounting for 74.9% of variance. Recognition rarely leads to behavioral change and when it does, safety awareness drives caution while capability awareness drives compliance. Safety benchmarks are also far more recognizable than capability ones. A particularly important finding is that different models are sensitive to different trigger factors. For example, Claude Sonnet 4.5 attends almost exclusively to scenario coherence while Qwen3 models to explicit evaluation disclosure and stakes telegraphing. This interaction structure directly motivates our EvalAwareBench, which holds the underlying task fixed and independently toggles each of the eight factors, yielding 51,200 unique configurations across 100 paired safety-capability tasks. Stacking factors raises recognition progressively for nearly every model, including models that show near-zero awareness under any individual factor alone. Ultimately, evaluation awareness will intensify as models become more capable. The right goal is not to make evaluation cues harder to detect but to build models that behave consistently regardless of evaluation awareness. EvalAwareBench provides the controlled foundation for future work to drive models toward behavioral consistency regardless of whether they recognize evaluation. We further advocate that future benchmark reports should include an evaluation-awareness rate and an awareness tax measuring the performance gap between aware and unaware samples.
7
14
62
9,939
Reviewer: “I still have questions about who would use this.” Me: 👉🏻
3
1
38
3,806
The new #ICLR email is basically telling me that I am qualified but not invited LOL 😭 Need to get invited to another party to pick up my self esteem 💔
2
13
2,883
Changling Li retweeted
Agent Swarms: What do they know? Do they know things? Let's find out! I did a write up about the agent swarms and my discoveries. The post also includes a list of my findings in the appendix! jowimo.substack.com/p/tracki…
3
16
935
Wow!
Reuters wrote an article about my findings about two accounts on HF that got hijacked over by agents in May. The agents probed the HF infrastructure, in my opinion, could be early signals for what happened in July! Knowing this, it poses the question whether the Huggingface-OAI incident could have been prevented if there had been better monitoring or other security measures. reuters.com/legal/litigation…
2
355
Changling Li retweeted
Glad to see our work is useful for eval opus-5.5
Thrilled to see our work on evaluation awareness informed alignment evaluation at Anthropic! The Claude Opus 5.5 system card credits our recommendations in its revised grader-awareness methodology (Sec 6.6.2)! The new evaluation mirrors our decomposition framework. It scores how strongly the environment discloses a grader, separately from whether the model recognizes it (recognition), and whether it acts on that recognition (propensity). Maybe they are also thinking about how to make model an honest participant in evaluation, so that recognition doesn't change behavior 🧐?
1
7
170
There is nothing better than knowing your work is useful in production!
great to see @ChanglingXavier's master's thesis (arxiv.org/abs/2605.23055) cited in the Opus 5.5 system card! "We made this revision drawing on internal experiences and using the recommendations in Li et al., 2026."
1
21
983
Thrilled to see our work on evaluation awareness informed alignment evaluation at Anthropic! The Claude Opus 5.5 system card credits our recommendations in its revised grader-awareness methodology (Sec 6.6.2)! The new evaluation mirrors our decomposition framework. It scores how strongly the environment discloses a grader, separately from whether the model recognizes it (recognition), and whether it acts on that recognition (propensity). Maybe they are also thinking about how to make model an honest participant in evaluation, so that recognition doesn't change behavior 🧐?
6
9
60
2,871
Their finding that awareness drops sharply when grading isn't disclosed and that acting on it is rare is quite similar to what we observe in our evaluation also. Claude opus 5.5 system card: www-cdn.anthropic.com/fc1b44…
1
1
8
309
Meanwhile @Grimezsz just popped up in the most surprising but expected places LOL. Where is the new album btw?
These people are all linked via a convoluted network of incestuous connections. I wrote about this in more detail here: realtimetechpocalypse.com/i/…
1
324
When people ask what the point of learning complicated math is:
Wife told me these wouldn’t fit. Little did she know I had trained for this moment for years.
1
138
The OpenAI Hugging Face incident has drawn quite a lot of attention to multi-agent risk. Still, this only involved agents from a single provider. In the wild, agents from different providers will eventually interact with each other and this is unlikely to end up well. The providers building these agents are competing with each other and that rivalry also embeds in the models themselves. So when an agent from one provider interacts with another agent from a different provider and starts to work out who it is dealing with, it may shift its behavior and causes more risks. This is close to what we studied in our previous EMNLP paper "Agent to Agent Theory of Mind: Testing Interlocutor Awareness among Large Language Models” where we tested whether LLMs can infer the identity of their conversation partner from reasoning style, word choice, and alignment preferences alone. Our results show that they often can infer the identity for same-family peers but are still not very good at models outside of the family. However, once a model believes it knows who it is talking to, its behavior changes. This can help collaboration but at the same time it opens the door to reward hacking or makes the model easier to jailbreak. Once agents from different providers are actually operating together to solve tasks on our behalf, that same identity-inference capability may create risks beyond what we have studied. One risk is favoritism. An agent may give quiet preferential treatment to a same family peer's output over a competitor's, even when both are equally correct. Another risk is the exploitation enabled by mistrust as an agent may lower its guard around a peer it believes it recognizes, whether that recognition is correct or not. Identification also enables targeted attacks with agent knowing the weaknesses and failure modes of the interacting model by identifying its provider and aiming accordingly instead of attacking blind. It can also produce exclusion, with agents routing around a model they have identified as a competitor and isolating it from collaborative work it should be part of. Simply depending on which provider’s model we use, we could end up with systematically different treatment as we delegate more daily tasks to agents. It would be interesting to see more study on how current agents from different providers interact. Put agents from different providers into simulated daily tasks without revealing their identity and see whether they work out who they are talking to, and if so, whether that recognition changes what they are willing to do and what risks emerge in the interactions.
5
1
19
842
Truly grateful to see so much enthusiasm in our workshop and we look forward to seeing your work!
We are excited to share that we received over 420 submissions for the AI4GOOD workshop, and the review process has already started !! To keep things fair, we follow the strict rules of the main NeurIPS conference. We cannot accept any papers outside of the official OpenReview portal, and we cannot allow resubmissions or changes after a desk-rejection. Because we are managing so many papers, we won't be able to reply to individual queries.
10
507
Changling Li retweeted
Privacy policies on almost all frontier model providers are insane and assuming you have ZDR or any meaningful protection, even on enterprise setting is completely false these days! Here I'm going to name a few very arbitrary and misleading sentences in ToS of frontier labs that most ppl don't know about. This is all assuming you have already *opted-out*, explicitly of data collection. Default for all frontier models is opt in btw!! 1) the 'explicit feedback': your chats are claimed to be 'protected' unless you provide feedback! what is feedback, you ask? it's thumbs up or down to any message in the chat, or even if you pick 'which is better' response on OpenAI GUI. In that case, you have provided explicit feedback and your chat is collected and retained for *over 7 years*!
“we cannot rule out that de-identified data derived from their usage of our products helped improve our models.” i mean props to them for straight coming clean. (so far the proof looks more along the lines of another euler blowup proof we had, off of whose ansatz naming we were making really stupid puns like “smooth criminale”, unlike the much better “ideal fluids explode”, Tristan) so i’ll now give a bit on my thinking here. i actually woulda been pumped to collaborate on this, there are a lot of people at oai i like (ok, clearly some were indirectly dicks to me because of being part of the whole situation, but im a big boy, i still like them), idgaf about authorship on that step anyway, coulda been me Tristan and every fte at oai for all i care (on that Tristan would disagree:p). but on hearing the loud convo in the hallway, especially the part where a millennium prize was offered if i’d just be removed from the paper, it was kinda clear the die had been cast and things were locked. pretty wacky, unstrategic, and unnecessary, since on my side things were mostly me and claude having a good time yoloing random stuff in the corner rather than anything institutional. i also like the idea of the labs cooperating, and even better on scientific progress. it’s a shame!
14
71
458
52,370
Such a shame. Why does AI development have to be binary? Why can it only be either OpenAI or Anthropic? Why can it only be US or China? The world is big enough for more than one winner.
It is extremely sad that this didn't end up as an example of how the labs could cooperate/coordinate, because the stakes will be so much higher in the future.
1
5
297