Associate Professor @ University of Oxford. AI Assurance & Multi-Agent Security. PI wittlab.ai

Oxford, England
Fantastic, systematic work on inter-agent influence just accepted to NeurIPS'26 - congrats to lead @ChandlerDSmith , @prpaskov as well as senior author @lrhammond and the others
Right now we learn how AI agents behave together after the fact, through hacks, logs, and headlines. Soon agents will interact with other agents on our behalf. Commerce will rely on those exchanges. How easily can agents persuade, deceive, and coerce each other? Better to find out in evaluations than in the logs. We started measuring. NeurIPS 2026, paper coming soon With @ChandlerDSmith Qi Guo @tilli_cecilia @SophiaHatz4 @DavidDAfrica @lrhammond @casdewitt @philiptorr
2
14
781
Christian Schroeder de Witt retweeted
Right now we learn how AI agents behave together after the fact, through hacks, logs, and headlines. Soon agents will interact with other agents on our behalf. Commerce will rely on those exchanges. How easily can agents persuade, deceive, and coerce each other? Better to find out in evaluations than in the logs. We started measuring. NeurIPS 2026, paper coming soon With @ChandlerDSmith Qi Guo @tilli_cecilia @SophiaHatz4 @DavidDAfrica @lrhammond @casdewitt @philiptorr
2
1
32
1,950
Super excited about co-hosting the AI4GOOD workshop at @NeurIPSConf in Paris - there will be a dedicated multi-agent security/safety track as well in collaboration with @krawiecka_kl @swapneel_mehta and others - the submission deadline is approaching fast!
🚨 8 days left to submit to the Trustworthy AI for Good (AI4GOOD) Workshop at #NeurIPS2026 in Paris We’re still accepting submissions. Join us in bridging AI safety, social good, and real-world impact, so that more capable models actually help society at scale. Our workshop is non-archival, and we will recognize outstanding papers and top reviewers with awards. Deadline: August 29, 2026 (AoE) More details and submission → trustworthy-ai-for-good.gith… 📧 Sponsorship & questions: zjingchen@cs.toronto.edu Let's bridge trustworthy AI and real-world impact. We hope to see you all in Paris. @TerryJCZhang @ChanglingXavier @_agirlyengineer @ozzaney0101, He (Shawn) Shuang, Jerick Shi, Prakhar Gupta, Kexin Li (Cassie), Wenjun (Wendy) Qiu @ZhijingJin @radamihalcea @MilindTambe_AI @david_lie @casdewitt @VectorInst @JinesisLab @EuroSafeAI @MPI_IS @UofTCompSci @TorontoSRI @CIFAR_News @ELLISInst_Tue @UMichCSE @michigan_AI @Harvard @ETH_en @CarnegieMellon @UniofOxford @NeurIPSConf @UofT
1
13
910
Christian Schroeder de Witt retweeted
🚨 8 days left to submit to the Trustworthy AI for Good (AI4GOOD) Workshop at #NeurIPS2026 in Paris We’re still accepting submissions. Join us in bridging AI safety, social good, and real-world impact, so that more capable models actually help society at scale. Our workshop is non-archival, and we will recognize outstanding papers and top reviewers with awards. Deadline: August 29, 2026 (AoE) More details and submission → trustworthy-ai-for-good.gith… 📧 Sponsorship & questions: zjingchen@cs.toronto.edu Let's bridge trustworthy AI and real-world impact. We hope to see you all in Paris. @TerryJCZhang @ChanglingXavier @_agirlyengineer @ozzaney0101, He (Shawn) Shuang, Jerick Shi, Prakhar Gupta, Kexin Li (Cassie), Wenjun (Wendy) Qiu @ZhijingJin @radamihalcea @MilindTambe_AI @david_lie @casdewitt @VectorInst @JinesisLab @EuroSafeAI @MPI_IS @UofTCompSci @TorontoSRI @CIFAR_News @ELLISInst_Tue @UMichCSE @michigan_AI @Harvard @ETH_en @CarnegieMellon @UniofOxford @NeurIPSConf @UofT
We are so excited to host the 2nd edition of the Trustworthy AI for Good (AI4GOOD) Workshop at #NeurIPS2026 in Paris, France 🇫🇷 We aim to bring together the AI safety, AI for social good, and AI policy/governance communities to connect what models can do as individual systems with what they do when deployed across populations. For full details, welcome to visit our website at: 🔗 trustworthy-ai-for-good.gith… 📍December 12, 2026 · Paris 🇫🇷 We have an amazing lineup of speakers and plans. More details 🧵
1
5
16
4,016
Proud to help launch Orbit v0 — a framework for multi-agent evaluations based on @AISecurityInst's Inspect, starting with security but built for safety and capability research too. Huge team effort, and excited to see where this goes. Hugecongrats to @BenHagag20 @wlanderson0 who have been developing this as part of my @MATSprogram, and @csrijac for mentoring. Stoked @coop_ai will be stewarding this going forward.
Today we're releasing Orbit v0: a framework for multi-agent safety and security research. If you've tried running multi-agent experiments, you know the drill - weeks of infra before your first result. Built on Inspect, Orbit lets you swap environments, threat models, and topologies without rebuilding anything.
1
2
14
1,916
Christian Schroeder de Witt retweeted
Work with @wlanderson0, @casdewitt, Sarah Scheffler as part of @matsprogram. We’re presenting ‘Architecture Matters for Multi-Agent Security’ at @icmlconf this week in Seoul, come find us!
1
1
2
271
Highly recommend @sumeetrm and @CharlieLondon02's upcoming talk on long-horizon reasoning in LLMs - this has been an exciting ride
If models can think for 100,000 tokens, why do they still lose the plot? Come join us for this AI4Science on alphaXiv talk: Long-Horizon Reasoning in LLMs. In this session, Sumeet Motwani (@sumeetrm) and Charles London (@CharlieLondon02) will share recent work on both training and evaluating models that can reason over much longer chains of thought. Their LongCoT benchmark tests whether models can handle long chains of dependent reasoning across different fields. Each step is solvable on its own, but the full problem requires planning, state tracking, backtracking, and avoiding compounding errors. Even the best models still score below 10%. They will also discuss h1, which trains long-horizon reasoning by chaining short problems into longer dependency graphs, then using RL with outcome-only rewards and a gradually harder curriculum. So if longer context windows are not enough, what does it actually take to make models reason reliably over long scientific and technical workflows? Whether you’re working on frontier LLMs, AI4Science, reasoning, or just curious about what current models still cannot do, you should definitely check this talk out! 🗓 Friday May 15th 2026 · 11 AM PT 🎙 Featuring Sumeet Motwani and Charles London 💬 Casual Talk + Open Discussion
2
764
Honoured to serve as Area Chair at NeurIPS 2026. @NeurIPSConf
2
49
5,087
Christian Schroeder de Witt retweeted
New mini experiment + blogpost + trajectories! tldr; we boost performance of RLM(GPT-5.2) to double the best performing number (38.7% --> 65.6%) on LongCoT-mini without any training! An example of the mismanaged geniuses hypothesis (MGH) we (@zli11010, @lateinteraction) proposed earlier this month. The LongCoT benchmark showed that frontier LMs and RLMs struggled to solve difficult compositional reasoning tasks. The paper generally attributes this to the RLMs inability to perform task decomposition, but we argue this is more our fault in how we prompt them; this capability is fully available to GPT-5.2 with an RLM harness! Building on @raw_works's insightful blogpost and @sumeetrm / @CharlieLondon02 et al.'s incredibly useful benchmark, where they originally found RLMs to be incapable of solving the MATH and CS splits altogether. We did not train anything since the release of the initial benchmark. To be fully transparent, these results are not meant to be added to their leaderboard either; benchmarks measure isolated capabilities, and we focus on showing (through different, rather specific prompting) that the capabilities required to solve these tasks are available to the models without additional training! It also has implications about how we would go about training these systems. Full blog below, it's a nice read :)
18
64
482
43,622
Christian Schroeder de Witt retweeted
LLMs will supposedly solve climate change and cure cancer, but in fact they can't even do multi-turn reasoning tasks effectively (SOTA models are < 10% on this benchmark). Interestingly, this work directly compares how much extra performance you get when you add an agentic harness (figure 7): a lot for simple optimization problems, 0% for math and chemistry.
10
12
103
21,893
Christian Schroeder de Witt retweeted
How can we test the "intrinsic" long-horizon reasoning capability of a model? We made a neat template-based problem construction, where each subproblem is easy, but their composition primarily makes any problem hard. Also avoids test saturation by scalable problem difficulty!
1
3
12
920
Proud to release LongCoT, a hard benchmark for long-horizon reasoning capabilities - measuring reasoning over hundreds of thousands of tokens. 🥳 Project led by my student @sumeetrm in collaboration with many others; excited about kicking off Oxford Witt Lab's collaboration with Ruben Glatt @Livermore_Lab
2
12
1,762
Christian Schroeder de Witt retweeted
Training multi-agent teams is hard. #AgentFlow comes to the rescue. We introduce Flow-GRPO, an efficient method to train multi-agent teams. Improves planning and tool use. Selected as an #ICLR2026 Oral (top 1%)🚀
🔥Introducing #AgentFlow, a new trainable agentic system where a team of agents learns to plan and use tools in the flow of a task. 🌐agentflow.stanford.edu 📄huggingface.co/papers/2510.0… AgentFlow unlocks full potential of LLMs w/ tool-use. (And yes, our 3/7B model beats GPT-4o)👇 🧩A team of four specialized agents coordinates via shared memory: Planner: plan reasoning & tool calls 🧭 Executor: invoke tools & actions 🛠 Verifier: check memory status ✅ Generator: produce final results ✍️ 💡The Magic: 🌀💫 AgentFlow directly optimizes its Planner agent live, inside the system, using our new method, Flow-GRPO (Flow-based Group Refined Policy Optimization). This is "in-the-flow" reinforcement learning. 📊The Results: AgentFlow (7B backbone) outperforms top baselines on 10 benchmarks, with average gains of: +14.9% on search 🔍 +14.0% on agentic 🤖 +14.5% on math ➗ +4.1% on science 🔬 🏆It even surpasses larger-scale models like Llama-3.1-405B and GPT-4o (~200B). Try it yourself! 🛠️Code: github.com/lupantech/AgentFl… 🚀Demo: huggingface.co/spaces/AgentF… 🤖Model: huggingface.co/AgentFlow/mod… 📊Visual: agentflow.stanford.edu/#visu… 💬Join our Slack: join.slack.com/t/agentflow-c… #agentic #llms #RL #tooluse
2
44
202
27,979
New work led by @aaronrose227 showing how to do interpretability in multi-agent settings
New paper: Detecting Multi-Agent Collusion Through Multi-Agent Interpretability LLM agents can secretly collude, even inventing steganographic signals that text monitors can't catch. We show you can detect this from their activations. w/@casdewitt 🧵 (1/n)
29
4,116
While we cannot always detect steganography directly, sometimes the effects of sharing information secretly can be observed relative to the subsequent behaviour of the agents - an important decision-theoretic approach to steganography detection in CoT settings pioneered by @usmananwar391 @j_piskorz_
✨New AI Safety work on Steganography and LLM monitoring✨ We propose ‘steganographic gap’: the first principled metric for detecting and quantifying encoded reasoning in LLMs, which can reveal hard-to-detect forms of steganography, e.g., paraphrasing-resistant steganography.
2
14
1,554
Christian Schroeder de Witt retweeted
The Red Team at @AISecurityInst is hiring! We work with frontier AI companies to red team their misuse safeguards, control measures, and alignment techniques. As the stakes rise, we need much stronger red teaming and many more talented researchers working within gov 🧵
2
34
233
74,234