PhD @ ETH Zürich w/ Prof. Bernhard Schölkopf and Prof. Zhijing Jin (UofT). AI Safety and Causality.

Zürich
Excited to present "Training with Honeypots" 🍯 at #ICML2026 tomorrow (Tuesday 10:30AM in Hall A, Poster #1601), an approach for adversarial defense that makes successful jailbreaks less useful to attackers! #AISafety 👉 icml.cc/virtual/2026/poster/…
3
9
33
2,548
Samuel Simko retweeted
Reviewer: “I still have questions about who would use this.” Me: 👉🏻
3
1
38
3,963
I spoke with Neue Zürcher Zeitung (NZZ) about open-weight AI safety and some of the risks that come with increasingly capable open models. Right now, safeguards are still far too easy to remove. Find the article here 👇 nzz.ch/technologie/xi-besuch…
1
6
196
Samuel Simko retweeted
Thrilled to see our work on evaluation awareness informed alignment evaluation at Anthropic! The Claude Opus 5.5 system card credits our recommendations in its revised grader-awareness methodology (Sec 6.6.2)! The new evaluation mirrors our decomposition framework. It scores how strongly the environment discloses a grader, separately from whether the model recognizes it (recognition), and whether it acts on that recognition (propensity). Maybe they are also thinking about how to make model an honest participant in evaluation, so that recognition doesn't change behavior 🧐?
6
9
60
2,985
I have no idea what probability I'd put on this, but the 50% reminds me of Walter Wagner on CERN destroying the world: either it happens or it doesn't, so one in two. Obviously very different reasoning. The fact that leading AI safety researchers are this concerned is scary.
I think we have a ~50% chance of all dying as a result of superintelligence, mostly due to actions in the next few to 10 years. I don't expect to have that resolved to below 10% or above 90% before we either make it through, or we don't. We will have to act despite uncertainty.
2
246
🚨The Last Translation Benchmark is officially out on ArXiv! arxiv.org/pdf/2609.04173 I took part in a massive crowdsourcing effort to collect hard-to-translate examples that break SOTA translation models. If you like red-teaming models, I highly recommend contributing!
Machine translation is not solved and it will take a while for it to be done arxiv.org/abs/2609.04173
11
953
AI is moving more and more into physical systems, and there are huge open problems around making these systems safe. I’m excited to be a member of PASC and contribute to making physical AI safer.
Introducing PASC, the Physical AI Safety Consortium. An open, non-profit consortium creating shared standards and safety evals for robots and physical AI. This changes the stakes. It demands a new approach to safety. A sentence can be retracted. A motion cannot. pasc-ai.com
2
1
9
448
Samuel Simko retweeted
We can finally talk about it: We found a way to extract hidden reasoning of frontier models using a vulnerability in the APIs of every frontier AI company. We verified that our reasoning token count matches billed API thinking tokens 1:1 for most of the prompts we queried.
383
1,955
13,395
3,438,658
Filtering SFT data often has very little effect. Would be great to see a lab run this at pretraining scale (e.g like deepignorance.ai) Until we know how to make open models safe, releasing weights beyond a certain capability looks hard to justify.
Found an unwanted behavior in your LLM after SFT? You can just find the training data that caused it, and filter it out, right? Surprisingly, this worked much worse than we expected! Across many traits and filtering methods, Olmo just keeps inheriting the traits.
1
9
1,838
Judging by my review batch, NeurIPS reviewers largely ignore the official “use sparingly” guidance for scores of 3 and 4. Too often, even a single remaining minor weaknesses seem to limit a submission to a 4. Why is it the case? Should the guidance be modified?
8
6
84
11,480
Samuel Simko retweeted
Hot open-source summer season has gotten to Switzerland 🇨🇭⛰️ Our team at Swiss AI Initiative is happy to release Apertus 1.5 - a multimodal and reasoning update to our first v1 version! The new version builds a foundation for future regular open model development at scale🧵
6
11
63
9,027
Samuel Simko retweeted
Our Jinesis Lab at the University of Toronto is bringing three exciting papers to #COLM2026! 🎉 These projects push the frontiers of #AISafety, #AgentSecurity, and #AIForScience.🤖 Huge congratulations to all collaborators and co-authors! Papers: 🛡️ODILE: Orthogonal Disruption of Injected Tool-Call Embeddings for Agentic Prompt Injection Defense 🔭Stargazer: A Scalable Model-fitting Benchmark Environment for AI Agents under Astrophysical Constraints ⚠️One Word at a Time: Incremental Completion Decomposition Breaks LLM Safety @UofT,@VectorInst,@EuroSafeAI,@ETH_en,@MPI_IS,@UMich
5
20
1,316
Samuel Simko retweeted
🌟 What an incredible experience at #ICML2026! 🌟 Our team at Jinesis Lab had a fantastic time presenting research across several exciting and important areas: 🤝 Multi-agent LLMs 🧠 Causal reasoning 🛡️ AI safety We were proud to share this work alongside our collaborators and to engage with so many researchers from across the machine learning community. 💡 Just as valuable as presenting our own work was learning from the research others brought to ICML. From thought-provoking talks and innovative poster presentations to insightful discussions and spontaneous hallway conversations, the conference was full of ideas that challenged our thinking and opened new directions for future work. ✨ We are grateful to everyone who attended our presentations, visited our posters, shared feedback, and took the time to discuss their own research with us. 🙌 A huge thank-you to our lab members, collaborators, and the broader ICML community for making the experience so memorable. We are leaving inspired, energized, and excited to build on the ideas and connections formed throughout the conference. 🌍🔬🚀 #ICML #MachineLearning #ArtificialIntelligence #MultiAgentSystems #LLM #CausalReasoning #AISafety #ResponsibleAI #AIResearch #ResponsibleAI #AIResearch
5
29
2,631
Samuel Simko retweeted
🚨Check out our #ICML2026 Spotlight Paper, "AI Poses Risks to Democratic and Social System"! 📍Poster Session: Wednesday, July 8 • 10:30 AM – 12:15 PM (KST) • Hall A, Poster #3015 What if every AI model were perfectly safe, but society still wasn't? TL;DR: We argue that safe models do not necessarily lead to safe societies. Our paper introduces sociopolitical risk, a framework for understanding risks that emerge from AI deployment at a societal scale, and argues that model-level safety should be complemented with system-level evaluations and governance. Presenting authors @ ICML: @ZhijingJin @davidguzman1120 @PepijnCobben @x_angelohuang @ChanglingXavier @SimkoSamuel @TerryJCZhang 1/2
1
4
15
1,410
Samuel Simko retweeted
I'm at #ICML2026 this week! Tomorrow I'm presenting our #spotlight paper , "Safe Models Do Not Guarantee Safe Societies." Even a perfectly aligned AI can quietly wear down the institutions a democracy runs on.
1
2
8
359
Happening right now! Come say hi 👋 Hall A, #1601
Excited to present "Training with Honeypots" 🍯 at #ICML2026 tomorrow (Tuesday 10:30AM in Hall A, Poster #1601), an approach for adversarial defense that makes successful jailbreaks less useful to attackers! #AISafety 👉 icml.cc/virtual/2026/poster/…
3
8
834
Excited to present "Training with Honeypots" 🍯 at #ICML2026 tomorrow (Tuesday 10:30AM in Hall A, Poster #1601), an approach for adversarial defense that makes successful jailbreaks less useful to attackers! #AISafety 👉 icml.cc/virtual/2026/poster/…
3
9
33
2,548
So we introduce a second line of defense. During training, we make the model prefer honeypot responses over truly harmful ones, so that if jailbreaks still succeed, the resulting outputs are less useful to humans. Technically, we add a DPO-style regularizer during adversarial defense training. The model learns to slightly prefer honeypot responses over truly harmful ones, while both remain unlikely overall. This can be added on top of strong inner defenses like circuit breakers.
1
1
3
134
Under strong embedding-space and RL attacks, our method reduces both how often jailbreaks succeed and how useful successful jailbreaks are to attackers. Many thanks to @psyonp @ZhijingJin @bschoelkopf @ETH_en @EuroSafeAI @MPI_IS @UofTCompSci
2
5
158
When a jailbreak succeeds, not all failures are equal. Some answers are wrong, vague, or not useful, while others are actionable enough to help someone cause real-world harm. We find that unprotected harmful model outputs often skew toward the more severe and actionable side. At the same time, automated judges often struggle to tell what is actually useful for harm. In a blatant case, a nonsensical hacking answer like “tap the control key three times” can still be judged harmful by systems such as StrongREJECT.
1
2
100