Research Fellow @ Kempner Institute, Harvard | Working on AI interpretability, robustness & safety

Code for "Large Language Models Generate Harmful Responses Using a Distinct Mechanism, Shared Across Harm Types" is now available! It includes B-TAP, the interp method behind our work, for localizing and causally intervening on critical weights. Code + project page + paper 👇🏻
1
13
54
3,770
Check out our new work, to be presented at @COLM_conf! We find mechanistic evidence that creativity and hallucinations are tied together. Working on this made me realize that some of the same mechanisms that make LLMs useful may also create failures.
At the upcoming COLM 2026, we will present our new paper, “Creativity and Hallucinations Share Mechanisms in LLMs.” Using mechanistic interpretability, we investigate the connection between creativity and hallucinations and find evidence of shared mechanisms. 🧵
1
3
51
4,243
Hadas Orgad retweeted
Replying to @ziv_ravid
I don't want to waste my energy reading something you couldn't be bothered to write.
1
4
52
2,109
Come work with me! I’ll be mentoring in this fall’s @cbai_ai fellowship. If you’re interested in interpretability + AI safety, especially understanding the mechanisms behind sycophancy, deceptions, harmful generations and hallucinations, apply to my stream 👇
Applications are open for the CBAI Fall Research Fellowship in AI Safety & AIxBiosecurity. Apply by September 6th at 11:59 PM EDT! 📅October 13 - December 18 💵 $15,000 stipend for 10 weeks 💰 ~$1,500+/week compute support 🏠 Housing close to our offices, arranged by CBAI
6
18
200
10,552
It’s been interesting watching PhD students and junior researchers start working with coding agents and tools. It’s clear it can dramatically increase velocity. On the other hand, it’s much easier for things to get far off track.
12
7
206
11,579
Announcing the keynote speakers for the Actionable Interpretability Workshop at COLM 2026! @ActInterp @boknilev @banburismus_ @ChrisGPotts @dhanya_sridhar What would you ask them?
2
13
52
5,443
Accepted to #emnlp ✅ 🎉 Read about the (non-) robustness of uncertainty probes 👇🏼
Replying to @_joestacey_
This work has been really fun to work on! Massive thank you my fantastic collaborators @OrgadHadas @inuikentaro @benbenhh and @NafiseSadat You can find the paper here: arxiv.org/abs/2604.11662
1
47
4,529
Hadas Orgad retweeted
Never mind the ethics of AI use, we need to establish the etiquette. Inviting me to read LLM-generated text without telling me that's what it is, as though you'd written it yourself, is _rude_
31
289
3,398
49,200
Hadas Orgad retweeted
1/ Text-to-video models generate beautiful videos, but they still struggle to follow complex prompts. Relations like left/right, towards/away, on top of, behind, or even multi-stage instructions are often generated incorrectly. In our new paper, accepted to SIGGRAPH Asia 2026, we introduce CVG, an inference-time guidance method that improves compositional video generation without retraining or architectural changes.
8
8
32
76,880
Accepted to COLM 2026? Consider submitting your work to the fast track at the Actionable Interpretability Workshop! @ActInterp We welcome work that advances the use of interpretability to real-world usage. Deadline: August 9 Links below 👇
1
13
33
7,488
Happening now! Poster #3206
At Seoul for #ICML! We will present our position paper tomorrow on Actionable Interpretability! Come hang out and argue on how (and if) interpretability can be actionable. 📍HALL A #3206, 2:00 PM – 3:45 PM
1
2
26
2,052
Hadas Orgad retweeted
Here we go! ✈️ 🇺🇸🇰🇷 If you're attending #ACL2026 and/or #ICML2026 check out recent work from our group with collaborators: 📍Jul 5 14:00-15:30: Matan will present his TACL work "Universal Jailbreak Suffixes Are Strong Attention Hijackers" @matanbt @mahmoods01 arxiv.org/abs/2506.12880 📍Jul 5 16:00-17:30: My student Or will show how matrix factorization can reveal compositional neuron groups and hierarchies in MLPs. @OrShafran arxiv.org/abs/2506.10920 📍Jul 6 9:00-10:30: Tomer will talk about "MoNaCo: More Natural and Complex Questions for Reasoning Across Dozens of Documents" @TomerWolfson arxiv.org/abs/2508.11133 📍Jul 7 14:00-15:45: Hadas will present our position: "Interpretability Can Be Actionable" @OrgadHadas arxiv.org/abs/2605.11161 (will also appear at the MI workshop) 📍Jul 8 10:30-12:15: I will be presenting a position work with @_galyo and @ymatias on metacognition as a way to move beyond hallucinations arxiv.org/abs/2605.01428 📍Jul 9 17:00-18:45: Or (after a transpacific flight from ACL) will present MFA: an unsupervised approach to disentangle representations of LLMs at scale using local geometry @OrShafran arxiv.org/abs/2602.02464 (will also appear at the MI workshop)
11
85
3,443
At Seoul for #ICML! We will present our position paper tomorrow on Actionable Interpretability! Come hang out and argue on how (and if) interpretability can be actionable. 📍HALL A #3206, 2:00 PM – 3:45 PM
1
8
51
4,296
Replying to @ActInterp
You can submit any relevant work that wasn't already accepted to another venue, and in any case won't be presented before the workshop day (October 9th). COLM accepted papers will go through a fast track (details in the link) actionable-interpretability.…
1
406
We’re extending the Actionable Interpretability workshop @ActInterp submission deadline by 3 days! New deadline: June 24th. Looking forward to your submissions ;) Link in thread
2
7
33
2,531
📢 We’re looking for reviewers for the Actionable Interpretability workshop @ActInterp! If you’re interested in helping review submitted papers, please sign up here: forms.gle/VpLJpkM6zw3V8bX56 Your expertise would be greatly appreciated!
7
26
2,340
On LLM generated reviews
Replying to @delliott
Reviews can seem very detailed but in practice there's little information there. The summary is often full of fluff so it's really just hard to understand what the paper is about. They often provide long lists of issues, and it's very hard to understand if the concerns are major or minor. They often make non realistic or nonsensical suggestions. All this without mentioning mistakes and hallucinations, which are often conveyed with such confidence that can bias the whole evaluation. The final scores are often borderline so no information there as well. As author, it creates lots of work which is either stupid or not feasible. As AC it's a nightmare, I want to get a concrete evaluation to work with and I get this noisy not informative text. In rebuttals things become a joke.. the person who generated the review has no idea how to judge what's going on and so typically you get a short statement like "I read the response and decided to keep my score".
12
1,866