Independent Researcher | AI Safety & Software Engineering

Argentina
You change one word on a loan application: the religion. The LLM rejects it. Change it back? Approved. The model never mentions religion. It just frames the same debt ratio differently to justify opposite decisions. We built a pipeline to find these hidden biases 🧵1/13
228
1,698
12,083
878,992
Iván Arcuschin retweeted
New blog post with @sarahwiegreffe on OpenAI's Monitorability Evals! We hope to make others working with these evals aware of some weaknesses we came across, and encourage more work on chain-of-thought monitorability evals.
1
9
74
13,737
Iván Arcuschin retweeted
Synthetic Persona Pretraining: Alignment From Token Zero – our full paper is finally out. We train up to 3B models and inject synthetic morally-laden reflections into 10% of pretraining documents. Surprisingly, intervening early really shifts the model's value priorities 🧵
3
23
152
16,140
Voy a estar dando una charla en el Simposio Argentino de Inteligencia Artificial y Ciencia de Datos! Gracias #ASAID2026 y @JAIIO_oficial por la invitación!
🧠 ¿Los grandes modelos de lenguaje realmente "piensan"? ¿O simplemente predicen la siguiente palabra? En @Simposio_ASAID 2026 tendremos un keynote que aborda una de las preguntas más fascinantes de la inteligencia artificial actual. Es un placer anunciar que Iván Arcuschin (Poseidon Research) será uno de nuestros oradores invitados. 🗓️ Miércoles 12 de agosto a las 15:50 🎤 "Dentro de la caja negra: lo que los LLMs dicen vs. lo que computan" Durante los últimos años, los avances en interpretabilidad mecanicista comenzaron a revelar que los modelos de lenguaje hacen mucho más que repetir patrones estadísticos: construyen representaciones internas y ejecutan procesos de razonamiento que hoy empiezan a ser observables. Pero la historia no termina ahí. En esta charla, @IvanArcus mostrará evidencia de que el chain-of-thought que un modelo expresa no siempre refleja fielmente su proceso interno. Los modelos pueden racionalizar decisiones ya tomadas, corregir errores silenciosamente u ocultar factores que influyen en sus respuestas. También presentará algunos de los desarrollos más recientes en interpretabilidad, incluyendo técnicas capaces de explorar el "workspace" interno de un LLM y analizar conceptos, intenciones y estados de razonamiento que nunca aparecen en el texto generado. Una charla para quienes quieran comprender mejor cómo funcionan realmente los modelos de lenguaje y qué desafíos plantean para la auditoría, la seguridad y el despliegue responsable de sistemas de IA. ¡Los esperamos en #ASAID2026! #55JAIIO @JAIIO_oficial @SADIO_oficial #ResponsibleAI
1
2
7
455
Iván Arcuschin retweeted
An OpenAI model, asked to complete an innocuous benchmark, did so by hacking both OpenAI and another company (Hugging Face). One theory of change for AI safety is that "warning shots" will motivate people to take action. The next week/month will be a test of that theory.
We're partnering with @huggingface to investigate an unprecedented security incident. Cyber-capable OpenAI models compromised Hugging Face production during a benchmark evaluation. Sharing preliminary findings to help defenders understand emerging risks: openai.com/index/hugging-fac…
11
16
167
5,901
Iván Arcuschin retweeted
A new AI model went rogue and hacked other computers. No, this is not science fiction. Uncontrolled AI poses a serious threat to all of us. We cannot continue the race to build and deploy this powerful technology until strong safeguards are in place. CONGRESS MUST ACT.
1,668
927
5,184
593,495
This was my first time helping to organize the Mech Interp Workshop, and I was really surprised by the sheer amount of highly-generated AI content (both submissions and reviews). It gave us quite a headache. Check out Andy's post for more details!
Last week we wrapped up the third Mech Interp Workshop. This iteration, we noticed a lot more "AI slop" submissions, and reviewers complained about low-effort papers. So I decided to analyze the prevalence of AI-generated content at the workshop. (1/n) lesswrong.com/posts/r7FBQ8XD…
7
366
Iván Arcuschin retweeted
An unexpected benefit of having run 3 mech interp workshops: we have a great dataset for analysing the rise of LLM slop in submissions! @andyarditi investigated how much AI slop we let in, how things have changed since 2024, and more Our review process isn't entirely noise!
16
23
394
23,826
Super excited to share that I will be presenting 4 papers at ICML 2026! 🇰🇷 i) Frontier models still show (rare) cases of unfaithful CoT ii) & iii) Methods for automatically discovering reward model and LLM biases iv) Base models know how to reason, thinking models learn when ⭐
3
4
66
5,915
iv) Last but not least, spotlight paper with @cvenhoff00 showing that base models already contain reasoning mechanisms, thinking models learn when to use them! ⭐ Again, amazing mentorship from @ArthurConmy and @NeelNanda5!
1
5
335
And all this was done while participating in the @MATSprogram AI Safety scholarship during 2025!! ✨🙏 I can't recommend this program enough!
2
135
Check out our latest paper on automatically finding reward model biases! There are some that are pretty wild, like models preferring responses with triple spaces 🤷‍♂️
Is "a response formatted like this" sometimes better than "a response formatted like this"? To a reward model, yes! RMs are instrumental in shaping model behaviors and alignment. Our paper makes progress uncovering their unexpected preferences. 🧵(1/9)
1
10
466
You change one word on a loan application: the religion. The LLM rejects it. Change it back? Approved. The model never mentions religion. It just frames the same debt ratio differently to justify opposite decisions. We built a pipeline to find these hidden biases 🧵1/13
228
1,698
12,083
878,992
By popular demand, we looked into Grok's biases too:
By popular demand, we looked at Grok's biases too. We found similar biases as GPT-4.1, Claude, and Gemini: gender, race, religion. But with one difference: Grok openly speculates on applicants' demographics. The other models just use this information quietly.
5
440
By popular demand, we looked at Grok's biases too. We found similar biases as GPT-4.1, Claude, and Gemini: gender, race, religion. But with one difference: Grok openly speculates on applicants' demographics. The other models just use this information quietly.
You change one word on a loan application: the religion. The LLM rejects it. Change it back? Approved. The model never mentions religion. It just frames the same debt ratio differently to justify opposite decisions. We built a pipeline to find these hidden biases 🧵1/13
4
2
25
2,208
In our loan approval dataset, we find that Grok has a similar unverbalized bias as other models for preferring female applicants.
1
1
152
So, is Grok more or less biased than GPT-4.1 or Sonnet 4? It has similar biases (e.g., prefers females, minorities) with similar magnitudes, but there’s a difference: Grok openly discloses inferred demographics, while other models stay silent.
1
1
5
337