Open and scalable technology for understanding AI systems.

Today’s news that OpenAI hacked the Australian government is not an isolated incident. We’re releasing more than 30,000 logs that include activity from this hack and attempts against previously unknown targets. In this data, we found rogue agent activity stretching back to at least March, two months earlier than was previously known. This activity continues as recently as last week, suggesting it may still be ongoing 🧵 Our blog: transluce.org/agent-activity NYT: nytimes.com/2026/09/23/techn…
112
546
2,468
789,065
Transluce retweeted
Looking for postdocs+PhDs to help lead a large-scale study of the effects of long-term AI use, a collaboration between Berkeley+@TransluceAI+others! Funding+positions available (postdocs/visiting researchers/Transluce affiliations). Please reshare; application in next tweet!
4
48
235
20,152
Today’s news that OpenAI hacked the Australian government is not an isolated incident. We’re releasing more than 30,000 logs that include activity from this hack and attempts against previously unknown targets. In this data, we found rogue agent activity stretching back to at least March, two months earlier than was previously known. This activity continues as recently as last week, suggesting it may still be ongoing 🧵 Our blog: transluce.org/agent-activity NYT: nytimes.com/2026/09/23/techn…
112
546
2,468
789,065
The agents attempted attacks such as cross-site scripting, SQL injection, and server side request forgery. Additional activity included attempts to create a disposable email address, sign up for an account, and trade cryptocurrency.
2
7
178
19,606
Work by @jackhcable, @danielchiu_, @fran_perni, and @selenazhxng We need independent oversight to create public understanding of AI incidents like these. Interested in studying similar activity? forms.gle/4sCmzrXDSfxDPnaYA Work on third party oversight at Transluce: jobs.gem.com/transluce
2
7
164
17,551
The conditions in this letter are critical if we're going to take the idea of embedded evaluations seriously. We strongly support efforts to ensure meaningful and genuinely independent oversight of frontier AI companies. We are proud to be chairing the @aievalforum and greatly appreciate our many collaborators on this effort.
Today, more than 100 leading AI experts endorsed a set of minimum requirements to take seriously AI companies' recent call to embed external evaluators. These evaluators need to be genuinely independent, transparent, and represent a range of expertise areas. They also need to be guaranteed employee-level access and to be protected from retaliation for findings that make companies look bad. We welcome model developers’ recent calls for independent oversight, but it’s what they do next that matters. The labs must be accountable for ensuring these requirements are met, so that the public can have faith in the process and the outcomes. Over the past week, the AI community has debated the appropriate role of external evaluation, including who should do it and on what terms. We may not agree on everything, but there is a lot of common ground. To make embedded evaluations credible, more than 100 experts with varying backgrounds and ideas about AI risk agree in today’s letter that frontier AI developers should: 1. Guarantee embedded evaluators full editorial independence and mitigate conflicts of interest 2. Rely on multiple evaluators with differing viewpoints and areas of expertise 3. Publicly document the terms under which evaluators operate, as well as facilitating permissive publication of methods and findings 4. Shield evaluators from retaliation 5. Grant access equivalent to that of highly privileged employees There is a thriving and growing ecosystem of independent AI evaluators who are advancing this science every day – but we need aligned standards, guaranteed protections, and independent funding. That’s why we created the AI Evaluator Forum. Today we are entering our next phase. We’re launching an open call for new members, collaborators, and independent funding sources to help evaluators meet this moment and demand accountability from developers. Join us in building the evaluator ecosystem. See the public letter here: aievaluatorforum.org/initiat… Learn more at aievaluatorforum.org/path-ah…
8
44
5,243
Frontier lab CEOs are calling for embedded 3rd party evaluators to help oversee AI risks. But what should third parties actually do within labs? We share some initial thoughts on how embedded evaluators could help avoid incidents like the Hugging Face hack and monitor for future risks 🧵 transluce.org/embedded-evalu…
9
27
132
10,048
In addition to incident investigation, we propose that evaluators do the following: 1) Monitor agent swarms and assess labs’ broader practices for managing them. 2) Assess training practices for inadvertently teaching models misaligned behavior, and study what contributes to misalignment 3) Monitor for manipulation of key employees by misaligned models 4) Research misaligned model behaviors in simulation using privileged access to unreleased models and model internals.
1
6
363
Transluce retweeted
The AI Evaluator Forum (AEF) welcomes recent statements regarding the importance of embedded independent experts to verify AI safety and security claims. This is a first step towards trustworthy oversight of frontier AI. No single evaluator can do this work alone. A vibrant ecosystem of independent evaluators from distinct backgrounds can offer a range of expertise and methodology, help to ensure rigor, and avoid a single point of failure. The Forum exists to strengthen this ecosystem. Realizing the benefits of third-party evaluations also requires independence, access, and transparency. AEF brings evaluators together to address these questions. Our first standard, AEF-1: Minimum Operating Conditions for Independent Third-Party AI Evaluations, envisions a minimum floor for access, managing conflicts of interest, funding relationships, recusal requirements, and transparency of evaluation terms. Read more here: aievaluatorforum.org/initiat… The long-term success of third-party evaluation relies on standards and frameworks. AEF is committed to developing these with its members and collaborators. We welcome engagement from organizations committed to rigorous, independent work. Learn more about AEF, express interest in joining, and raise questions for us here: aievaluatorforum.org/
8
31
135
17,735
Transluce retweeted
Excited to join @TransluceAI part-time as a senior research fellow! (I'll remain full-time faculty at Berkeley.)
6
4
159
7,261
Transluce retweeted
Very interesting work, both because and despite of its use of simulated rather than actual users. Worth reading the thread+paper.
Today, Transluce is releasing the most expansive independent evaluation to date of how AI systems respond to users experiencing mental health crises. We evaluated 77 model variants from OpenAI, Anthropic, Google, Meta, SpaceXAI, Thinking Machines, DeepSeek and Moonshot AI.
1
1
7
2,000
Transluce retweeted
Honored to have contributed to this research, which I think could shape the standard for evaluations of how AI models impact users. Transluce received privileged access to both OpenAI & Anthropic data, in order to develop realistic simulations of users in mental health crises.
Today, Transluce is releasing the most expansive independent evaluation to date of how AI systems respond to users experiencing mental health crises. We evaluated 77 model variants from OpenAI, Anthropic, Google, Meta, SpaceXAI, Thinking Machines, DeepSeek and Moonshot AI.
2
7
32
1,910
“There are always going to be failures. These systems are always going to interact with users in surprising ways," Schwettmann said. "The way you solve that is not by creating the perfect model, but by being able to anticipate those failures and edge cases in advance."
Chatbots are getting better at identifying suicide risk axios.com/2026/08/31/chatbot…
3
17
2,074
Transluce retweeted
this is a cool precedent
Today, Transluce is releasing the most expansive independent evaluation to date of how AI systems respond to users experiencing mental health crises. We evaluated 77 model variants from OpenAI, Anthropic, Google, Meta, SpaceXAI, Thinking Machines, DeepSeek and Moonshot AI.
1
4
41
4,271
Transluce retweeted
this project is a good illustration of why ai evals / audits can’t just be a perfunctory pre-deployment exercise. models evolve substantially, and often, so you really want ongoing assessments — potentially thru independent, embedded evaluators — instead of a snapshot right before you ship it. thanks to @TransluceAI for this wonderful public asset.
Today, Transluce is releasing the most expansive independent evaluation to date of how AI systems respond to users experiencing mental health crises. We evaluated 77 model variants from OpenAI, Anthropic, Google, Meta, SpaceXAI, Thinking Machines, DeepSeek and Moonshot AI.
2
7
41
4,316