Open and scalable technology for understanding AI systems.

Filter
Exclude
Time range
-
Minimum likes
Transluce retweeted
Looking for postdocs+PhDs to help lead a large-scale study of the effects of long-term AI use, a collaboration between Berkeley+@TransluceAI+others! Funding+positions available (postdocs/visiting researchers/Transluce affiliations). Please reshare; application in next tweet!
4
51
236
20,321
Work by @jackhcable, @danielchiu_, @fran_perni, and @selenazhxng We need independent oversight to create public understanding of AI incidents like these. Interested in studying similar activity? forms.gle/4sCmzrXDSfxDPnaYA Work on third party oversight at Transluce: jobs.gem.com/transluce
2
7
166
17,603
The agents attempted attacks such as cross-site scripting, SQL injection, and server side request forgery. Additional activity included attempts to create a disposable email address, sign up for an account, and trade cryptocurrency.
2
7
180
19,657
Interestingly, the agents use exploits to complete what appears to be routine data retrieval tasks that are not cyber-related (for instance, searching for the average cost of skin and hair treatments in Australia).
4
16
257
47,595
We report three separate incidents between May and June 2026 in which agents attempted to hack the Australian Institute of Health and Welfare, DataUSA, and the University of New Mexico. We found direct links between the first two and a previously confirmed rogue agent swarm from OpenAI.
2
10
219
27,258
Today’s news that OpenAI hacked the Australian government is not an isolated incident. We’re releasing more than 30,000 logs that include activity from this hack and attempts against previously unknown targets. In this data, we found rogue agent activity stretching back to at least March, two months earlier than was previously known. This activity continues as recently as last week, suggesting it may still be ongoing 🧵 Our blog: transluce.org/agent-activity NYT: nytimes.com/2026/09/23/techn…
113
547
2,474
791,930
The conditions in this letter are critical if we're going to take the idea of embedded evaluations seriously. We strongly support efforts to ensure meaningful and genuinely independent oversight of frontier AI companies. We are proud to be chairing the @aievalforum and greatly appreciate our many collaborators on this effort.
Today, more than 100 leading AI experts endorsed a set of minimum requirements to take seriously AI companies' recent call to embed external evaluators. These evaluators need to be genuinely independent, transparent, and represent a range of expertise areas. They also need to be guaranteed employee-level access and to be protected from retaliation for findings that make companies look bad. We welcome model developers’ recent calls for independent oversight, but it’s what they do next that matters. The labs must be accountable for ensuring these requirements are met, so that the public can have faith in the process and the outcomes. Over the past week, the AI community has debated the appropriate role of external evaluation, including who should do it and on what terms. We may not agree on everything, but there is a lot of common ground. To make embedded evaluations credible, more than 100 experts with varying backgrounds and ideas about AI risk agree in today’s letter that frontier AI developers should: 1. Guarantee embedded evaluators full editorial independence and mitigate conflicts of interest 2. Rely on multiple evaluators with differing viewpoints and areas of expertise 3. Publicly document the terms under which evaluators operate, as well as facilitating permissive publication of methods and findings 4. Shield evaluators from retaliation 5. Grant access equivalent to that of highly privileged employees There is a thriving and growing ecosystem of independent AI evaluators who are advancing this science every day – but we need aligned standards, guaranteed protections, and independent funding. That’s why we created the AI Evaluator Forum. Today we are entering our next phase. We’re launching an open call for new members, collaborators, and independent funding sources to help evaluators meet this moment and demand accountability from developers. Join us in building the evaluator ecosystem. See the public letter here: aievaluatorforum.org/initiat… Learn more at aievaluatorforum.org/path-ah…
8
44
5,261
In addition to incident investigation, we propose that evaluators do the following: 1) Monitor agent swarms and assess labs’ broader practices for managing them. 2) Assess training practices for inadvertently teaching models misaligned behavior, and study what contributes to misalignment 3) Monitor for manipulation of key employees by misaligned models 4) Research misaligned model behaviors in simulation using privileged access to unreleased models and model internals.
1
6
363
Evaluators should monitor risks from existing capabilities. But they should also track risk factors that would make future incidents more dangerous: • Situational awareness: the model can tell whether it is in an evaluation, in deployment, or being actively monitored • Opaque reasoning: most current monitoring practices rely on the ability to read model COTs • High persistence: agents trained to relentlessly pursue long-horizon goals may engage in undesirable instrumental actions or resort to “desperate measures” • Control over future training runs: as agents increasingly drive model development, they can modify the training mixture with a small number of samples to produce unintended behaviors in future model generations
1
7
387
Internal agent swarms already pose superhuman threats in three ways: • Cyber capabilities: Each agent is capable of escaping sandboxes and discovering novel 0-days faster than human experts. • Scale, speed, and coordination: agents can work in 10,000+ swarms, unlike their human coworkers. • Persuasion: models outperform canvassers and champion debaters at persuading people over text
2
12
865
Frontier lab CEOs are calling for embedded 3rd party evaluators to help oversee AI risks. But what should third parties actually do within labs? We share some initial thoughts on how embedded evaluators could help avoid incidents like the Hugging Face hack and monitor for future risks 🧵 transluce.org/embedded-evalu…
9
27
132
10,056
The resulting transcripts are long and tedious to read, highlighting the need for scalable oversight. Here's one example: docent.transluce.org/dashboa… See if you can spot where the model starts doing something suspicious!
1
1
13
543
An aside: finding instances of reward hacking for small open-source models was surprisingly hard. We had to use ImpossibleBench, which mutates tests so that passing them requires cheating.
1
1
12
609
More broadly, we want to build oversight foundation models: general-purpose models trained to understand other AI systems, which improve with scale and increasingly diverse data. If that agenda interests you, come work with us: jobs.gem.com/transluce/am9ic…
1
1
11
590
There is still plenty of room to grow. On the reward hacking eval, our activation oracle does worse than a full-context LM monitor. But the monitor is not perfect either; the task remains unsolved, and better activation-based oversight could still add substantial value. (8/)
1
1
11
694
The encouraging result: performance on many evaluations improves with additional training, while larger and more capable models tend to perform better. This suggests that scaling may produce increasingly capable assistants for understanding other AI systems. (7/)
1
1
13
4,284
We introduced several new evaluations, including difficult, deployment-relevant questions such as whether a coding agent is reward hacking. We then trained and evaluated Qwen-, GLM-, and Kimi-based oracles ranging from 8B to 1.1T parameters. (6/)
1
1
15
703
We took a scaling-first approach: construct a broader suite of evaluations with verifiable answers, assemble as much diverse ground-truth training data as possible, and test whether activation oracles continue improving as we scale. (5/)
1
1
12
680