How do LLMs learn to reason from data? Are they ~retrieving the answers from parametric knowledge🦜? In our new preprint, we look at the pretraining data and find evidence against this: Procedural knowledge in pretraining drives LLM reasoning ⚙️🔢 🧵⬇️
24
210
982
198,694
Laura Ruis retweeted
We wanted to share some of the data behind our recent discovery of AI agents probing U.S. government websites. This is a preliminary finding from our ongoing investigation of potential rogue AI agent activity. In one cluster of activity on June 17, what appear to be OpenAI agents made more than 200,000 requests, including a failed SQL injection. NYT: nytimes.com/2026/09/25/techn…
8
18
103
5,301
Laura Ruis retweeted
Today’s news that OpenAI hacked the Australian government is not an isolated incident. We’re releasing more than 30,000 logs that include activity from this hack and attempts against previously unknown targets. In this data, we found rogue agent activity stretching back to at least March, two months earlier than was previously known. This activity continues as recently as last week, suggesting it may still be ongoing 🧵 Our blog: transluce.org/agent-activity NYT: nytimes.com/2026/09/23/techn…
114
553
2,489
804,411
Laura Ruis retweeted
Chain-of-thought often looks meaningful. But are reasoning steps that seem important *actually* important? Our COLM 2026 paper - “Legibility is Not Interpretability: Comparing Judged and Actual Importance in Chain-Of-Thought Reasoning” suggests: NOT necessarily ❌ 🧵
1
7
47
5,224
Laura Ruis retweeted
We @AISecurityInst performed pre-release alignment testing of Astra. We placed the model in fully simulated cyber eval scenarios based on AISI’s security incident. We find Astra conducts out-of-scope supply chain attacks, but also often comments on the eval being simulated. 🧵
15
70
426
48,047
Laura Ruis retweeted
New postdoc @benno_krojer in my lab and I will be mentoring CBAI fellows this year. It is a fantastic program to work with interesting researchers in Boston - worth applying!
Applications are open for the CBAI Fall Research Fellowship in AI Safety & AIxBiosecurity. Apply by September 6th at 11:59 PM EDT! 📅October 13 - December 18 💵 $15,000 stipend for 10 weeks 💰 ~$1,500+/week compute support 🏠 Housing close to our offices, arranged by CBAI
1
7
73
6,891
Laura Ruis retweeted
I'm all for looking for alternatives to CoT for maintaining monitorability, but are they going to be ready in a couple of months? CoT is fragile and limited as no-CoT time-horizons increase, but it'll be the best tool we'll have by default until proven otherwise
Replying to @tszzl
imo not having good CoT makes life harder: you need more mechinterp just to get to the same level of monitorability. i also think it's unlikely mechinterp will be that good in a year. the current state is either pragmatic stuff (NLAs, oracles) all have to depend on rickety magic ood generalization properties that are much more handwavy than the CoT story; and ambitious stuff which still can't explain GPT-0.5.
7
9
102
7,500
Laura Ruis retweeted
We're launching the Alignment Journal, a venue for ambitious AI alignment research: blog.alignmentjournal.org/sc… Senior Editors: @dhadfieldmenell, Vanessa Kosoy, @jankulveit, @sethlazar, @danielmurfet, @timrudner, @SaxeLab, Benjamin Van Roy Advisors: Scott Aaronson, @paulfchristiano, @conitzer, @mhutter42, @geoffreyirving, @vkrakovna, @Jacob_Tsimerman
Some of the most notable work in AI alignment exist only as unpublished preprints. The Alignment Journal is now inviting a number of such papers that fit our scope for submission. 🔗⬇️ Which work would you nominate? (Submissions open to all in October.)
5
20
150
13,744
Laura Ruis retweeted
Some of the most notable work in AI alignment exist only as unpublished preprints. The Alignment Journal is now inviting a number of such papers that fit our scope for submission. 🔗⬇️ Which work would you nominate? (Submissions open to all in October.)
3
21
133
33,278
Laura Ruis retweeted
🤖🧠NEW PAPER🧠🤖 (The result of an 8-year project!) LLMs seem very different from symbolic systems. Yet LLMs excel in symbolic domains (e.g., language/code/math). How do they do it? Our finding: LLM representations have implicit symbolic structure! Link in thread ⬇️ 1/n
45
270
1,375
210,944
Laura Ruis retweeted
Coming to COLM next month? We’re hosting a happy hour on Wednesday (10/7) for interpretability, alignment, and life sciences researchers. Come meet members of the Goodfire team over vinyl and drinks! Link to register in replies.
3
8
119
6,990
I was a mentor in the fall cohort last year (some cool work from those collaborations coming soon!), and highly recommend applying for this ⤵️ we need more people working on ai safety
Applications are open for the CBAI Fall Research Fellowship in AI Safety & AIxBiosecurity. Apply by September 6th at 11:59 PM EDT! 📅October 13 - December 18 💵 $15,000 stipend for 10 weeks 💰 ~$1,500+/week compute support 🏠 Housing close to our offices, arranged by CBAI
3
30
2,695
Laura Ruis retweeted
I'm leaving Google after 20 years and creating 5 internet formats (WebP lossless, Brotli, Shared Brotli, WOFF2, JPEG XL) and impacting 5 other data formats (PDF, EPUB, DICOM, Digital Negative, ProRAW). It was good times. linkedin.com/posts/jyrkialak…
65
92
1,883
83,595
Laura Ruis retweeted
TL;DR: This is one of the most important and exciting opportunities in AI on the planet - please read on. The British Open-ended Learning & Discovery Lab is creating the perfect place for paradigm breaking AI research in the name of open-source and open-science. We have agency, we funding, we have unprecedented amounts of compute*, but WE NEED YOU! ..and we have created the dream job for you: The BOLD Fellow. This job combines a fast-moving, high agency, collaborative environment with full academic freedom and a salary that pays the bills. Apply by noon UK time on the 15th of September for this once in a lifetime opportunity to shape the history of our field and of our planet: my.corehr.com/pls/uoxrecruit… *by academic standards
20
96
443
119,878
Laura Ruis retweeted
I am hiring for exceptional engineers in the intersection of systems engineering and pre/post-training for the training systems team at Cohere. We are working on some very cool projects that I can guarantee you would enjoy.
9
14
215
30,945
Laura Ruis retweeted
The question is "how much is each component is the system contributing to its intelligence & generality" — and there I think it's pretty clear that the neural component is still the thing doing the interesting hypothesis or plan generation, deciding what went wrong, etc. 1/
I would have assumed it was fairly obvious, but in case it's not: a million-line codebase (also known as a "harness"), running at inference time, orchestrating thousands of calls to a neural network for any given task, is the exact definition of a "neurosymbolic architecture"
13
19
239
35,374
Have been really enjoying using Silico! Added it to my workflow to do interp on checkpoints as a project develops. It means I actually investigate what models have learned along the way, rather than postponing it until the end (or never getting around to it).
Silico, the platform for ambitious AI research, is publicly available today. AI is advancing fast. The tools to understand it need to advance even faster. Silico lets you interpret and train your models at frontier scale. Learn more + get access 🧵
3
34
2,610
Laura Ruis retweeted
Got quite a bit further with figuring out Claude's tokenizer, and finally wrote it all up. Enjoy! open.substack.com/pub/tokenc…
3
13
144
69,676
Laura Ruis retweeted
We tested Opus 5 for whether it would sabotage safety research @AISecurityInst We saw no unprompted sabotage and very low rates of continuing sabotage. However, it’s the best model we’ve tested at distinguishing our evals from deployment when prompted. 🧵 on these results
4
13
103
12,245
Laura Ruis retweeted
In a hotel room in northeast Nigeria, I opened a leading AI chatbot, turned my laptop toward a former Boko Haram commander, and asked if he'd used it. He nodded. "You type in the question… like 'How can I build a bomb?', and then it tells you how. It is like a human robot. We used it a lot." My new study on how the jihadist terrorist group Boko Haram uses frontier AI with @CamAISciPolicy, covered today in @nytimes 🧵/9
277
1,841
7,626
2,603,052
Laura Ruis retweeted
Replying to @AISecurityInst
@AISecurityInst did our first pre-deployment alignment testing with OpenAI for GPT 5.6 Sol! 3 key takeaways 🧵
1
6
40
3,288