Teaching AI the joy of invention at @Recursive_SI, Professor of AI @AI_UCL, PI @BOLD_Lab_AI, Fellow @ELLISforEurope. Ex @GoogleDeepMind @AIatMeta @CompSciOxford

London, England
My @iclr_conf 2025 Keynote on "Open-Endedness, World Models, and the Automation of Innovation" is now publicly available: iclr.cc/virtual/2025/invited…
13
44
394
67,049
Tim Rocktäschel retweeted
We are again recruiting @bold_lab_ai - please share the love 🙏. We are looking for: 1) Postdocs (my.corehr.com/pls/uoxrecruit…) -- deadline 9th of Oct at noon 2) research assistants (my.corehr.com/pls/uoxrecruit…) -- deadline 9th of Oct at noon 3) Strategic Partnership Project Manager (my.corehr.com/pls/uoxrecruit…) -- deadline 14th of Oct at noon
61
50
272
639,761
Impressive. Luckily we can turn @NetHack_LE / balrogai.com into an ASI benchmark in an instant. The human world record is 61 consecutive NetHack ascensions (teddit.net/r/nethack/comment…). Also curious to see if Astra can ascend in all roles (nethackwiki.com/wiki/Z-score).
9
9
86
9,938
Aaaaaaaand it's apparently solved 🫣😅
The NetHack Learning Environment (@NetHack_LE) has been around since 2020. Since 2024, balrogai.com is evaluating frontier LLM performance on it. We are slowly seeing progress on it, but it still seems far from being solved.
19
41
523
83,977
Tim Rocktäschel retweeted
Looks like it may be AGI now. kenforthewin.github.io/blog/…
"Is it AGI" flow chart. Developed with @_rockt at NeurIPS 2022.
12
25
340
62,298
Tim Rocktäschel retweeted
We’ve invested in @Basecamp_Res as part of its $140m round. Basecamp has built a database of genetic information from 1m+ unstudied organisms, using it to design new medicines, including cancer and autoimmune therapies. The potential to transform medicine discovery is huge.
4
6
30
4,690
Tim Rocktäschel retweeted
Many of the comments are about nonlinear time lines. See the 2014 mother of all nonlinear time lines in Sec. 20 of [1], starting with the Big Bang, converging soon: people.idsia.ch/~juergen/dee…
4
11
66
16,237
Tim Rocktäschel retweeted
AI moves so fast that I thought this book was going to be a year too late. It turns out it's right on time. While I was writing, I kept having to change chapters from "someone should do this" to "someone has done this," and then go talk to the people who had done it. A good example is the chapter on virtual cells. I'd written that we needed large initiatives to build one. By the time I finished, I was writing about the ones that exist — the Chan Zuckerberg Initiative's is now the largest and best-funded push to model a human cell. The Eureka Machine is out today. It's about building a full-stack scientific superintelligence: a living map of human knowledge, a model of physical reality, high-fidelity simulations, autonomous labs, and an agent swarm of AI scientists on top. There's no better science to start with than the science of AI itself, because that's where the feedback loop is tightest. Then you expand the aperture to physics, chemistry, and especially pre-clinical biology. This will be the most meaningful application of AI, and one of the most important things humanity does. And it may be the last organic invention we need to make. After that, it starts inventing everything else for us to solve our most pressing problems. While writing, I couldn't let the idea go. Eight of us started @Recursive_SI to go build it. AI is going to massively accelerate science, and unlike most of what gets argued about right now, that isn't zero sum. I hope you all enjoy the book. I look forward to discussing it with you all.
22
46
185
26,481
The NetHack Learning Environment (@NetHack_LE) has been around since 2020. Since 2024, balrogai.com is evaluating frontier LLM performance on it. We are slowly seeing progress on it, but it still seems far from being solved.
7
6
109
57,009
Tim Rocktäschel retweeted
The UK’s most powerful government-built AI supercomputer, Isambard-AI, cost £225m The Botley Road bridge cost roughly the same Apparently our frontier AI capacity and one bridge are comparable national investments
Cars and buses go from Oxford Station down Botley Road for the first time in three years! The project took longer than the construction of The Titanic and saw costs balloon to £237m. Looks good though.
17
64
846
72,508
Tim Rocktäschel retweeted
Two more entries in the BALROG's leaderboard, once again thanks to @creus_roger 🙏 Claude Opus 5 and GPT 5.6 Terra both with max reasoning effort. Terra places neatly in between Luna and Sol, as expected. Opus 5 max gets close to GPT-6 astra, but not quite there.
1
4
17
1,984
Tim Rocktäschel retweeted
At Tesla, we learned the concept of the 'idiot index' which is retail price / BOM cost. The higher the index, the more the need to understand how that thing was built and what drove the cost. These motors have a raw BOM cost of about $25. In the US, they are sold for $125-$300, by most incumbents; an idiot index of >5 in many cases. To understand why this happens, we visited some of these incumbents' lines. Three reasons: 1. Not AI native on the ops side: Still using spreadsheets and expensive software to manage inventory, schedule production, and handle customer support. No hunger to move fast and rip out these systems. 2. Locked into expensive lines; high retail price helps with ROI: Some of these companies have bought automated, rigid lines for relatively commoditised motors (3115s). Now they face a price squeeze and can't really lower prices since they want to pay off the line. They can't move to better-margin motors because lines are fixed. 3. Manual lines, $55/hour landed labour cost: Folks who saw manufacturers struggle with 2 above are keeping lines manual and flexible, but need to deal with the high cost of labour, which is passed on to the customer. What happens when we can get robots to automate this assembly end-to-end? Do dark factories mean product prices drop to COGS + energy? Excited to find out.
90
131
2,227
258,718
Great progress. Not yet AGI though.
BALROG’s leaderboard has three new entries, courtesy of @creus_roger Very interesting that GPT 5.6 Sol at max effort is still within error of Gemini 3 and 3.1 pro. GPT Astra 6 on max reasoning effort however reaches new heights, and a @NetHack_LE avg. progression of 13% 🏰
3
2
34
6,934
Tim Rocktäschel retweeted
BALROG’s leaderboard has three new entries, courtesy of @creus_roger Very interesting that GPT 5.6 Sol at max effort is still within error of Gemini 3 and 3.1 pro. GPT Astra 6 on max reasoning effort however reaches new heights, and a @NetHack_LE avg. progression of 13% 🏰
6
9
56
13,043
Tim Rocktäschel retweeted
Okay, first run of 3× GPT-6 Astra agents coordinating on Alem (Hard) is looking pretty interesting 👀 25.4% Base, 25.2% Coord. (current best Hard Coord. on the leaderboard is ~19%) Just one seed so far, so very preliminary.
2
3
31
3,648
Tim Rocktäschel retweeted
I like this article in MIT Tech Review South Korea. It covered many important topics, and allowed me to describe what I feel like is a new type of RL that is only recently possible, yet powerful. Curious to hear what everyone thinks of my answers. jeffclune.com/media/shared/J…
3
6
44
6,821
Tim Rocktäschel retweeted
if you don't own your model, the model owns you.
everyone getting sniped by the personal drama, but missed the more interesting unanswered question: can these labs see all your work and scoop you when the stakes are "high enough"? 🤷🏻‍♀️
4
15
173
27,140
Tim Rocktäschel retweeted
There are a lot of things wrong with this world… but too much intelligence is not one of them.
55
459
2,924
339,648
Tim Rocktäschel retweeted
I wrote down my thoughts about the last several weeks of AI hell. ⬇️
14
88
513
153,133