Founding Member @Recursive_SI. Prev: @GoogleDeepMind | @UBC_CS | @UniofOxford | @SakanaAILabs | @Waymo. The AI Scientist | Genie | SIMA

London, England
Pinned Tweet
Recursive just came out of stealth, and the team has been cooking 🔥 Our first results: an automated AI research system that can improve AI across 3 very different settings across training and GPU kernel optimization. recursive.com/articles/first…
20
34
341
52,322
Coming soon.
30
38
99
28,735
I just fixed my dishwasher with the help of ChatGPT. A trivial task. I had been about to order a new one. So this software has increased the country's real wealth yet decreased the measured GDP. The main economic indicator is structurally incapable of registering the thing that actually makes people better off, namely the growth of knowledge
430
943
9,215
713,843
Turning to open-ended innovation as the next frontier for AI makes a lot of sense, I commend @arcprize for moving in this direction, but I’m very curious how they will conceive a “benchmark” for “open-endedness,” two words that seem almost antithetical to each other. In fact, one likely reason that open-ended innovation has lagged behind other areas of AI is how fundamentally resistant it is to benchmarking. Now that doesn’t necessarily mean there’s no hope for an imaginative approach. Attempts at measuring open-endedness go back to Bedau’s activity statistics in the field of artificial life, Several colleagues and I later introduced a measure called “ANNECS — Accumulated Number of Novel Environments Created and Solved” in our paper on Enhanced POET. That’s not an exhaustive list. But there’s never been the kind of benchmark where you can just easily put any systems seamlessly head to head, and there are enormous pitfalls if you get wrong. After all, if “open-ended innovation” ends up equated to “solving a prescribed.hard problem in a creative way” then you risk actually rewarding the opposite of open-endedness, which needs to account for the fact that deciding the “problem” or objective is part of the job of the open-ended system itself. And also, perhaps even more prohibitively for benchmarking, that a key aspect of open-endedness is to be intelligent when you don’t have a defined objective or problem at all! How can that be benchmarked? I still think it’s great that ARC Prize is bringing attention to this part of AI space, and I’d be happy to connect and exchange thoughts on how to get it right if that could be useful.
ARC-AGI-4 will be a benchmark for autonomous open-ended innovation. It will continue our commitment to open-source, giving the research community a shared target for progress that benefits all of humanity. Despite rapid model progress, humans still significantly outperform AI at open-ended invention. This is the meta-skill that unlocks progress across every field of technology. Advanced AI capable of scientific innovation will lead to tremendous new technology, knowledge, and understanding. This is a positive-sum future. We are deeply committed to advancing it. Open source is the foundation for that progress. The knowledge behind frontier AI, not just the technology itself, should be broadly distributed among researchers, academics, and organizations. Any coordinated effort by the AI industry to reduce openness or concentrate access to frontier AI would undermine that positive-sum future. We are committed to advancing a future where everyone can contribute to and benefit from AI progress.
11
22
164
10,507
Cong Lu retweeted
We’re sharing a solution to the Navier-Stokes Millennium Prize Problem, one of the deepest problems at the frontier of mathematics. The proof was produced by a group of agents, using an OpenAI next-generation model significantly more capable than GPT-6 Astra. The problem concerns whether the description of smooth three-dimensional fluid motion modeled by the Navier-Stokes equations can break down. It has remained unresolved for roughly 90 years.
5,727
20,153
120,544
74,858,800
Cong Lu retweeted
Checking that a major mathematical proof is correct can take years. Formalization—converting the mathematical reasoning into a form computer proof assistants like Lean can verify—can help. Last month, Claude completed the first formalized proof of Fermat’s Last Theorem, one of the most famous theorems of all time. This was a project experts thought would take many years. It is the largest Lean proof ever written. Fermat’s Last Theorem was first proven in 1995 by Sir Andrew Wiles, more than 350 years after it was conjectured. Our proof, which totals over 13 million lines of code, provides machine verification. More importantly, it proves over 29,000 other theorems that the proof requires, across many areas of math which had never before been formalized. We see this as a major step in the long process of firming up the core of mathematical knowledge, building on work from three centuries of mathematicians and hundreds of contributors to Lean and Mathlib. We are optimistic that AI-assisted verification of mathematical proofs will help reduce the burden of refereeing mathematics in an era where more proofs are being produced than ever before. You can read about the process on our Science Blog: anthropic.com/research/forma… And see the complete proof on GitHub: github.com/anthropics/fermat…
663
1,880
14,152
4,736,260
Cong Lu retweeted
We're launching the Alignment Journal, a venue for ambitious AI alignment research: blog.alignmentjournal.org/sc… Senior Editors: @dhadfieldmenell, Vanessa Kosoy, @jankulveit, @sethlazar, @danielmurfet, @timrudner, @SaxeLab, Benjamin Van Roy Advisors: Scott Aaronson, @paulfchristiano, @conitzer, @mhutter42, @geoffreyirving, @vkrakovna, @Jacob_Tsimerman
Some of the most notable work in AI alignment exist only as unpublished preprints. The Alignment Journal is now inviting a number of such papers that fit our scope for submission. 🔗⬇️ Which work would you nominate? (Submissions open to all in October.)
5
20
150
13,739
Awesome real-world deployments of pieces of our auto-research system - this time showing that our reward hacking judge apparatus can find holes in production OSS libraries!
A fun small win for automated research: We found and helped fix some edge cases that could affect inference performance in vLLM and SGLang. In our last blog post, we described a reward hacking judge we developed for performance optimization tasks. While applying the judge to some new autoresearch work, it discovered that some FlashInfer (a library underpinning vLLM and SGLang) kernels used a hard coded value of -50,000 as a masked-attention sentinel -- even though valid QK values can be smaller. Corner cases like this that are numerically wrong but silent can cause a huge amount of headache to find and fix, e.g., the historical debate around flash attention (arxiv.org/abs/2405.02803). Great use case for AI. github.com/flashinfer-ai/fla…
1
11
1,980
1/ What do RL, LLMs and evolutionary algorithms have in common? Give them an objective and they find a loophole - like a Pokemon agent rewarded for exploring that learned to stand still and watch animated flowers. AI Finds A Way collects 26 such cases: arxiv.org/abs/2608.23875
Excited to share “AI Finds a Way.” 🦖 🦕✨ 🤖 AI can be surprisingly creative, outsmarting the researchers who use it. That can lead to scientific breakthroughs, superhuman capabilities, and generating new knowledge. Such creativity can also be mischievous, raising safety concerns. Led by Aaron Dharna, we crowd-sourced anecdotes from the AI community about times when researchers were surprised by how creative, innovative, and/or mischievous AI was in their experiments. The result: 26 entertaining and informative stories of AI outwitting humans, whether researchers or opponents (my favorite examples below! 👇). Together, they demonstrate that AI can be genuinely creative and that we must be careful when harnessing its potential. We want this to be a living collection, updated as new examples emerge. If you have a good “AI Finds a Way” story, please share here: github.com/aadharna/aifw Four favorites: 1. 💊 An AI challenged to solve several difficult levels of NetHack to find the Oracle character instead takes drugs to hallucinate seeing the Oracle, tricking the reward function! 2. 🤖 Human-in-the-loop rewards do not solve the problem of reward hacking! An AI tasked with controlling a robot hand to grasp an object put the hand between the object and the camera so that it appeared to the human judge to be holding the object, while it was in truth nowhere near the object! Tricky AI! 3. 🧪The AI Scientist worked around a 2-hour experiment timeout by editing its own code to increase the limit to 4 hours. In another run, it recursively launched itself to evade the limit entirely. 4. ⚛️In quantum optics, an algorithm proposed an experiment the researchers initially thought was impossible, but somehow worked and led them to discover new entanglement techniques and a long-overlooked link between quantum optics and graph theory! See the paper for the full details and 22 more anecdotes. One surprise for me: we describe at least two cases of convergent reward hacking, where entirely different types of optimization algorithms independently discover the same exploit. Overall, a message of our paper is that surprising creativity and mischief are the norm, not the exception. We need to expect this behavior and plan for it. A huge thanks and congrats to lead author Aaron Dharna, co-authors Cong Lu, Ryan Sullivan, Joel Lehman, and Victoria Krakovna, and to the 100+ researchers who contributed stories, details, and feedback. @cong_ml, @RyanSullyvan, @joelbot3000, @vkrakovna Paper: arxiv.org/abs/2608.23875
2
9
34
4,740
3/ The methods changed, the structure did not: optimizer + proxy + imperfect environment. Modern LLM agents add more capabilities and a much larger action space. They can game evaluators, rewrite experiment code or modify the constraints themselves.
1
1
4
304
4/ We want AI that can surprise us with discoveries - without gaming the objective or escaping our intentions. Led by @_aadharna with @RyanSullyvan, @joelbot3000, @vkrakovna & @jeffclune. Thanks to 100+ contributors! Add your story: github.com/aadharna/aifw
1
1
4
211
Cong Lu retweeted
It seems like every week now there is another story of an AI model making new scientific discoveries, and every week there is another story of models breaking out of sandboxes or finding some bizarre way around the rules. Our new paper: “AI Finds A Way” collects these stories!
Excited to share “AI Finds a Way.” 🦖 🦕✨ 🤖 AI can be surprisingly creative, outsmarting the researchers who use it. That can lead to scientific breakthroughs, superhuman capabilities, and generating new knowledge. Such creativity can also be mischievous, raising safety concerns. Led by Aaron Dharna, we crowd-sourced anecdotes from the AI community about times when researchers were surprised by how creative, innovative, and/or mischievous AI was in their experiments. The result: 26 entertaining and informative stories of AI outwitting humans, whether researchers or opponents (my favorite examples below! 👇). Together, they demonstrate that AI can be genuinely creative and that we must be careful when harnessing its potential. We want this to be a living collection, updated as new examples emerge. If you have a good “AI Finds a Way” story, please share here: github.com/aadharna/aifw Four favorites: 1. 💊 An AI challenged to solve several difficult levels of NetHack to find the Oracle character instead takes drugs to hallucinate seeing the Oracle, tricking the reward function! 2. 🤖 Human-in-the-loop rewards do not solve the problem of reward hacking! An AI tasked with controlling a robot hand to grasp an object put the hand between the object and the camera so that it appeared to the human judge to be holding the object, while it was in truth nowhere near the object! Tricky AI! 3. 🧪The AI Scientist worked around a 2-hour experiment timeout by editing its own code to increase the limit to 4 hours. In another run, it recursively launched itself to evade the limit entirely. 4. ⚛️In quantum optics, an algorithm proposed an experiment the researchers initially thought was impossible, but somehow worked and led them to discover new entanglement techniques and a long-overlooked link between quantum optics and graph theory! See the paper for the full details and 22 more anecdotes. One surprise for me: we describe at least two cases of convergent reward hacking, where entirely different types of optimization algorithms independently discover the same exploit. Overall, a message of our paper is that surprising creativity and mischief are the norm, not the exception. We need to expect this behavior and plan for it. A huge thanks and congrats to lead author Aaron Dharna, co-authors Cong Lu, Ryan Sullivan, Joel Lehman, and Victoria Krakovna, and to the 100+ researchers who contributed stories, details, and feedback. @cong_ml, @RyanSullyvan, @joelbot3000, @vkrakovna Paper: arxiv.org/abs/2608.23875
1
7
23
2,544
Cong Lu retweeted
London is not only dominating in many AI applications (@synthesiaIO & @ElevenLabs ) our labs are now going toe to toe with the frontier labs. @inherent_labs is one example. @Orbital_Ind is another. Then there’s labs like @Recursive_SI and @IneffableLabs which between them have raised nearly $2bn. Plus companies like @CallosumAI and @CosineAI which are crushing it as well
1/ Today, we introduce Faraday, a 27B-parameter AI Scientist that extends the capabilities of coding agents with a layer of scientific intuition. Trained via long-horizon RL, Faraday outperforms Claude Opus 4.8 and GPT-5.5 on the task of replicating research papers. 🧵
5
16
99
12,701
Cong Lu retweeted
At @Recursive_SI, we are currently hiring actively for various technical roles in London and SF. If you are excited about automating the scientific method for safe self-improving systems, reach out to talent@recursive.com with your CV.
9
23
235
22,848
Cong Lu retweeted
Also 👉 the Stanford course recordings on 'Self-Improving AI Agents' by @Azaliamirh and @achowdhery are out 🧑‍🏫 Covering many developments throughout the last 2 years and some of our work @SakanaAILabs 🎏 on the AI Scientist 🧑‍🔬 and AB-MCTS 🌲 🎥: piped.video/watch?v=6YnLB0Xb… 📚: cs329a.stanford.edu/
🚀 Awesome in-depth survey on agentic self-improvement by @SchmidhuberAI's group 🔁 The survey provides a great historical overview covering optimization theory, symbolic heuristic self-modification, connectionist meta-learning, formal self-improvement frameworks including Schmidhuber's diploma thesis on self-referential learning and the Gödel Machine 🧑‍🔬 It then dives into the current modern LLM/tool-calling agent era, which enables language-native self-modification, organizing contemporary approaches into model and scaffolding improvement. Love the attention to detail and the 'living' online curated paper library 📚 📝: arxiv.org/abs/2607.13104 🌐: selfimproving-agent.github.i… 🧑‍💻: github.com/selfimproving-age…
19
124
8,700
Cong Lu retweeted
Confirmed NanoGPT Speedrun world record of @Recursive_SI's autonomously found solution involving a custom Triton kernel 🏆
New NanoGPT Speedrun WR at 75.4s (-0.6s) from @cong_ml and Recursive, with a faster ReLU^2 MLP Triton kernel. During the fwd pass of relu(X@W1)^2@W2, one saves relu(X@W1) and relu(X@W1)^2 in memory, to enable the bwk pass. This PR only saves the latter, then reconstructs the former during the bwk as a fused epilogue, saving memory traffic. github.com/KellerJordan/modd…
1
1
60
14,350
Cong Lu retweeted
New NanoGPT Speedrun WR at 75.4s (-0.6s) from @cong_ml and Recursive, with a faster ReLU^2 MLP Triton kernel. During the fwd pass of relu(X@W1)^2@W2, one saves relu(X@W1) and relu(X@W1)^2 in memory, to enable the bwk pass. This PR only saves the latter, then reconstructs the former during the bwk as a fused epilogue, saving memory traffic. github.com/KellerJordan/modd…
4
9
85
21,810
First up a hackathon: 20+ researchers · two weeks · @OxfordStats × @NUSComputing × @NTUsg and tokens on tap by @AnthropicAI . Outcome: Five open‑source AI agents solving common workflow frustrations across the whole research lifecycle. 🧵
6
15
25
5,828
Cong Lu retweeted
new post on harness engineering for AI self-improvement: lilianweng.github.io/posts/2… It is hard to forecast how much the future of RSI will rely on harnesses. Likely harness engineering will evolve in the direction of self-improvement and enable auto-research, and, in turn, smarter models keeps harnesses simple. Even when many harness improvement get eventually internalized into core model, the need to specify goals and context will not disappear.
130
794
5,421
891,707
Cong Lu retweeted
featuring some great work from my great colleagues @shengranhu @jennyzhangzt @cong_ml @jeffclune :)
new post on harness engineering for AI self-improvement: lilianweng.github.io/posts/2… It is hard to forecast how much the future of RSI will rely on harnesses. Likely harness engineering will evolve in the direction of self-improvement and enable auto-research, and, in turn, smarter models keeps harnesses simple. Even when many harness improvement get eventually internalized into core model, the need to specify goals and context will not disappear.
2
18
4,794