FR/US/GB AI/ML Person, Director of Research at @GoogleDeepMind, Honorary Professor at @BOLD_Lab_AI, @ELLISforEurope Fellow. All posts are personal.

London, United Kingdom
🧵 Time for a short end-of-2025 wrap up 🧵 Genuinely aiming for this to be a short one for two reasons: 1. I'm doing it at the last moment 😅 2. Most of what I was involved in is not stuff that can be shared publicly (yet… or ever?). Or maybe I was just lazy... Let's go [1/12]
7
21
229
83,007
ngl this is a tune
The claude folks should just own this 🤣
8
5
160
58,676
RT @kevinwty: We GOT our fantastic paper IN!! 🎉🎉 #NeurIPS2026 Acceptance! 🐐 Want to know how to tackle long-horizon agentic tasks with R…
1
79
Edward Grefenstette retweeted
The arXiv paper on AIDE² is out. AIDE² improved its own research agent beyond the version we hand-tuned for two years, and the gains hold on benchmarks the outer loop never saw. New in the paper are transfer across models and comparisons with more AI research agents.
We're releasing the arXiv paper on AIDE². The RSI system where AI research agents improve their own research efficiency. It includes new results on transfer across models and comparisons with more AI research agents: [1/4]
5
17
1,652
Okay guys I think AI moviemaking is gonna make it.
We Asked AI To Simulate How The Founding Fathers Would React To Our Time And They Were Shocked
3
1
36
13,034
Edward Grefenstette retweeted
How aggressive a cancer is and how large is the benefit from chemotherapy are different estimands. Genomic assays were built decades ago for the first and repurposed for the second. Today, we released 3 preprints on estimating the benefit of escalated treatment directly.
1
21
36
3,981
Edward Grefenstette retweeted
Despite the recent progress in AI for mathematics, I don't think AI will replace the role of mathematicians. Instead, it makes mathematics closer to the natural sciences. In the natural sciences, we invent theories, comprehensible models, that explain empirical observations. When an AI answers a mathematical question, it provides humans with an effectively empirical observation about the answer. It is effectively empirical because the answer is a measurement from a system outside of immediate human theoretical understanding. This answer, as with any new glimpse of reality, remains mysterious until humans make sense of its underlying structure and provide an explanation of it in our own frameworks for understanding the world. An AI proof is stronger than simply an empirical observation, but it may be so complex that it provides little deep or novel insight beyond establishing the answer. The role of human mathematics is then to distill these effectively empirical findings into human-legible explanations. Seen this way, much of human mathematics may begin to look like a branch of interpretability, and AI, a new kind of telescope.
26
50
420
68,569
I don't want to take sides on this but OAI not coming across as the adults in the room here...
6
1
155
12,492
Disappointed to hear the @TfL announcement system in King's Cross blast an @audible_com announcement formatted as a normal announcement. Misleading, tasteless, and unwelcome.
18
3,313
"SPCX to the moon 🚀🚀🚀" bros have really gone silent, haven't they 😅
3
1
5
3,832
Edward Grefenstette retweeted
TL;DR: This is one of the most important and exciting opportunities in AI on the planet - please read on. The British Open-ended Learning & Discovery Lab is creating the perfect place for paradigm breaking AI research in the name of open-source and open-science. We have agency, we funding, we have unprecedented amounts of compute*, but WE NEED YOU! ..and we have created the dream job for you: The BOLD Fellow. This job combines a fast-moving, high agency, collaborative environment with full academic freedom and a salary that pays the bills. Apply by noon UK time on the 15th of September for this once in a lifetime opportunity to shape the history of our field and of our planet: my.corehr.com/pls/uoxrecruit… *by academic standards
20
96
443
119,879
So... solving reward hacking necessary for RSI, you say? 🤔😉
It's recursive self-improvement if it works, reward hacking if it doesn't.
2
19
5,334
Edward Grefenstette retweeted
While I believe in recursive self improvement, I think we are grossly underestimating how hard it might be. Results from evals like PostTrainBench suggest you can train AI to make incremental progress but novel ideas are harder to generate.
32
14
262
23,573
Edward Grefenstette retweeted
thanks for sharing i agree this should be done asap - the earlier the better it's also weak sauce that researchers are incentivized to transfer to california to escape the gardening leave here but imo, there is a finite (and not very long) time window when banning gardening leave is an obvious net positive for the uk frontier ai ecosystem once the cost of experiments goes high enough, it's plausible that a ban becomes sufficiently damaging that the top ai labs leave jurisdictions (even sunny california) that enforce a ban at that point a ban would be negative for the uk ecosystem from the perspective of building frontier labs that are v. capital intensive (though not necessarily negative for the long tail of other ai startups) obviously there are many factors that feed into this tradeoff one factor i think is underrated: right now siloing costs labs velocity, because researchers need broad context to be useful. if ai gets good enough to hold and route that context, labs can compartmentalise people without paying the same tax at that point any one departure carries much less of the org's r&d, so 6 months of topiary protects much less with silos, gardening leave doesn't matter so much without silos, it's a big deal i primarily view this as a reason to get this done now rather than wait
What UK AI startups need: @KanishkaNarayan In California, one can start a company the day after leaving Google. My colleagues Jeff Dean, Oriol Vinyals et al made this very clear recently with their impressive speed. In the UK, the same American companies impose 1 year garden leaves on senior AI researchers and 6 months on junior researchers. @GoogleDeepMind for example forced people to sign these contracts at the time of promotion, not the time of hiring. What this means is that researchers cannot start new companies for up to 1 year, cannot easily hire, and in short: they cannot compete. American VCs cannot understand why we move so slow. It is time for UK Gov to do the right thing: Make garden leaves optional for employees. If you’re an AI entrepreneur or VC in the UK or Europe, I would appreciate your comments here. You’ve all confided in me. It is time you let government know the urgency of this — everyone should speak up.
3
31
5,506
Edward Grefenstette retweeted
1/ Today, we introduce Faraday, a 27B-parameter AI Scientist that extends the capabilities of coding agents with a layer of scientific intuition. Trained via long-horizon RL, Faraday outperforms Claude Opus 4.8 and GPT-5.5 on the task of replicating research papers. 🧵
75
227
1,758
367,377
Amazing launch, @redpony and team! Great to see Google ship technology in this space 🥰
I'm so happy to be able to announce the first general purpose sign-language-to-text translation (SL2T) model from my team that's out today and powering a new ASL input feature on @Android phones (and don't worry, we're not stopping at ASL).
2
1
18
3,645
Edward Grefenstette retweeted
A sad thing I see is when companies start becoming reactive over proactive. Its a clear external sign that the leadership has lost vision. They make decisions they feel are safe instead of through conviction of a future. The problem is that "safe" almost always means following someone else. This isn't a clear sign of death, but it is a very difficult position to be in... For me personally, I hate seeing it because so many companies start with a genuinely interesting vision. Then the world changes, they lose confidence in that vision, and instead of finding a new way to adapt and express it, they start chasing whatever everyone else is doing. It happens during every major secular shift. AI is the obvious current example. A lot of companies that looked exciting just a few years ago now feel like they're playing catch-up instead of leading. Its a tough thing to watch.
78
129
1,996
125,503
Feeling BOLD with @bold_lab_ai
2
1
61
3,705
Kavinsky died?? 😢
5
16
8,126
This Leopold / Situational Awareness story is crazy. > Go from $1.5B to $20B in under two years. > Over 1,000% since launch, 270% ytd to May alone. > Jane Street which almost never allocates to external managers becomes an investor. > $24B near the highs, levered as much as 4x. Up 439% for H1'26. > July 2026: AI names go down as much as 50%. > Compounded with up to 4x margin... > Leopold sends letter to LPs for additional capital reported by FT yesterday, says that this is a buying opp. > Today: CNBC reports Situational Awareness has exited all public equity positions. > One buyer takes everything, both L/S books, in one enormous block. > Jane Street?? Who buys a whole levered spread book in one go? Isn't that is literally Jane Street's business? > As an LP, Jane Street had the inside view of exactly his book. Which would make this even more ironic... > If it was JS, the arc completes: Leopold's rarest backer would be the buyer of his whole book lol. > What's left of Leopold? Whatever's left in the private book after selling Anthropic stake and the capital raise, if any.
6
4,678