Sr. Staff RS @GoogleDeepMind. 🇯🇵-born 🇨🇳🇨🇦. ex: Omni Thinking, Gemini 1.5-3, GPT-4 @OpenAI, GoogBrain Robotics (JP: @shanegJP). Personal opinions.

Mountain View, CA
I had the pleasure of catching up with Geoff Hinton yesterday to celebrate his birthday. As we took a long, winding walk along the coast, I was reminded of: not all tokens are born equal. Just as his insights and words from a decade ago still trigger new chains of thought for me, yesterday’s conversation was exceptionally fruitful. We spent some time discussing Tim Lillicrap, my current manager and co-lead of Gemini thinking, whose research on random matrix feedback from ten years ago actually inspired my very first paper with Ilya Sutskever. Geoff offered his perspective on exactly why that method worked, dissecting the phenomenon with precise intuitions grounded in mathematics. Our discussion ranged from the technical to the philosophical. He remains as intellectually open-minded as ever, still imagining for true novel breakthroughs that go beyond the current paradigms of backpropagation. Our last (brief) chat was shortly after his foward-forward algorithm in late 2022. We also explored global perspectives, touching on the dynamics between China, Japan, and the West. We discussed how the philosophical roots of Asian cultures might be better suited for the societal shifts of a post-AGI world, and what the West could learn from the East to ensure that the future is positive. On a personal note, I was touched by his genuine concern for the next generation. When I mentioned my daughter, he agreed that raising children is the ultimate way to reshape one’s priorities--a grounding force that liberates you from immediate 'capitalist FOMO' and refocuses you on the long term. His words carry the weight of years of high-quality processing; they aren't just outputs of raw intelligence, but of deep, sincere care for humanity. It is clear he cares about far more than any single technology or nation. His reasoning and stories serve as a powerful reminder to understand and explore the broader world—proof that a vast perspective is what shapes greater wisdom and enhances our humanity. I attended NeurIPS 2024 hoping to congratulate him on his Nobel Prize but missed the chance. I am so glad I was finally able to do so this year on his birthday. And finally, he left me with a few more important predictive tokens: 'Google and Gemini will win' :)"
12
30
543
120,924
I felt a similar despair 4-5 years ago as I was still partially working on robotics. After co-authoring "LLMs are Zero-Shot Reasoners" and "LLMs Can Self-Improve" in 2022, it became clear that digital/symbolic AGI had to precede, and would ultimately accelerate, physical AI. While video models are powerful zero-shot reasoners (video-zero-shot.github.io/), LLMs remain the indispensable core in every aspect of physical AGI development. Over the past year, I've actively collaborated across Gemini, Omni, and generative media to test my hypotheses regarding physical AGI. The pieces are finally coming together. It is clear that frontier LLM labs, alongside the rest of the robotics community, are uniquely equipped to drive this next transition in embodied AI.
I’m sensing deep despair in academics over the past week. Astra, Fable, Muse are zero-shotting benchmarks in robotics & world models. Uneasy pill to swallow, but this is what step jumps in progress looks like.
11
36
379
92,620
Beyond RLHF valley
2
754
Mujoco is so back
GPT 6 Astra is capable of orchestrating two robots juggling with one ball at least airborne throughout the exchange This is an actual MuJoCo physics simulation at 1x speed
4
10
91
18,831
Shane Gu retweeted
We’re also introducing Gemini 3.8 Flash Cyber, our most capable cybersecurity model. It shows frontier-level performance in discovering vulnerabilities and patching them at scale, with Flash-level speed & pricing. That includes achieving 86.2% on the important CyberGym industry benchmark, plus 47.2% on CWE-Bench for patching. We saw a 70%+ success rate in discovering vulnerabilities across 20 programming languages on our internal benchmark.
339
405
4,571
408,320
Can Gemini count the number of claps? 👏 Accurately counting rapid movements is a notoriously tricky task for AI. Because static processing ingests video at a fixed 1 FPS by default, split-second movements like a clap easily get missed entirely or get confused with a snap or click. Watch Gemini 3.7 Flash use the new agentic video understanding capability to accurately identify and count every single clap by automatically adapting the processing speed as needed:
22
85
849
135,820
Mujoco is back.
Replying to @koraykv
3.7 Flash brings a big jump in agentic performance and coding accuracy. To demonstrate, we set up a 3-agent team to autonomously train a robotics control model from scratch. We hope you like 3.7 Flash, and you can read more here: blog.google/innovation-and-a…
8
6
154
29,982
Shane Gu retweeted
Today we're launching Gemini 3.7 Flash - our latest workhorse model for coding and agentic workflows, with an introductory price at half the original cost of 3.6 Flash. ⚡️ We have been iterating rapidly with the Flash series, going from 3.5 to 3.7 in just 3 months, making it more helpful across a wide range of tasks: • Software Engineering (DeepSWE v1.1): 37.0% ➔ 65.3% • Web Development (Code Arena Elo): 1506 ➔ 1588 • Enterprise Automation (AutomationBench): 13.4% ➔ 30.4%
121
243
2,366
592,644
3 new members have joined my team. Excited to work with them! 1) @rm_rafailov – Co-developed DPO at Stanford. Founding member and Tinker co-creator at Thinking Machine. 2) @arvind_io – Ex-Brain. Co-developed GPT-3 and GPT-4 at OpenAI 3) Jack Ji – RL PhD from Princeton
20
7
561
50,504
Rafael joined from outside Google, and Arvind / Jack are internal transfers within GenAI unit of GDM. It's more productive / fun with work with a smaller, flexible team these days.
1
45
4,509
In fact, the first test-time compute I read is the test-time training (TTT). It was coined "dynamic evaluation" by Mikolov and used by Alex Graves / Ilya in old RNN/LSTM days. SoTA compression algorithms like PAQ-8 also had weight updates.
ssi are ploughing the fields of a new paradigm shift, TTT (test time training). it will allow the models to update a subset of their weights on the fly. so your digital co worker won’t forget everything about your business needs in each new chat, continual learning is here. much like giving models a scratch pad to think and allocating test time compute, ilya has found the next shift and pushed hard. here’s a cute silly image to show you the third axis we are scaling. the laws. they hold 🥹
4
14
222
29,723
Shane Gu retweeted
Re: Demis, Jeff Dean moves: I think a fair few folks are treating this as bearish for GDM and that is imho a misread. The prospect of reaching AGI and ASI is beginning to look increasingly overdetermined. Being in operational leadership roles to preside over an overdetermined outcome is no longer as high-leverage as being in a leadership role on the next frontier. Early AI/AGI/ASI leads will, over the next year, begin leaving what look like important leadership posts to go place their bets on what they think the most important thing will be.
28
22
533
67,919
The "true AGI is the friends you made along the way" slide
Replying to @JeffDean
One more fun slide from our pitch deck.
3
16
339
38,274
Shane Gu retweeted
Our self-improving agents optimized the full @vLLM_project inference stack, with up to 16% more throughput and interactivity for @deepseek_ai v4 Pro and @Zai_org GLM 5.2 on B200s (no MTP). Every change was verified and our agents got better and faster at it with each iteration.
12
32
253
113,511
Shane Gu retweeted
It was an honor to give one of the invited workshop talks at ICML. The thesis: whoever turns messy real-world outcomes into reliable, scalable reward signals trains models where foundation models alone are stuck. It's also how we pick startups. Talk built with @HenryYin_ Coding agents suddenly started working because code shipped with free verifiers, like compilers and type checkers. We did reinforcement learning on those verifiers, and as we scaled compute, capability compounded. The next unicorn startup are ones that build verifiers for more challenging domains. Periodic Labs grades models with a physics fit against real measurements from its autonomous lab, and the capability ceiling moves. Applied Compute audits LLM judges across tens of thousands of rubric criteria before any RL runs, then trains open models to state of the art on Harvey's legal benchmark. Elorian is building verifiers that give models native visual reasoning capabilities, teaching models to actually see structure, where today's frontier models fail at counting more than ten objects.
Excited to co-organize the 1st "RL from World Feedback" workshop at ICML on July 10, featuring: - @MillionInt : ex VP of RL @ OpenAI - @chelseabfinn : Stanford/PI - Ben Eysenbach: Princeton - @brianzhan1 : Partner @ striker.vc - @robertarail : GDM - Jesse Zhang: UW
4
3
64
13,677
ICML RL from World Feedback workshop: Today 8:00 AM–5:00 PM KST @ Grand Ballroom 101-102!
Excited to co-organize the 1st "RL from World Feedback" workshop at ICML on July 10, featuring: - @MillionInt : ex VP of RL @ OpenAI - @chelseabfinn : Stanford/PI - Ben Eysenbach: Princeton - @brianzhan1 : Partner @ striker.vc - @robertarail : GDM - Jesse Zhang: UW
4
3
27
12,333
Back in the UK after 8 years! Working from GDM London (HQ) and took a trip at my old PhD lab in Cambridge. I miss the European academic ecosystem. Its rich history tunes out short-term FOMO. European AI researchers will keep thriving. My talk title: "From Academia to AGI".
2
6
141
15,717
It's important to go beyond classic ML research and keep finding any areas for "fundamental impact". AI has gone into societal impact phase. E.g. Carl Rasmussen continued David MacKay's legacy to improve climate change policies.
10 years ago today, we lost Sir David MacKay FRS. Physicist. Mathematician. Polymath. Gone at 48. I was working on my PhD at Cambridge, and attended some of his last lectures and symposium. He was a reason that attracted me to Cambridge over MIT in 2014. His textbook, Information Theory, Inference, and Learning Algorithms, was the first ML book I ever read — recommended to me by none other than Geoff Hinton. He used that same information theory to build Dasher — a text entry system where users steer through a continuous stream of letters flowing toward them, with a probabilistic language model making likely next letters larger and easier to reach, so that any tiny movement — a finger, a gaze — becomes efficient writing. It was the first ML application that truly blew my mind, and sent me deep into a rabbit hole: arithmetic coding, PAQ8 compression, nonparametric models. A journey I partly owe to his PhD student Christian Steinruecken, who also happened to share my love of Japan. As Chief Scientific Advisor to the UK's Department of Energy & Climate Change, he brought a physicist's clarity to policy. In Sustainable Energy – Without the Hot Air, he ran the numbers on our entire energy diet — and made me confront an uncomfortable truth. One of the biggest single factors? Beef — roughly 1,000 days of cow-time per steak. Hard to argue with the data. Hard to act on it when you were born and raised in Japan. I'm still working on that one, David. At his final symposium in Cambridge — just a few weeks before his passing — the room told the full story. Geoff Hinton and his Caltech PhD advisor John Hopfield — both Nobel Prize winners in Physics 2024 — gave tributes. Environment policy advisers spoke. Dasher users sent video messages of thanks from around the world — people who found their voice because of him. It was extraordinary to witness, in one room, just how many minds and lives a single person had touched. The story of how Hinton first noticed him: at a conference workshop poster session, among everyone who stopped by, it was the young MacKay who asked the sharpest, most penetrating question. Hinton remembered it. That's how it begins. I've always liked physicists who cross into ML — they bring a groundedness, a refusal to hide behind formalism without meaning. David MacKay and Max Welling are the role models I point to. Not just for the mathematics they built, but for how they carried it: with humility, curiosity, and a stubborn insistence on reaching beyond academia. He seemed to know his time was limited, and gave everything anyway. His legacy stays.
11
3,531
Excited to co-organize the 1st "RL from World Feedback" workshop at ICML on July 10, featuring: - @MillionInt : ex VP of RL @ OpenAI - @chelseabfinn : Stanford/PI - Ben Eysenbach: Princeton - @brianzhan1 : Partner @ striker.vc - @robertarail : GDM - Jesse Zhang: UW
2
18
197
43,014
Some goals for this workshop: 1) promote RL from world-grounded signals (e.g., efficiency, safety, and economic outcomes) beyond noisy RLHF 2) bridge academia and industry. Great to have @MillionInt @brianzhan1 to speak on latest RL startup trends of Silicon Valley + beyond
2
2
16
2,736
Flying to ICML in Seuol, Korea and will be around July 8-10 (then a few days in Tokyo). Excited to meet researchers looking into RL scaling + continual learning + recursive self-improvement. Also co-organizing the 1st "RL from World Feedback" workshop sites.google.com/view/rlxf-i…
7
4
128
11,314
workshop:
Excited to co-organize the 1st "RL from World Feedback" workshop at ICML on July 10, featuring: - @MillionInt : ex VP of RL @ OpenAI - @chelseabfinn : Stanford/PI - Ben Eysenbach: Princeton - @brianzhan1 : Partner @ striker.vc - @robertarail : GDM - Jesse Zhang: UW
1
3
2,687