Research at @GoogleDeepMind. Gemini Omni Team. Priors: GNNs, Structured World Models, Neural Assets, Veo Ingredients/References, Veo Robotics

San Francisco, CA
My PhD thesis "Deep Learning with Graph-Structured Representations" is now available for download: hdl.handle.net/11245.1/1b63b… -- It covers a range of emerging topics in Deep Learning: from graph neural nets (and graph convolutions) to structure discovery (objects, relations, events)
43
599
3,167
Thomas Kipf retweeted
Leaving aside the arguments over the reasons why this has happened, it is shocking that Europe does not have a single frontier AI lab, nor even a near-frontier lab nor even an effort that could likely lead to building up a frontier lab in the future.
250
211
2,836
146,641
Now is the time to step up the game in terms of AI cybersecurity before it’s too late. Great article and vision by @Shalev_lif and team.
We are approaching cyber-superintelligence. But our systems aren't ready. We must secure our model weights and infrastructure against three new threats: sabotage, escape, and theft. There is a way forward. Secure Acceleration: A Cyberdefense Strategy for Superintelligence
1
2
1,570
🙋‍♂️
there are many many people who studied physics and came to the conclusion that working on machine learning would probably create more scientific progress than studying physics including many of the original chatgpt crew like john schulman, liam fedus, amodei & jared kaplan at ant
1
1
29
5,627
Thomas Kipf retweeted
Today we have 3 new DeepMind Institute essays: How can we control misbehaviour in agent swarms? How should we orchestrate complex networks of AIs and people? The case for making AGI’s benefits equitably distributed. Get them here -> bit.ly/deepmind-institute
52
130
891
76,145
Thomas Kipf retweeted
I was encouraged this week to see the leaders of the frontier labs agree on the need for them to slow down the pace of AI development. Given the stakes, it’s a good and necessary first step.   But I’m even more encouraged by the growing recognition that how this powerful new technology develops should be at the center of our public debate.   I’ve been watching the progress on AI for over a decade now, and one thing that’s clear to me is that the potential impact of this technology is not overhyped. It’s also moving at lightning speed – and even faster than those who are engineering it can keep up with.   I’m not an AI accelerationist who believes it will lead to some techno-utopia, and I’m not a doomer who thinks it will inevitably lead to humanity’s destruction. But whether this technology results in amazing breakthroughs in medicine, energy and education or unleashes huge economic disruptions, greater inequality, and potential catastrophe will depend on the choices that we make right now – choices that should be made not just by the companies involved, but by all of us.
3,831
4,514
33,080
4,688,725
Continual learning a.k.a. learning on the fly
I and @nicochristie ran the fly connectome and found the group of neurons (hΔH, hΔA, hΔI and hΔG) that could allow the fly to navigate using fast synaptic weight updates, not neural activations. This is fast-weight continual learning in a fly, something current LLMs don't do!
5
1
75
7,213
We’re hiring! This is a rare opportunity to work with the folks that built Veo, Genie, and Gemini Omni — and advance the frontier of world models. More details below 👇
15
40
363
51,451
Research Scientist opening (Mountain View / San Francisco): google.com/about/careers/app…
12
2,483
Thomas Kipf retweeted
1) Intelligence is the process of minimizing the generator-verifier gap 2) The generator-verifier gap is a fundamental property of the universe 3) ASI is the optimal such process
If a solution is fundamentally easier to be verified than to be generated, it probably means that there is a learning signal. Lots of problems fall in this category.
5
6
93
15,155
Hope everyone in SF got up at 5:30am this morning to secure their NeurIPS ticket
5
2
42
12,731
it’s a good model (and still insanely fast)
Introducing Gemini 3.8 Flash, another jump in Gemini's agentic + coding capabilities, and our 3rd updated Flash model in only 6 weeks... This model has been a ton of fun to work with, excited to see what you all think!
23
3,257
We just released an update to Gemini Omni: video references, extension, 360p, 4K & frame interpolation — and it’s a better model overall ✨
Gemini Omni 1.1 Flash is our newest multimodal model for video generation and editing. It delivers a new suite of creative capabilities and controls for developers 🎥 With this update you can: 🎬 Extend your scenes 🎯 Specify starting and ending frames of a shot ➕ Add video input references ✨ Upscale your favorite takes up to 4K ⚡ Test ideas quickly in 360p See these in action 🧵
5
4
44
4,508
Thomas Kipf retweeted
Multimodal sensors are indispensable. By combining inputs from cameras, lidar, and radar, the Waymo Driver creates a rich, redundant world view that no single sensor can replicate. • Lidar captures 3D geometry with millimeter precision • Cameras read street signs and light colors • Radar tracks velocity and cuts through rain, fog, and dust
180
82
1,070
751,296
This is huge. I always thought it would take until the 2030s until we see fully autonomous cars on the roads in Germany due to the famously slow and conservative regularity approach. Auf geht’s!
Servus, München! 🇩🇪 We’re bringing the Waymo Driver to Germany. Over the coming weeks, our vehicles will arrive in Munich to begin laying the groundwork to launch our fully autonomous ride-hailing service in late 2027! Read more: waymo.com/blog/2026/08/waymo…
8
2
149
18,712
Mike’s recirculation paper and Jie’s post about scaling laws coming out on the same day couldn’t have been timed better
4
3
58
6,930
Thoughts About Scaling Law Scaling, but not only of parameters. Every model release now ends with the same question: how many parameters? It isn't a question that can be answered on its own. Parameter count is only meaningful alongside three others — how much data you have, where you intend to spend your compute, and who will run the model, under what conditions. The field learned this the hard way. Kaplan et al. (2020) fit an exponent that told everyone to grow parameters faster than data — roughly 2.7:1 — and the industry complied: GPT-3, Gopher, MT-NLG. Hoffmann et al. (2022) redid the experiment across four hundred models and found the compute-optimal split is closer to 20 tokens per parameter, and that with sufficient compute the two should grow at the same rate rather than drifting apart. The error in the earlier fit compounded with every order of magnitude of compute, which is why the largest models of that generation were the most misallocated. The trillion-parameter round was, in retrospect, a detour the whole field took together and then reversed. Chinchilla wasn't the end either. It optimized training compute for models that would be trained once and evaluated. Today a model is called billions of times a day and inference dominates lifetime cost. Put inference into the objective and the optimum moves toward smaller models trained far longer — deliberate over-training, which is what Llama-2-7B and Gemma-2-9B were doing at roughly 290 and 889 tokens per parameter. Sparsity moved the target again. In a MoE model two quantities have to be kept apart: total parameters govern roughly how much the model can hold — knowledge, facts, the long tail — while activated parameters and effective depth govern roughly how far it can think, how many steps of a causal chain it can carry before it comes apart. A dense 20:1 ratio does not transfer. And the ratio isn't a single number at all: Roberts et al. (2025) find the optimal tokens-per-parameter is task-dependent, with memorization favoring more parameters and reasoning favoring more data. Follow-up work on MoE observes that at fixed TPP, pushing total parameters higher actually degrades reasoning, while activating more experts reliably helps it. This matters for what we are building toward. Finding a vulnerability is not a retrieval problem. It doesn't come from having memorized more CVEs; it comes from carrying a twenty-step chain of inference to the end without losing the thread. That capability does not live in total parameter count. Which brings us to this release. Total parameters appear to matter up to a threshold — enough to hold the world — after which additional capability comes from scaling elsewhere: effective depth per forward pass, and above all post-training. GLM-5.3 is our controlled experiment on that claim. Same base, same architecture, same total and activated parameters as GLM-5.2. One month of scaling long-horizon environments and RL. The gains are not marginal. Well, scaling has more than one dial. We turned the post-training one this time because it had the most slack left in it — not because the others are finished. Base model size, pretraining data, compute spent per forward pass: all of them are still on the table, and we will come back to each. What this experiment taught us is that the dials do not have to be turned together, and that the one worth turning next is rarely the one that was worth turning last. We are not done scaling. Next time, maybe mid-training, pre-training, and even more.
1
5
2,366
What if a foundation model could tell us how to modify its architecture to boost inference and reasoning instantaneously—without retraining? What if that tweak incurred near zero latency cost during generation and supported indefinite state tracking? arxiv.org/abs/2608.17981
4
1,924
This has been a long time in the making — I remember Mike telling me about this idea over a year ago. Simple trick, big results, and cool paper.
What if a foundation model could tell us how to modify its architecture to boost inference and reasoning instantaneously—without retraining? What if that tweak incurred near zero latency cost during generation and supported indefinite state tracking? arxiv.org/abs/2608.17981
65
14,582
Thomas Kipf retweeted
I feel like right now is one of the best times ever to be a machine learning researcher. The pace of progress and modern tools and compute available to us is unprecedented. With the field moving this quickly, it almost feels irresponsible to not take part or do something else.
16
17
393
25,562
Thomas Kipf retweeted
if you haven't given the new Gemini 3.7 Flash a fair shake, you're seriously missing out when you factor in token throughput, 3.7 flash is actually in a league of its own. it's seriously not even close coding benchmarks will put its total task time near frontier models like 5.6 sol, but in the planning/discussion stage gemini 3.7 flash is straight up 2-6x faster than comparable options. speed at this stage keeps you in flow and leads to better alignment with your agents and better execution gemini 3.7 as your "front door" model to clarify your scope with fable 5/gpt 5.6 sol as an advisor is a really ergonomic and effective setup, it's what i'm currently using
17
9
216
32,354