PhD in seeing things @Princeton | prev @Meta Superintelligence Labs @GoogleDeepMind, @berkeley_ai | @SFGiants @warriors guy

Berkeley, CA
Today seems to be a fitting day for @GoogleDeepMind news, so I'm excited to announce our new preprint! Prior work suggests that text & img repr's are converging, albeit weakly. We found these same models actually have strong alignment; the inputs were too impoverished to see it!
11
25
144
36,507
Tyler Zhu retweeted
The most profound paper of my PhD so far. We truly did something special to make LMs play/explain chess, needing both architectural & algorithmic innovation (natural-language analogue of the Alphazero algorithm) + applicable to many other domains. Please read on and share! 1/10
56
167
1,508
101,917
Tyler Zhu retweeted
Right-Click Anything: a little demo combining Muse Spark, SAM 3.1 and Muse Image -- now all on @Meta Model API! Muse Spark captions the scene and generates noun-phrases. These are used to prompt SAM 3.1 which predicts instance masks for each object. Then, right-click any object mask as you would with text on any web page! “Look Up” hits Muse Spark web search for real product listings, “Copy” gives you the alpha-matted PNG, “Remove” paints the SAM 3.1 mask into the image and asks Muse Image to inpaint it away. SAM 3.1 also supports video so many more possibilities there too! Have fun combining models across Meta's API platform!
4
11
141
8,677
Tyler Zhu retweeted
given a choice between steak and agi the group chat chose steak (non beef eaters chose the side of potatoes)
2
1
14
1,740
“The incentives may now be pointing in the direction of no longer sharing any promising research directions with the broader community, which would reverse centuries of traditions of open science and do serious long-term damage to the future of the field” - Terence Tao
The irony of this drama would be that math becomes dead not by any AI model per se, but by researchers corrupting the very _ethos_ of academic collaboration that has kept math alive in spirit.
2
1
19
1,200
The irony of this drama would be that math becomes dead not by any AI model per se, but by researchers corrupting the very _ethos_ of academic collaboration that has kept math alive in spirit.
drama summary for those confused: - Aug 15th: Tristan Buckmaster & Levent Alpöge make progress on a few important math problems - they do NOT have a proof for the $1,000,000 Millenium Prize problem. BUT, they do claim to have a proof for a similar (non-Millenium) Navier Stokes problem that could help lead the way there - Levent works at Anthropic, but this research was independent of his work there, with a mix of GPT and Claude models. Tristan is not related to Anthropic. - Early Sep: Rumor spreads to OpenAI that Anthropic has solved a major problem. Tristan emails OpenAI to clarify. without revealing the problem they solved or how they did it. - After hearing of the rumor, OpenAI started researching Navier Stokes with a new internal model. - Sep 6th: OpenAI's Sebastien Bubeck tells Tristan that they solved the $1,000,000 Millenium Prize Navier Stokes problem. The approach is very similar to Tristan & Levent's approach to the non-Millenium problem. - Tristan is suspicious of the timing, as only few others were trying this approach. OpenAI says the model didn't access his user data directly, but leaves unanswered whether Tristan's chat conversations were part of the training. - OpenAI says they would partially credit Tristan for the $1,000,000 discovery (even though Tristan did not solve the $1,000,000 problem) — but only if they remove Levent as an author, as he works for Anthropic. - Sep 8th: Tristan refuses to remove Levent, and rushes to publish their results independently. Currently unclear is whether Anthropic had a separate solution for the $1,000,000 problem, or whether the rumor was about Tristan & Levent's independent research.
2
1
53
3,983
Tyler Zhu retweeted
We live in a multimodal world. We see, talk, act, and dream. Yet most LLMs still start with language pretraining. Why not train them natively with multimodal I/O from scratch? Because it’s SUPER HARD, adding modalities triggers training instability, design complexity, and often modal competition So what’s the path forward? Introducing: Towards Physics of Multimodal Pretraining (junlinhan.github.io/projects…) We unpack the underlying mechanics of multimodal pretraining across 4 aspects: Knowledge Flow, Modality Synergy, Early Unification, and Recipe.
28
171
1,031
109,357
Tyler Zhu retweeted
>the truth is that the vast majority of what we call thinking is also just regurgitating or recombining thoughts we've seen already, what frontier models can already do is simply not that different from what we do nearly all the time Fully agree with this.
there are people who have load-bearing beliefs about human intelligence generally and their intelligence specifically that increasing AI capabilities render increasingly obviously fake, and they are not going to have a good time. the truth is that the vast majority of what we call thinking is also just regurgitating or recombining thoughts we've seen already, what frontier models can already do is simply not that different from what we do nearly all the time. your highly original thoughts that demonstrate the depth and breadth of your soul are mostly warmed-over reheated leftovers of leftovers of thoughts that were first articulated centuries if not millennia ago. this is completely fine and ordinary and not a problem, unless you were basing your ego defenses on this very popular but largely fake misunderstanding of what thoughts are and where they come from. you cannot own a thought. they were never yours to accrete your sense of self around. they are farts that gently pass through the butthole of your mind, full of sound and fury, signifying nothing
31
61
1,057
65,475
Tyler Zhu retweeted
Can video generation models do for vision what LLMs did for language? Introducing GenCeption from @GoogleDeepMind: one feed-forward video model for various vision tasks — SOTA, data-efficient, and emerging behaviors (ECCV 2026) 🌐 genception.github.io (1/8)
17
99
468
118,032
Tyler Zhu retweeted
I'll be at #ICML2026 to present our *spotlight* work! TLDR: LLMs can learn to self-generate curricula for problems they can't yet solve, using self-play with meta-RL. Please reach out to chat about self-improving agents, synthetic data & environments, curriculum learning, or anything else! We've updated the paper with some fun additions ⬇️
Can a model learn to break its own reasoning plateau? In our new paper, we show that LLMs can be taught with meta-RL to generate their own "stepping stones" that kickstart learning on hard math problems (0/128 success rate) where direct RL fails. Paper 📝: arxiv.org/abs/2601.18778 Blog post 🌐: ssundaram21.github.io/soar/ (1/n)
7
21
150
19,138
Tyler Zhu retweeted
Presenting this work today at @CVPR from 11:45 AM – 1:45 PM, Poster 451! Swing by if interested, would love to chat! #CVPR2026
What if the best visual reasoning steps are ones humans can’t specify? 🤔 Existing VLM reasoning is often constrained by language, pixels, and human-designed intermediates. We introduce Latent Implicit Visual Reasoning, where we show that VLMs can discover the best visual reasoning steps by themselves — no bboxes, no intermediate images, no extra supervision. Presenting this week at CVPR! (1/n)🧵
1
4
12
2,385
Tyler Zhu retweeted
Congrats to the team for wining CVPR Best Paper Award!! 🏆 Come to our oral session (Mile High Ballroom 13:00-14:15) and poster (16:00-18:00) today for more details 🚀
A SINGLE encoder + decoder for all the 4D tasks! We release 🎯 D4RT (Dynamic 4D Reconstruction and Tracking). 📍 A simple, unified interface for 3D tracking, depth, and pose 🌟 SOTA results on 4D reconstruction & tracking 🚀 Up to 100x faster pose estimation than prior works
22
37
297
43,528
Tyler Zhu retweeted
Excited to share ID-Sim, our identity-focused similarity metric, presenting at #CVPR2026 this week in Denver! 🎉 Humans are remarkably good at distinguishing highly similar objects across different contexts. We asked: can we train a metric that does the same?
2
12
46
7,603
Tyler Zhu retweeted
I will be presenting BeyondObjects this week at #CVPR2026. Come check out our poster on Thursday 4:30-5:30pm at the syndata4cv workshop (Room 607) or Saturday 4:45-6:45pm at ExHall A & F poster location 99. Happy to chat!
Text-to-image (T2I) models can generate rich supervision for visual learning but generating subtle distinctions still remains challenging. Fine-tuning helps, but too much tuning → overfitting and loss of diversity. How do we preserve fidelity without sacrificing diversity (1/8)
1
5
17
3,071
Tyler Zhu retweeted
FLUX.2's @bfl_ml text tokens aren't just holding your prompt. During image editing, they absorb reference image content, and some of that absorbed content, like color and style, causally drives the output appearance. New paper 🧵👇
7
37
207
30,117
What comes after today’s visual backbones? At T4V @CVPR 2026, we’re bringing the community together for a focused half-day workshop on Transformers for Vision and Multimodal AI — covering image, video, 3D, MLLMs, efficient attention, SSMs/Mamba, and the next generation of visual architectures. 📍 Wed June 3 · Room 607 🕐 1:45–5:40 pm (Denver local time) Invited speakers: @RanjayKrishna, @thoma_gu, @sherryyangML, @jcniebles, @liuzhuang1234, and @TongPetersb. Join us at CVPR: sites.google.com/view/t4v-cv… @UNC @NVIDIAAI @NVIDIAAIDev @AIatMeta @ImagineEnpc @NJU1902 @BaskinEng #CVPR2026 #T4V #MultimodalAI #NVIDIA
12
42
7,179
Tyler Zhu retweeted
Popular self-supervised models like V-JEPA v2 avoid pixel reconstruction, claiming it's "irrelevant" for high-level semantics. But what if pixels were never a problem? Introducing Recurrent Video Masked Autoencoders (RVM) to be presented at #CVPR2026 arxiv.org/abs/2512.13684 🧵
2
29
194
26,264
We'll be presenting this work on video-text alignment this week at #ICLR2026! Come check out our poster in Pavillion 4 #3405 in the Thursday AM session from 10:30am to 1pm. Happy to chat about PRH, repr learning, video understanding, or how vision plays into text models!
Today seems to be a fitting day for @GoogleDeepMind news, so I'm excited to announce our new preprint! Prior work suggests that text & img repr's are converging, albeit weakly. We found these same models actually have strong alignment; the inputs were too impoverished to see it!
2
5
40
4,475
Tyler Zhu retweeted
How should we talk about LLMs? Does it matter if we frame them as a machines 📠, tools ⚒️, or companions 👥? In our #CHI2026 paper, that these framings can alter what people believe about LLMs and how they use them. See 🧵for more!
5
16
49
5,869
[1/9] Glad to share another project I have done during my time at @GoogleDeepMind, which actually was finished a half year ago. Shall we use sequential sampling or parallel sampling for test-time-scaling? We find that parallel sampling typically tends to be better than sequential sampling (by prompting LLMs to self-correct or continue solving the problems) in thinking models. We have three hypotheses and found that the less exploration in sequential sampling is the key. We hope to shed the light on LLMs creativity! 🧵roll ...
3
24
169
19,547