DPhil in Computer Vision, University of Oxford || Clarendon Scholar @Oxford_VGG @OxfordTVG || Experience with @Snap @AIatMeta

Noxus
Runjia Li retweeted
We share H3-World 🌍 The first to turn MiniMax-H3 itself into world model. No new action module. We directly convert H3’s pretrained language understanding into world control. Only 8K samples + 0.199% Trainable Params. 📄 huggingface.co/papers/2609.0… 💻 danzer1xxxxchan.github.io/H3…
41
159
1,268
225,580
Runjia Li retweeted
We live in a multimodal world. We see, talk, act, and dream. Yet most LLMs still start with language pretraining. Why not train them natively with multimodal I/O from scratch? Because it’s SUPER HARD, adding modalities triggers training instability, design complexity, and often modal competition So what’s the path forward? Introducing: Towards Physics of Multimodal Pretraining (junlinhan.github.io/projects…) We unpack the underlying mechanics of multimodal pretraining across 4 aspects: Knowledge Flow, Modality Synergy, Early Unification, and Recipe.
28
171
1,031
109,366
Runjia Li retweeted
We trained a video world model on just 15 hours of video of a single-arm robot. 🧵 It generalizes zero-shot to unseen embodiments (and even orangutans). And it picked up something we never trained for: give it the object motion you want, and it synthesizes the robot motion that produces it. The trick is representing actions as masked video.
13
50
295
58,491
Runjia Li retweeted
🎁 We are pleased to introduce Syn4D, a large-scale multiview synthetic dataset for dynamic scenes, accepted at #ECCV2026. The complete dataset, including its dense geometric annotations, is now publicly available. 🌐 Project: jzr99.github.io/Syn4D/ 💻 GitHub: github.com/jzr99/Syn4D 🤗 Dataset: huggingface.co/datasets/Syn4… 🧵1/8
4
22
135
14,492
Runjia Li retweeted
AutoPartGen code is now available! 🚀 We are excited to finally release a reimplementation of AutoPartGen, as promised! Check out the code, models, project page, and paper here: 💻 Code: github.com/facebookresearch/… 🤗 Models: huggingface.co/facebook/auto… 🌐 Project page: silent-chen.github.io/AutoPa… 📎 Paper: arxiv.org/abs/2507.13346 We sincerely apologize for the long delay, and we greatly appreciate your patience and continued interest in our work!
4
26
209
95,671
Runjia Li retweeted
SAE interventions can be unreliable. 🧠🔒 We show that even when features are clamped, bad behaviors can still return through alternative residual-space directions. 🧩↩️ Feature control ≠ behavior control. 🚨 Paper: huggingface.co/papers/2606.1… Page+Code: mingyuee88.github.io/sae-pos…
1
9
69
6,355
Runjia Li retweeted
🚀 Introducing Instruct-Particulate, our new model for inferring articulated structures from static 3D meshes, with significantly improved generalization to novel object categories and support for kinematic prompting. To achieve this, we scaled our training data 40× and redesigned the model to follow kinematic prompts. The result: diverse, realistic, simulator-compatible articulated 3D assets can now be generated directly from real-world images! 🔗 Project page: instruct-particulate.github.… 🤗 Demo: huggingface.co/spaces/rayli/…
3
19
120
32,757
Runjia Li retweeted
🌙 Open-sourcing Yume -- a programmable, explicit world model on Godot. You build a game by describing the world as JSON -- the things in it + the rules for how they behave -- and one fixed engine runs it. github.com/kamwoh/yume
4
12
61
259,943
Runjia Li retweeted
Gamma-World Generative Multi-Agent World Modeling Beyond Two Players
8
14
102
31,990
Impressive!!
Introducing VGGT-Ω: scaling feed-forward reconstruction across static and dynamic scenes, and studying whether the learned geometric representations transfer beyond reconstruction.
1
182
Runjia Li retweeted
🚀 Introducing Articraft, a coding agent for articulated 3D asset creation. Articraft writes code, executes it, receives validation feedback, and refines the result into simulation-ready 3D assets with parts, joints, and motion. We’re also releasing Articraft-10K: 10,000+ articulated objects across 250 categories, unlocking large-scale interactive scenes for robotics simulation and physical AI. 🔗 Project page: articraft3d.github.io/ 💻 Code: github.com/mattzh72/articraf…
22
108
742
191,956
We made an interactive client-server viewer for LagerNVS with @JonathonLuiten! You can now interactively explore scenes from just a photo capture - no optimization, no 3D Gaussians, just load your images, run the model on a cloud GPU and stream the renders to your local browser. Check out the video below for some spaces I recently captured in Oxford, London and beyond!
5
26
175
17,358
Runjia Li retweeted
We’re reimagining a 50-year-old interface - the mouse pointer - with AI. 🖱️ These experimental demos show how people can intuitively direct Gemini on their screens using motion, speech, and natural shorthand to get things done 🧵
449
1,035
8,468
1,686,360
Runjia Li retweeted
We scaled up Lyra to generate explorable 3D worlds! 🚀 Introducing Lyra 2.0 — turning a single image into a 3D world you can walk through, look back, and even drop a robot into 🤖 Code and Model available today! 🌐 Website: research.nvidia.com/labs/sil… (1/N)
27
115
856
1,148,303
Introducing ActionParty: the first video world model that controls up to 7 players simultaneously on the same screen across 46 game environments. We tackle the action binding problem in video diffusion, ensuring each player's action is applied to the right subject. 🧵
6
9
53
10,085
Runjia Li retweeted
Dropping an exciting new demo of MosaicMem! 👀🔥 A friend brought up a great question: why not combine long-horizon navigation video generation, promptable world events, and scene concatenation? Fair point — so we gave it a shot. 🎬✨ For more technical details, check this thread 🧵👇 x.com/GnosisYu/status/203502… #WorldModel #GenerativeAI #VideoGeneration #InteractiveAI #Genie3 #EmbodiedAI #GameAI
World models have made impressive progress in video generation, yet they still struggle with a fundamental challenge: memory. In long rollouts, the camera trajectory gradually drifts from the user-specified motion and revisited scenes no longer align with earlier observations. These errors accumulate over time, causing the generated world to steadily lose coherence. 🚀Excited to share our solution MosaicMem 🌍🧠 — our new hybrid spatial memory for video world models. Project Page: mosaicmem.github.io/mosaicme… Paper: huggingface.co/papers/2603.1…
21
107
8,905
🎉EgoEdit @Snapchat has been accepted to CVPR 2026! 🏆👻 We are bringing high-quality, real-time editing to egocentric videos. Our massive 100k video dataset and benchmark are ALREADY PUBLIC! 🔓🚀 🏠 Project Page: snap-research.github.io/EgoE… 🤗 Dataset: huggingface.co/datasets/ligu…
EgoEdit Dataset, Real-Time Streaming Model, and Benchmark for Egocentric Video Editing
5
8
105
21,973
The work was done in a joint collaboration with @WilliMenapace during my internship @Snap. Many thanks to @moayedhajiali, @ashmrz10, Chaoyang Wang Arpit Sahni, @isskoro, Aliaksandr Siarohin, @JakabTomas, @han_junlin, @SergeyTulyakov, @philiptorr
4
372
Replying to @Snapchat
Many thanks to coauthors! And thank @_akhaliq for posting our paper!
1
236
Runjia Li retweeted
Mode Seeking meets Mean Seeking for Fast Long Video Generation paper: huggingface.co/papers/2602.2…
5
18
119
20,418