Pinned Tweet
I'm so excited that our @theworldlabs team has achieved a major milestone today! Introducing Atlas - a first of its kind multimodal world model trained from scratch! 🚀 Atlas is capable of generating frames with pixel-perfect camera control, reconstructing large scenes from as few as one single input image, simulating space-time by reframing videos, natively outputting 3D spaces from one or more input images, composing multiple posed images into a consistent 3d world, and more! This is the best camera conditioned world model ever, opening doors to many possible use cases from VFX to robotics. I'm so so so proud of our team!♥️
Introducing Atlas: The world's first multimodal world model that generates image and video frames with pixel-perfect camera control and reconstructs them in 3D. Model the world, move the camera, and simulate space & time.
325
1,102
9,749
1,247,579
The goal of building any technology, AI included, should be bettering human lives and society.
“Any threat to human society, including existential, is within ourselves,” says World Labs Technologies CEO Fei-Fei Li as she discusses the risks surrounding AI, and the responsibility humans have in shaping how the technology is developed and used. Listen to our full interview here: bloom.bg/3V74HgP
85
151
1,003
104,489
Agents can be AI, but agency is human. It is every individual’s responsibility to empower our own agency, with the help of tools. But the North Star should always remain human centered.
189
373
2,361
162,152
Fei-Fei Li retweeted
What happens when you give a world model 32 images of a place we know very well? World Labs put Atlas to the test on Voyager. Atlas brings text, images, video, and 3D into a shared spatial context, enabling one model to generate new views, reconstruct scenes, and simulate worlds. With Voyager, you can now see that in action in real time. Pretrained from scratch on NVIDIA Blackwell GPUs. Take a look around 👇
From 32 input images to real-time flight through @nvidia's Voyager headquarters. Trained on NVIDIA Blackwell GPUs, Atlas uses these images as 3D spatial context to generate new views, letting you explore with pixel-perfect camera control. Take a look around.
30
42
293
53,879
Fei-Fei Li retweeted
From 32 input images to real-time flight through @nvidia's Voyager headquarters. Trained on NVIDIA Blackwell GPUs, Atlas uses these images as 3D spatial context to generate new views, letting you explore with pixel-perfect camera control. Take a look around.
52
73
484
219,154
Fei-Fei Li retweeted
Atlas beta access is rolling out to users, some ideas for API developers - give Astra "eyes" in Blender, let it place and render new camera views with Atlas - 3D reconstruction (this video uses 9 image inputs of my studio to make a scene, any Zillow apartment works) - graybox-to-world, use depth inputs and generate multi-view renders of a scene
20
53
594
45,186
Incredible opportunity for robotic learning researchers/engineers to join @theworldlabs! ❤️‍🔥
We're hiring in robot learning at @theworldlabs! Join me, @drfeifei, and the team to define and scale the next generation of world models for robot learning! Atlas for Robotics: worldlabs.ai/blog/atlas#robo… Real-to-Sim-to-Real: worldlabs.ai/blog/real-to-si… Apply: jobs.ashbyhq.com/worldlabs/8…
33
70
814
124,956
Fei-Fei Li retweeted
SparkJS 2.2.0 is out packing tons of bug fixes an improvements. Enjoy! 🎉 🥳 20% faster splat sorting, 14% FPS gains in some systems, 50% smaller bundle size and more. Dive into the full detailed change log github.com/sparkjsdev/spark/…
7
33
11,089
Fei-Fei Li retweeted
University leaders should make sure they reason about the broader trend. Around 2020, Fei-Fei and others at Stanford identified the growing compute divide between industry and academia, and the need for policy response in pushing for the National Research Cloud. In 2021, the same point was made when applied to model training in the foundation models paper. In 2022, the same point was made when applied to model evaluation in HELM among other works. In 2026, the same point was made when applied to model use. OpenAI spent more than $6.5 million to resolve Navier Stokes (Astra prices for 130B output tokens only). The stakes are not just the academic computer scientists doing frontier research on how to build AI models. The stakes are about the whole university system doing frontier research. Coping with "scarcity is the mother of invention" won't cut it, and we should trust the "godmother of AI" on that.
The world of R&D is forking into two paths: the token-abundant research, and the token-starved research. The future is in the former - evidentially, the progress by today's top AI industry teams and neolabs is breathtaking, where researchers’ human brilliance is super charged by AI’s assistance. Every research university president should be reading this report and reflecting on what the future of higher education research should be. openai.com/index/research-ac…
4
9
96
33,651
The world of R&D is forking into two paths: the token-abundant research, and the token-starved research. The future is in the former - evidentially, the progress by today's top AI industry teams and neolabs is breathtaking, where researchers’ human brilliance is super charged by AI’s assistance. Every research university president should be reading this report and reflecting on what the future of higher education research should be. openai.com/index/research-ac…
83
404
2,613
320,050
Fei-Fei Li retweeted
Real time streaming of next view prediction model that has accurate camera positioning and is 3D consistent. Unbelievable.
One more thing…
4
9
123
31,459
Fei-Fei Li retweeted
For those who remember @theworldlabs original RTFM demo: Atlas allows us to break out of these small spaces. Here I took the kids playroom and using Atlas-turbo prompted a new outside area into existence.
RTFM, a new generative interactive world model by World Labs, generates real-time video from a single image for exploring 3D worlds. Trained on large-scale video data to predict the next frame. Achieves unbounded persistence via spatial memory. try it: rtfm.worldlabs.ai
2
4
48
18,251
Fei-Fei Li retweeted
atlas lets you walk around any zillow listing in realtime
16
20
309
29,076
Fei-Fei Li retweeted
Our new Atlas model now runs real-time after some intense inference optimization! Generating and navigating worlds is so much more visceral when it's interactive. Inference runs on both B200 and MI355X with similar performance. Both devices are incredible workhorses!
One more thing…
6
13
152
26,812
Atlas runs on real-time!!🚀
One more thing…
44
70
842
97,764
Fei-Fei Li retweeted
World Labs co-founders Justin Johnson and Dr. Fei-Fei Li say LLMs use next-token prediction, but spatial intelligence has its own equivalent: Justin: "The soft definition of AI-completeness is there's this fundamental primitive that's an AI task. But if I could solve this AI task in its full, broadest generality, it would solve any intelligence problem." "The classic example in LLMs is that next-token prediction is AI-complete... I think from Ilya: there's a mystery novel, the thing has to read the whole novel, and the final sentence is, 'And the killer was.' Predict the next token. You could basically frame any kind of intelligence task in terms of that." "So clearly next-token prediction is something people believe is AI-complete." "New-view prediction, this primitive that we have in Atlas, especially generative new-view prediction, is also AI-complete." "I want to have a world where Martin is writing a proof of the Riemann hypothesis on the blackboard, and then the camera pans over to the next whiteboard." Fei-Fei: "Evolution had to solve new-viewpoint prediction by making animals move. Nature gave animals eyes, but nature didn't give trees eyes. Why? Because when you move, you see a new viewpoint... We do believe very strongly that next-viewpoint prediction is the equivalent of next-token prediction." @jcjohnss @drfeifei @martin_casado
World Labs co-founders Fei-Fei Li, Justin Johnson, Ben Mildenhall, and a16z's Martin Casado on Atlas, a world model for spatial intelligence: LLMs are built on next token prediction. Video models are built on next frame prediction. Atlas is built on new view prediction, and it's the first model to unify pixel generation and pixel reconstruction, two problems computer vision has kept in separate tracks for half a century. The practical result is a 50 to 100x reduction in what it takes to digitally capture a 3D representation of a space. Previously, you needed 100 to 300 photos of a single room. Atlas can work from just three. In this conversation, they get into the slow motion shot from The Matrix that took hundreds of cameras and now takes three iPhones, the overnight Slack message that made them bet the company in five seconds, why robotics is bottlenecked on data rather than chips, and the case that new view prediction is AI-complete. 00:00 Intro 01:50 The Matrix slow motion scene now takes three iPhones 02:48 Why new view prediction is the primitive 07:10 Unifying generation and reconstruction 11:15 Gaussian splats became the bottleneck 14:17 Dense capture used to mean 300 photos 17:30 Why reconstruction needs generation to fill the gaps 18:44 The LLM lesson image models missed 23:39 The video that made them go all in 28:04 3D design is 95% revisions 30:50 The problem in robotics is data, not chips 32:48 Why a robot policy can't be trained like an image model 34:44 When the simulator becomes the planner 36:45 Frozen time required footage full of movement 40:57 Why new view prediction is AI-complete 42:43 Nature gave animals eyes but not trees YouTube: piped.video/watch?v=qn1QDDBn… @drfeifei @jcjohnss @BenMildenhall @theworldlabs @martin_casado
24
37
352
88,356
Fei-Fei Li retweeted
>blogs about ray tracing for a CS project in 2012 >co-founds World Labs in 2024 >ships Atlas in 2026, a spatial intelligence model that turns a few photos into a whole 3D world curiosity compounds
World Labs co-founders Fei-Fei Li, Justin Johnson, Ben Mildenhall, and a16z's Martin Casado on Atlas, a world model for spatial intelligence: LLMs are built on next token prediction. Video models are built on next frame prediction. Atlas is built on new view prediction, and it's the first model to unify pixel generation and pixel reconstruction, two problems computer vision has kept in separate tracks for half a century. The practical result is a 50 to 100x reduction in what it takes to digitally capture a 3D representation of a space. Previously, you needed 100 to 300 photos of a single room. Atlas can work from just three. In this conversation, they get into the slow motion shot from The Matrix that took hundreds of cameras and now takes three iPhones, the overnight Slack message that made them bet the company in five seconds, why robotics is bottlenecked on data rather than chips, and the case that new view prediction is AI-complete. 00:00 Intro 01:50 The Matrix slow motion scene now takes three iPhones 02:48 Why new view prediction is the primitive 07:10 Unifying generation and reconstruction 11:15 Gaussian splats became the bottleneck 14:17 Dense capture used to mean 300 photos 17:30 Why reconstruction needs generation to fill the gaps 18:44 The LLM lesson image models missed 23:39 The video that made them go all in 28:04 3D design is 95% revisions 30:50 The problem in robotics is data, not chips 32:48 Why a robot policy can't be trained like an image model 34:44 When the simulator becomes the planner 36:45 Frozen time required footage full of movement 40:57 Why new view prediction is AI-complete 42:43 Nature gave animals eyes but not trees YouTube: piped.video/watch?v=qn1QDDBn… @drfeifei @jcjohnss @BenMildenhall @theworldlabs @martin_casado
12
23
345
72,829
Fei-Fei Li retweeted
World Labs co-founder Justin Johnson says Atlas can recreate the famous "Bullet Time" shot from The Matrix with iPhones: "The way they did that shot is they had a ring of hundreds of cameras. [Neo] fell over in the studio, they had hundreds of cameras viewing that angle on a green screen, and then they used those hundreds and hundreds of cameras to make that famous shot in The Matrix." "Now with Atlas, we can do this with as few as three cameras. No studio capture, no green screen, no expensive calibration. We can literally stick three iPhones on tripods and use these to take video of something happening." "From those three iPhone videos, we can then reframe the shot and imagine, like, freeze time, have the camera fly in as the milk is splashing up, and get these amazing frozen-time views. We can do this with just a couple cameras." @jcjohnss
World Labs co-founders Fei-Fei Li, Justin Johnson, Ben Mildenhall, and a16z's Martin Casado on Atlas, a world model for spatial intelligence: LLMs are built on next token prediction. Video models are built on next frame prediction. Atlas is built on new view prediction, and it's the first model to unify pixel generation and pixel reconstruction, two problems computer vision has kept in separate tracks for half a century. The practical result is a 50 to 100x reduction in what it takes to digitally capture a 3D representation of a space. Previously, you needed 100 to 300 photos of a single room. Atlas can work from just three. In this conversation, they get into the slow motion shot from The Matrix that took hundreds of cameras and now takes three iPhones, the overnight Slack message that made them bet the company in five seconds, why robotics is bottlenecked on data rather than chips, and the case that new view prediction is AI-complete. 00:00 Intro 01:50 The Matrix slow motion scene now takes three iPhones 02:48 Why new view prediction is the primitive 07:10 Unifying generation and reconstruction 11:15 Gaussian splats became the bottleneck 14:17 Dense capture used to mean 300 photos 17:30 Why reconstruction needs generation to fill the gaps 18:44 The LLM lesson image models missed 23:39 The video that made them go all in 28:04 3D design is 95% revisions 30:50 The problem in robotics is data, not chips 32:48 Why a robot policy can't be trained like an image model 34:44 When the simulator becomes the planner 36:45 Frozen time required footage full of movement 40:57 Why new view prediction is AI-complete 42:43 Nature gave animals eyes but not trees YouTube: piped.video/watch?v=qn1QDDBn… @drfeifei @jcjohnss @BenMildenhall @theworldlabs @martin_casado
16
26
176
44,009
Fei-Fei Li retweeted
Old footage sitting in your camera roll could now be reconstructable as a 3D scene. World Labs co-founders Ben Mildenhall and Fei-Fei Li on how Atlas got there: Ben: "In a casual sense... I took three photos of this object, or six photos of this room. I look at the photos, I can understand in my mind how those piece together. I can fill in the gaps and get it." "But there's never really been any reconciliation between those data-driven priors and the brute force dense reconstruction, which is much more akin to scientific or medical imaging... When we say dense, we really mean dense." "This room, I want like 100, 200, 300 photos to capture it. And what we're trying to do is bring that down to like three. We're saying like 50, 100x reduction." "At that scale it completely flips that calculus on its head of what type of captures you reconstruct. You can go back to existing imagery you have. You can go to stuff you find on the internet and even build scenes out of that. You can go to casual videos and unearth a lot of footage that in the past we would never have treated as reconstructable, and bring it to life as 3D." "This is something we've been playing around with a lot with Atlas. Taking old clips. I've taken a bunch of my own old captures that never worked before and put them through the system and seen a reconstruction for the first time." Fei-Fei: "The Stanford demo is underappreciated. Anywhere between 3 to 25 images, you can reconstruct that entire Stanford quad... Everything you see is generated, but according to the laws of reconstruction. And this is really magical." @BenMildenhall @drfeifei
World Labs co-founders Fei-Fei Li, Justin Johnson, Ben Mildenhall, and a16z's Martin Casado on Atlas, a world model for spatial intelligence: LLMs are built on next token prediction. Video models are built on next frame prediction. Atlas is built on new view prediction, and it's the first model to unify pixel generation and pixel reconstruction, two problems computer vision has kept in separate tracks for half a century. The practical result is a 50 to 100x reduction in what it takes to digitally capture a 3D representation of a space. Previously, you needed 100 to 300 photos of a single room. Atlas can work from just three. In this conversation, they get into the slow motion shot from The Matrix that took hundreds of cameras and now takes three iPhones, the overnight Slack message that made them bet the company in five seconds, why robotics is bottlenecked on data rather than chips, and the case that new view prediction is AI-complete. 00:00 Intro 01:50 The Matrix slow motion scene now takes three iPhones 02:48 Why new view prediction is the primitive 07:10 Unifying generation and reconstruction 11:15 Gaussian splats became the bottleneck 14:17 Dense capture used to mean 300 photos 17:30 Why reconstruction needs generation to fill the gaps 18:44 The LLM lesson image models missed 23:39 The video that made them go all in 28:04 3D design is 95% revisions 30:50 The problem in robotics is data, not chips 32:48 Why a robot policy can't be trained like an image model 34:44 When the simulator becomes the planner 36:45 Frozen time required footage full of movement 40:57 Why new view prediction is AI-complete 42:43 Nature gave animals eyes but not trees YouTube: piped.video/watch?v=qn1QDDBn… @drfeifei @jcjohnss @BenMildenhall @theworldlabs @martin_casado
9
16
176
48,805
Next view prediction is the key to Atlas, enabling us to unify pixel-level generation and reconstruction. @jcjohnss @BenMildenhall @martin_casado and I had a deeper discussion on some of the most exciting technical innovations of Atlas, our newly released world model for spatial intelligence!
World Labs co-founders Fei-Fei Li, Justin Johnson, Ben Mildenhall, and a16z's Martin Casado on Atlas, a world model for spatial intelligence: LLMs are built on next token prediction. Video models are built on next frame prediction. Atlas is built on new view prediction, and it's the first model to unify pixel generation and pixel reconstruction, two problems computer vision has kept in separate tracks for half a century. The practical result is a 50 to 100x reduction in what it takes to digitally capture a 3D representation of a space. Previously, you needed 100 to 300 photos of a single room. Atlas can work from just three. In this conversation, they get into the slow motion shot from The Matrix that took hundreds of cameras and now takes three iPhones, the overnight Slack message that made them bet the company in five seconds, why robotics is bottlenecked on data rather than chips, and the case that new view prediction is AI-complete. 00:00 Intro 01:50 The Matrix slow motion scene now takes three iPhones 02:48 Why new view prediction is the primitive 07:10 Unifying generation and reconstruction 11:15 Gaussian splats became the bottleneck 14:17 Dense capture used to mean 300 photos 17:30 Why reconstruction needs generation to fill the gaps 18:44 The LLM lesson image models missed 23:39 The video that made them go all in 28:04 3D design is 95% revisions 30:50 The problem in robotics is data, not chips 32:48 Why a robot policy can't be trained like an image model 34:44 When the simulator becomes the planner 36:45 Frozen time required footage full of movement 40:57 Why new view prediction is AI-complete 42:43 Nature gave animals eyes but not trees YouTube: piped.video/watch?v=qn1QDDBn… @drfeifei @jcjohnss @BenMildenhall @theworldlabs @martin_casado
35
115
976
109,635
More on Atlas!
Introducing Atlas: The world's first multimodal world model that generates image and video frames with pixel-perfect camera control and reconstructs them in 3D. Model the world, move the camera, and simulate space & time.
3
4
22
16,344