Today we sharing Atlas, our new multimodal world model.
One reason this model is so special to me is that it combines two core visual intelligence tasks I have worked on for over a decade: generation and reconstruction
Introducing Atlas:
The world's first multimodal world model that generates image and video frames with pixel-perfect camera control and reconstructs them in 3D.
Model the world, move the camera, and simulate space & time.
Sep 1, 2026 · 5:35 PM UTC
28
50
625
79,246
Way back in 2018 the last paper of my PhD was on image generation from scene graphs. I was very excited to be able to generate 128x128 images with a few barely recognizable objects
My new paper on generating images from scene graphs using graph convolution and GANs is up on arXiv! To appear at CVPR2018, with @agrimgupta92 and @drfeifei arxiv.org/abs/1804.01622
4
3
39
3,326
The same prompt on Atlas is way too easy and undersells the model's capabilities. And text-to-image isn't even its main capability!
1
20
766
Even earlier, I used to work on neural style transfer. This used to require specialized algorithms and bespoke models; but it was one of the earliest signs of life for the generative AI boom that was to come
1
15
592
Atlas can render images in many styles, controllable with a simple text prompt, without any special machinery to support it. Just simple algorithms scaled up.
Generality and scale over specialist methods has been the recipe for success over time; and Atlas epitomizes it.
1
10
396
Post-PhD, I got obsessed with the idea of reconstructing full 3D worlds from a single image.
In my first attempt in 2019, we were able to generate a few 256x256 images where we could move the camera maybe a foot, just enough for a whiff of 3D
Excited to share our work on view synthesis! We generate new views of a scene from a single image. The model reasons about 3D structure without 3D supervision, trained end-to-end on image pairs
Led by @oliviawiles1, with @georgiagkioxari and Rick Szeliski
arxiv.org/abs/1912.08804
2
16
597
Atlas blows the doors off this problem. From a single image it can orbit a room, move down the hallway, and imagine the next room; while modeling reflections, in 1440p, with perfectly controllable camera motions. This is the kind of result I was dreaming of way back when.
2
1
16
593
Reconstruction understands what is; generation imagines what could be. Both are needed for modeling the world, and Atlas brings them together in one unified world model.
Atlas has so many capabilities; we have just scratched the surface. Read more here:
worldlabs.ai/blog/atlas
1
16
542

































