AMI Labs | ex-Meta FAIR | @Princeton CS '19 Building the next revolution of AI models that understand the real world.

New York City
Pinned Tweet
[1/9] What happens when you treat vision as a first-class citizen during multimodal pretraining? To find out, we studied the design space of training Transfusion-style models that input and output all modalities, from scratch. Here is what we learned about visual representations, data, world modeling, architecture, and scaling behavior! Paper: arxiv.org/abs/2603.03276 Website: beyond-llms.github.io/ @TongPetersb, @DavidJFan, @__JohnNguyen__, @ellisbrown, @GaoyueZhou, @JasonQSY, @boyangzheng, @webalorn, @han_junlin, @rob_fergus, @NailaMurray, @gh_marjan, @ml_perception, Nicolas Ballas, @_amirbar, Michael Rabbat, Jakob Verbeek, @LukeZettlemoyer, @koustuvsinha, @ylecun, @sainingxie
13
60
310
76,906
Nice work Junlin!!
We live in a multimodal world. We see, talk, act, and dream. Yet most LLMs still start with language pretraining. Why not train them natively with multimodal I/O from scratch? Because it’s SUPER HARD, adding modalities triggers training instability, design complexity, and often modal competition So what’s the path forward? Introducing: Towards Physics of Multimodal Pretraining (junlinhan.github.io/projects…) We unpack the underlying mechanics of multimodal pretraining across 4 aspects: Knowledge Flow, Modality Synergy, Early Unification, and Recipe.
1
20
2,041
Nice work!! Really glad to hear WebSSL works well for robot policies 😁 Representation matters
Your policy doesn't need 7B params. It simply needs dense features. Introducing Patch Policy: pretrained ViT + small transformer beats OpenVLA-OFT with 0.7% of its params, and trains on a 5090. Here it inserts a cable (~2mm tol), and does it again as we unplug mid-rollout. 🧵
2
2
23
4,845
David Fan retweeted
Our world modeling team at FAIR is hiring! Do reach out / spread the news!
World models are one of the most exciting frontiers for robotics and embodied AI right now and we're growing the team at Meta FAIR to push on it. Hiring Research Scientists in: 🇺🇸 Menlo Park 🇨🇦 Montreal Links below 👇 or find me at #RSS2026
1
3
77
12,376
David Fan retweeted
(1/n) Excited to share our new paper from @AIatMeta FAIR: "EgoBabyVLM: Benchmarking Cross-Modal Learning from Naturalistic Egocentric Video Data" 👶Human infants learn language from sparse, noisy multimodal input. Today's VLMs can't. We built a benchmark + challenge to close that gap. w/ @rust_phillip, @angelvillar96, @juanmiguelpino, @mcxfrank, Emmanuel Dupoux and many other amazing collaborators. 🧵
2
11
60
14,465
David Fan retweeted
In Oct last year, Representation Autoencoders provided an elegant solution to unified tokenization for understanding and generation. Today we make them a bit more simple. a bit more general. Result: >10x faster convergence, better reconstruction, better generation. And yes we test them on T2I and world models :) Introducing RAEv2
20
55
706
2,296,940
David Fan retweeted
Today we released the code for our CVPR 2026 paper, Flowception. Flowception bridges fully bidirectional sequence modeling and autoregressive generation by inserting frames via learned order, then denoising them with continuous flow. Website: flowception-meta.github.io Code: github.com/facebookresearch/…
I thought the path to variable-length video was frame-wise autoregressive with complex forcing schedules, but I was wrong! The solution is simple! Flowception, using frame insertions to model any-order video generation.✅ Arxiv: arxiv.org/abs/2512.11438 Project page: flowception-meta.github.io Great collaboration with Jakob, @inthebrownbag, and @RickyTQChen on this work, and led by Tariq!
4
16
93
19,257
Congrats Dr. @TongPetersb!! Your research journey has clearly culminated in a very cohesive and inspiring narrative that unifies several areas of work with much scope for future expansion :D It's a testament to your work ethic, good taste in problems, and attention to detail. I'm so proud of you! Very lucky to learn from you
Congrats Dr. Tong! Really glad to be a part of your PhD journey @TongPetersb
36
12,328
Check out the code and data for DexWM! Training + Inference Code: github.com/facebookresearch/… Robocasa Data: huggingface.co/datasets/face…
The code for DexWM is now publicly available: github.com/facebookresearch/…. The repository includes the full training and evaluation pipelines, along with custom dexterous manipulation datasets generated in RoboCasa, making it easy to reproduce our results and build on top of this work.
1
20
3,535
This week I joined AMI Labs as a founding member of the NYC lab! I'm super excited to build AI systems that truly understand the physical world. I'm also excited to help build the research agenda, culture, and team from the ground up, and learn what it takes to build a company. It's a real privilege to be here. The last ~2 years at FAIR have been the most rewarding of my professional life so far. When I first started doing research almost 10 years ago, joining FAIR felt like a pipe dream. I became a better researcher, open-sourced for the first time (V-JEPA 2, WebSSL, MetaMorph, DexWM), grew as a person, made life-long friends, and even rekindled old hobbies — like playing clarinet with the Meta NYC orchestra. I want to thank Mike Rabbat, Maryam Fazel-Zarandi, the JEPA team, and all my amazing colleagues + collaborators who made this dream come true. The research world is small, so I know our paths will cross again. Please stay in touch!
Advanced Machine Intelligence (AMI) is building a new breed of AI systems that understand the world, have persistent memory, can reason and plan, and are controllable and safe. We’ve raised a $1.03B (~€890M) round from global investors who believe in our vision of universally intelligent systems centered on world models. This round is co-led by Cathay Innovation, Greycroft, Hiro Capital, HV Capital, and Bezos Expeditions, along with other investors and angels across the world. We are a growing team of researchers and builders, operating in Paris, New York, Montreal and Singapore from day one. Read more: amilabs.xyz/ AMI - Real world. Real intelligence.
35
16
486
59,843
[1/9] What happens when you treat vision as a first-class citizen during multimodal pretraining? To find out, we studied the design space of training Transfusion-style models that input and output all modalities, from scratch. Here is what we learned about visual representations, data, world modeling, architecture, and scaling behavior! Paper: arxiv.org/abs/2603.03276 Website: beyond-llms.github.io/ @TongPetersb, @DavidJFan, @__JohnNguyen__, @ellisbrown, @GaoyueZhou, @JasonQSY, @boyangzheng, @webalorn, @han_junlin, @rob_fergus, @NailaMurray, @gh_marjan, @ml_perception, Nicolas Ballas, @_amirbar, Michael Rabbat, Jakob Verbeek, @LukeZettlemoyer, @koustuvsinha, @ylecun, @sainingxie
13
60
310
76,906
[9/9] On a personal note: I grew a lot from this project. More challenging than the technical hurdles was navigating organizational dynamics, advocating the research vision, and securing the compute resources. A huge thanks to the incredible team that made this a reality, to FAIR for the support, and especially to @TongPetersb + @__JohnNguyen__. We started this journey together and stuck by each other’s side through all the highs and lows. This work builds upon all of our collective experiences in and long-term beliefs about visual representations + multimodal modeling. I see this as just the first step in a long-term research agenda. I couldn't ask for a better partnership!
2
1
16
1,331
Also many thanks to the numerous people who gave feebdack on the work and supported us through this journey! @JimmyTYYang1, @sharut_gupta, @JiachenAI, @jiawzhao, @Hu_Hsu, @em_dinan, @XiaochuangHan, @TusharNagarajan, @TianhongLi6, @jilin_14, @Aniket_d98, @shushengyang, @jihanyang13, @WeijiaShi2, @gene_ch0u, Daniel Bolya, just to name a few
16
1,227
David Fan retweeted
Humans communicate through language and interact with the world through vision, yet most multimodal models are language-first. What happens when we go beyond language? 🤔 Beyond Language Modeling: a deep dive into the design space of truly native multimodal models Paper: arxiv.org/abs/2603.03276 Project: beyond-llms.github.io/
10
37
202
40,546
David Fan retweeted
Train Beyond Language. We bet on the visual world as the critical next step alongside and beyond language modeling. So, we studied building foundation models from scratch with vision. We share our exploration: visual representations, data, world modeling, architecture, and scaling behavior! [1/9]
35
219
1,048
219,751