We are releasing WorldDiT, a unified architecture for robotics world modeling and control. On the LIBERO benchmark, it performs the best among all publicly released methods that do not need a VLM to generate actions. Its size and performance sit on the reported Pareto frontier.
28
54
154
70,872
ROSCon is in Toronto this week. Tonight, we’re hosting Toronto Robotics Night for people here for the conference and the local robotics community. Come by after the sessions for drinks and appetizers. We start at 6:30, a 5 minute walk from the ROSCon venue. RSVP to join: luma.com/501ufk17
4
8
1,255
bagel.com retweeted
we're hosting a ROSCon special edition of Bagel's Toronto Robotics Night this week! we're gathering people who are globally shaping physical AI in one room to wrap up the conference together for researchers, founders, robotics engineers, and builders RSVP here: luma.com/501ufk17
4
3
8
252
bagel.com retweeted
The first toronto robotics night was a grand success!
Some of the most interesting work in robotics now sits between world models and real machines. We’re bringing the people working on both together this Thursday at the first ever Toronto Robotics Night. Come by if you're in town : luma.com/83qrbegr
1
4
24
1,911
bagel.com retweeted
guess who's back
1
2
7
286
Some of the most interesting work in robotics now sits between world models and real machines. We’re bringing the people working on both together this Thursday at the first ever Toronto Robotics Night. Come by if you're in town : luma.com/83qrbegr
3
4
18
5,744
bagel.com retweeted
Hardware is a big constraint in robotics, that's why we don't think the answer will be increasingly larger VLM models as action backbones. We tried to solve that issue with WorldDiT, which is at the pareto frontier for size-performance. Now I want to see how far we can scale it without losing this tradeoff. Here's a simulation demo.
We are releasing WorldDiT, a unified architecture for robotics world modeling and control. On the LIBERO benchmark, it performs the best among all publicly released methods that do not need a VLM to generate actions. Its size and performance sit on the reported Pareto frontier.
8
9
85
12,280
bagel.com retweeted
The new WorldDiT robotics model is really cool: It's very small (< 1B parameters), yet it can perform prediction and control in a single model. • It can predict how the world will change • It then decides what the robot should do Models are getting smaller and more capable. We are doing more with way less.
We are releasing WorldDiT, a unified architecture for robotics world modeling and control. On the LIBERO benchmark, it performs the best among all publicly released methods that do not need a VLM to generate actions. Its size and performance sit on the reported Pareto frontier.
6
6
23
15,206
We are releasing WorldDiT, a unified architecture for robotics world modeling and control. On the LIBERO benchmark, it performs the best among all publicly released methods that do not need a VLM to generate actions. Its size and performance sit on the reported Pareto frontier.
28
54
154
70,872
During inference, WorldDiT acts for a few steps, observes what changed, and replans. World modeling stays in training, while light-weight deployment remains action only.
1
8
584
Today we're releasing Paris 2.0, to our knowledge the first decentralized-trained video generation model. At Bagel Labs, we believe frontier models should not require homogeneous clusters of premium, supply constrained GPUs. Paris 1.0 proved this for image generation. Paris 2.0 extends the recipe to video generation and lays the substrate for global-scale world models. To test the approach, we trained two models head-to-head in an iso-FLOP, iso-data comparison. One was a monolithic model trained conventionally, on a single premium GPU cluster. The other was Paris 2.0, trained across an extreme mix of GPU types, generations, and vendors distributed around the globe. Against the monolithic model under matched data and compute, the results were: FVD: 561.04 → 279.01 (a ~2x improvement) CLIP text-video alignment and aesthetic score both improved. To our knowledge, this is the first distributed training architecture to surpass its monolithic counterpart under matched data and compute. Technical Report: arxiv.org/abs/2605.26064 Model Weights: huggingface.co/bageldotcom/p…
We're releasing Paris 2.0, which, to our knowledge, is the world's first decentralized trained video generation model. We benchmarked it against a monolithic model trained on the same data and compute budget, and Paris 2.0 outperformed the monolithic by ~2x on FVD benchmark.
6
6
35
7,876
bagel.com retweeted
We're releasing Paris 2.0, which, to our knowledge, is the world's first decentralized trained video generation model. We benchmarked it against a monolithic model trained on the same data and compute budget, and Paris 2.0 outperformed the monolithic by ~2x on FVD benchmark.
88
112
661
457,645
In town for NVIDIA GTC? If you're building generative world models or investing in the people who are - we're putting the right people in one room for you tomorrow night in Palo Alto. Co-hosted by Alumni Ventures. Signup link below.
5
6
21
11,558
bagel.com retweeted
we managed to extend DDM to heterogeneous objectives! one step closer to decentralized AI XD tldr: experts trained in complete isolation, with different objectives, no communication - and mixing them beats making them all train the same way. 20-48G memory per expert
Excited to share that Bagel Labs' paper got accepted at CVPR 2026. A lot of the most important diffusion model research has historically stayed inside frontier labs. We're bringing more of that in the open through open science and open infrastructure. In this work we showcase the very counterintuitive advantage of mixing different training objectives (DDPM and Flow-Matching) through an ensemble of diffusion models. This is one of the first ever works to successfully combine diffusion models trained with heterogeneous objectives. See details here: blog.bagel.com/p/heterogeneo…
3
10
2,721
bagel.com retweeted
Excited to share that Bagel Labs' paper got accepted at CVPR 2026. A lot of the most important diffusion model research has historically stayed inside frontier labs. We're bringing more of that in the open through open science and open infrastructure. In this work we showcase the very counterintuitive advantage of mixing different training objectives (DDPM and Flow-Matching) through an ensemble of diffusion models. This is one of the first ever works to successfully combine diffusion models trained with heterogeneous objectives. See details here: blog.bagel.com/p/heterogeneo…
4
7
21
5,217
Diffusion models are becoming the foundation for image, video, and world models. We are hosting a founders and investors gathering on that topic during NVIDIA GTC week, co-hosted by our friends at Alumni Ventures. Mar 16, Menlo Park. Sign up below. luma.com/nvidia-gtc-generati…
3
8
18
6,013
bagel.com retweeted
Being at the frontier - by the definition of it - means creating the frontier. You don't get to be at the frontier by following someone else. And creating the frontier often means discoveries that go against the established knowledge. We recently made such a discovery about distributed diffusion model training. A common way to optimize diffusion model training is by ensuring the numerical stability of their generation paths. We found that that's not true for the most efficient distributed diffusion model training architecture. We shared what works instead in our blogpost below. blog.bagel.com/p/stability-q…
14
33
450
52,412