Labor of love from the whole team, we are really excited to bring this to you ❤️
17
458
Sumith Kulal retweeted
Introducing FLUX 3 Action. An open weights 7B World Action Model that achieves first place on the RoboLab benchmark. It outperforms the previous best open model by 6.1 percentage points while using 56% fewer parameters and running up to 3.95x faster.⁠⁠ FLUX 3 Action removes the usual trade-off between world action model performance and VLA speed: it still predicts video and actions together, but plans more than twice as far ahead and runs faster per second of robot motion than the strongest open VLA. Teams can fine-tune FLUX 3 Action on their own demonstrations to create policies for a particular robot and task. Together with @nvidia, we also integrated FLUX 3 Action natively into @huggingface's LeRobot, with fine-tuning recipes included and edge deployment on NVIDIA Jetson. Beyond robotics, we’re also seeing promising results training task-specific policies for acting in simulated environments like gaming, controlling a vehicle, computer use, and wherever else a model needs to understand a visual environment and then choose what to do next. FLUX 3 Action builds on the same image, video, and audio pretraining as FLUX 3, but uses a smaller architecture designed for practical deployment. In midtraining, we trained the model to predict actions and future frames together. We’re releasing the weights, code, fine-tuning recipe, benchmarks, and reproducible examples so researchers and developers can build on the model with their own robots, environments, and tasks (see below).
77
234
1,790
166,270
Sumith Kulal retweeted
FLUX 3 Video now has fast, precise editing. The fastest and lowest cost video editing model in the world. In our internal benchmarking, it’s also ahead of other leading models in the accuracy and precision of edits. Edit an existing video by: • Adding, removing, or replacing objects and characters • Replacing backgrounds • Editing text • Changing colors, materials, visual effects • Changing or translating dialogue with accurate lip sync • Changing the events of a clip
82
110
908
77,601
Sumith Kulal retweeted
FLUX 3 Video is #2 in the world. To celebrate, FLUX 3 Video is free to use in our playground until Sunday 16th, 11:59pm PT (link in the thread). This is just the beginning of what’s coming. Up next: 4K, video editing, and multiple images & videos as reference input.
Last week, @bfl_ai launched their first video model, FLUX 3 Video. They have been testing an update on @arena, and the new version is ranked #2 in the Text-to-Video Arena! With 1496 pts, the updated FLUX 3 Video is just 16 pts behind the #1 spot, Gemini Omni Flash (1512 pts). This model is coming soon. Congrats to the @bfl_ai team on the release!
139
185
1,783
401,874
Sumith Kulal retweeted
Last week, @bfl_ai launched their first video model, FLUX 3 Video. They have been testing an update on @arena, and the new version is ranked #2 in the Text-to-Video Arena! With 1496 pts, the updated FLUX 3 Video is just 16 pts behind the #1 spot, Gemini Omni Flash (1512 pts). This model is coming soon. Congrats to the @bfl_ai team on the release!
FLUX 3 Video is here. Serious, fun, creative, real, cinematic, whatever you need it to be. Native audio, Text to Video, Image to Video with multiple frames, video continuation, dialogue in multiple languages. Comes with Draft mode so you can explore ideas fast at a fraction of the cost. Up to 20 seconds and 1080p native. Available in the API or in your favorite tool. 2K, 4K, and Open Weights coming soon.
38
61
547
309,004
Sumith Kulal retweeted
FLUX 3 Video is out today! This model can do awesome, creative and strange things no other model in the space is capable of. I had great fun in leading this and in working with the most amazing and talented team in the world @bfl_ai! Video is the start - much more to follow!
4
11
74
2,056
Sumith Kulal retweeted
Two years later our first video model - FLUX 3 Video - is finally released🥹. This is such a strong, fun and versatile model made by the best team in the world @bfl_ai. We will add further variants and generation from references in the near future. I am also super excited about the upcoming open weight, action prediction and image variants of FLUX 3.
38
51
500
25,977
Sumith Kulal retweeted
FLUX 3 Video is here. Serious, fun, creative, real, cinematic, whatever you need it to be. Native audio, Text to Video, Image to Video with multiple frames, video continuation, dialogue in multiple languages. Comes with Draft mode so you can explore ideas fast at a fraction of the cost. Up to 20 seconds and 1080p native. Available in the API or in your favorite tool. 2K, 4K, and Open Weights coming soon.
188
363
2,821
591,491
Sumith Kulal retweeted
A preview of FLUX 3 is now available on Hermes agent for 48h. Hermes chains shots into full short films, one prompt end to end. They're running a short film contest all week: @NousResearch has the details. Or try it in person: FLUX 3 live at our SF event, this Friday.
FLUX 3 Preview is now publicly available, only on Hermes Agent and free on all Nous Portal paid subs for the next 48 hours. Create a short film with it tagging @NousResearch and @bfl_ai - the 3 best entrants by 7PM PT on August 1st will receive: 1st place: 1 year of free FLUX 3 generation (20/day) + $2,000 Portal credits + Nous hoodie 2nd place: $1,000 credits + hoodie 3rd place: $500 credits + hoodie First 100 new signups using code L1YSMYDB get a free month of Nous Portal Plus with full video gen access. First 25 upgrades using code KEQHYO3X get $20 off. portal.nousresearch.com/sign… > hermes update
46
81
797
325,191
Sumith Kulal retweeted
Black Forest Labs × Nous Research Friday July 31st, San Francisco Join us: luma.com/071qvqom
50
34
621
151,213
Sumith Kulal retweeted
This is absolutely amazing with Flux 3’s split-screen rendering. More power to storytelling! Upscaled with Topaz Astra. Prompt : Split-screen video. Two equal vertical halves. Both halves show the SAME event, at the SAME time, frame-synchronized, filmed by two different cameras. SCENE: A quiet street, daytime. A long tall hedge wall runs along the sidewalk, too tall to see over. There is one narrow gap in the hedge ahead. A woman in a red jacket walks alone along the sidewalk on the near side of the hedge. On the FAR side of the hedge is an open grass field, where a large golden dog runs. LEFT HALF — CAMERA A: ground-level tracking shot on the sidewalk, following the woman from behind at shoulder height. IMPORTANT: from this camera, only the woman, the sidewalk, and the hedge wall are visible. The field, the dog, and anything behind the hedge are NEVER visible in the left half. The street looks calm and empty. RIGHT HALF — CAMERA B: aerial top-down drone shot, directly above, moving with the woman. This view shows BOTH sides of the hedge at once: the woman on the sidewalk, the hedge as a thin green line, and the golden dog sprinting across the field on the other side, on a converging path toward the gap in the hedge. ACTION — identical timing in both halves: 0–8s: The woman walks calmly. LEFT: peaceful, nothing unusual. RIGHT: the dog races closer and closer to the gap, its path and the woman's path clearly about to meet. 8–11s: The woman reaches the gap. The dog bursts through it. LEFT: the dog appears suddenly from nowhere, a total surprise; the woman flinches. RIGHT: the meeting looks perfectly predictable, two paths joining. 11–15s: The dog jumps up joyfully; it is her own dog greeting her. She laughs, kneels, and hugs it. Both halves show this ending. RULES: Same woman, same dog, same timing in both halves. The dog is visible ONLY in the right half until 8s. No cuts, no other people, no cars.
Made with AI
35
59
672
35,463
Sumith Kulal retweeted
It never stops being exciting what you can achieve with good data, scaling compute, and careful engineering. Incredible effort by the entire team.
Introducing FLUX 3. One multi-modal model for Image, Video, Audio and Action-Prediction. Creations are truer to life in every kind of style. FLUX 3 Video is now available in early access (link below). Jointly trained in one unified architecture, our model can be extended to predict actions for robotics. See our work with mimic and Audi in the thread.
5
4
59
5,516
Great pleasure to work with @elvisnavah & the whole @mimicrobotics team 🦾
Introducing FLUX-mimic, a next-generation Video-Action Model for general purpose dexterity, developed in partnership with @bfl_ai. Late last year we published mimic-video and introduced Video-Action Models (VAM): a new family of robotics foundation models built on top of video generation models. We showed that robot control reduces to visual prediction, and that robot capability is downstream of improvements in video modeling accuracy. The obvious implication was that advances in the video modeling frontier would directly translate to increased capabilities in end-to-end robot learning. FLUX-mimic is that thesis at frontier scale: We've applied our VAM architecture to the strongest video backbone available today, FLUX 3 from Black Forest Labs, and trained it on data from our own robots and wearables. General-purpose dexterity, running on a single GPU on premises. Because the model already understands world dynamics, it needs far fewer demonstrations to learn a new task. This is game-changing for our mission to deploy robots to factory floors, where industrial robot data is scarce and expensive to collect. We're now testing and deploying FLUX-mimic with manufacturing leaders like @Audi, on complex, multi-step manipulation long considered impossible for conventional automation.
3
1
16
872
Sumith Kulal retweeted
fascinating to watch the Black Forest Labs team casually automate audi's industrial manufacturing operations with their video models scaling laws for robotics are here
Black Forest Labs
12
27
253
38,067
Sumith Kulal retweeted
3
8
137
15,144
Sumith Kulal retweeted
Introducing FLUX 3. One multi-modal model for Image, Video, Audio and Action-Prediction. Creations are truer to life in every kind of style. FLUX 3 Video is now available in early access (link below). Jointly trained in one unified architecture, our model can be extended to predict actions for robotics. See our work with mimic and Audi in the thread.
335
877
6,254
1,214,857
Sumith Kulal retweeted
107
84
693
171,891
Sumith Kulal retweeted
Yesterday, I had the insane experience of joining the G7 summit in Evian and present our perspective on open innovation in AI to the G7+ world leaders. With openness under pressure around the world, it is vital that we preserve a culture that makes open and responsible development the norm, not the exception My full remarks below. > Thank you, President Macron, for organizing this important conversation, and for bringing together this excellent group of people. I think I represent the youngest company in this room. I am Robin Rombach, co-founder and CEO of Black Forest Labs: a 90-person frontier AI lab, built on both sides of the Atlantic in Germany and the United States. We develop “world models”. These are visual AI models that understand and simulate our complex physical world. Many of you have probably never heard of us. But if you’ve ever generated an AI image or AI video—I’m sure some of you have done so—you’ve probably encountered our research. Not far from here, at universities in Germany, our team pioneered and published many of the techniques that gave rise to visual generative AI. Our team has co-developed 3 of the top 5 most popular open generative AI models of all time, including Stable Diffusion, and the only open models as popular as DeepSeek. We’ve talked a lot about language models today. They are remarkable. But language is a compressed, and ultimately limited, representation of the world. We are building models that learn from thousands of years of video, and can reason in complex visual environments—the rich, messy, complicated world we actually live in. These are still early days, but this technology is going to transform every corner of the economy. For that reason, we believe that visual models, alongside language models, will soon be critical economic infrastructure. That’s why we are passionately committed to open innovation. Sharing our technology openly means that businesses around the world can build their own AI systems, rather than renting them from a handful of companies. Open technology is vital for transparency, competition, and strategic independence in AI. A culture of open innovation made AI possible in the first place, including the transformer paper that gave birth to today’s language models, and we need to preserve that culture. But we are fully aware that open innovation poses unique challenges, especially in image, video, and audio. Open AI models can be misused or modified to produce unlawful and deeply harmful content, and it can be difficult to withdraw these models once they're released. Our commitment is to show that innovation in visual AI can be both open and responsible, not one or the other. This is partly a technical challenge, and we are making good progress. For example, our latest open models demonstrated over 10 times fewer vulnerabilities for sexual deepfakes and child abuse material than the open models released by other Big Tech firms. But there is a role for governments too, from safety tools to evaluation standards to targeted regulation. For example, we have welcomed the leadership of the Trump Administration, as well as the European Union and the UK Government, in developing new legal strategies to combat sexual deepfakes. It’s crucial that we strike the right balance. The future of our societies depends on getting this technology out into the world safely. But a climate of fear around open technology, or a focus on suppression over diffusion, will leave the world reliant on a handful of firms for critical infrastructure. We are ready to work with you all to make open and responsible innovation the norm, not the exception.
13
19
173
20,661
Sumith Kulal retweeted
At the G7 today, sitting across from Presidents Trump and von der Leyen, our @robrombach made the case for open innovation in AI. "...It’s crucial that we strike the right balance. A climate of fear around open technology, or a focus on suppression over diffusion, will leave the world reliant on a handful of firms for critical infrastructure. @bfl_ai is ready to work with you all to make open and responsible innovation the norm, not the exception."
2
5
21
1,000
Sumith Kulal retweeted
So good to see this tradition grow stronger every year! The diffusion circle started with ~10 of us sitting on the floor of the eerily empty Baltimore Convention Center at ICML 2022, and we've kept it going ever since. We need more circles at conferences, diffusion or otherwise😁
Thanks all for coming and circling with us!
5
7
77
11,759