reflecting a bit on my (long) journey @huggingface - most big things start very small: > around 6 years ago @julien_c contacts me (on twitter) he's building a "hub for ml models community" (wtf is this). Still I feel it's interesting and I want to be part of it > day 18: we start iterating on design ideas and prototypes with the Hub team - trying a lot of stuff to understand what could work (and we are having fun doing it) > day 41: julien shares our mockups with the team, "everyone is pretty excited". 2 weeks later an offer email arrives, subject: "Offer 🤗😱" > day 167: we write "The AI community building the future", 2 days before launch, not sure anyone will care > day 169: launch day, julien: "we launch now or tomorrow?" me: "now, I'm done". we launch the new Hugging Face Hub > day 318: clem posts traffic charts in slack at 1am: "very nice growth in august!". Glimpses of something happening 👀 > day 687: it's not small anymore... also on this day dalle-mini goes viral, for a week the whole internet is on huggingface > years pass by, we keep shipping uninterrupted, the Hub team grows a bit (with exceptional people). The community shows up, 1M models, then 3M... > day 2221 (yesterday): surprise all-hands, Jensen Huang joins the call. We're joining NVIDIA.
38
19
318
61,869
Opus 5.5 designing LEGO 👀 I asked it to design a Microduck I can build with real LEGO pieces. It: > designed it life-size using 1113 real LEGO parts > verified: 3,204 connections, 0 collisions, every step buildable, centre of mass inside the feet 🤯 > made a 141-page LEGO-style booklet (237 steps) > priced every piece in the browser and prepared the orders on BrickLink
324
628
8,276
1,249,060
btw if this wasn't clear ofc I'm building it irl with my kids Claude found all pieces for me (total cost €95) - will it really work? we'll find out :)
8
155
19,378
How to get to the second floor 🦆
26
33
724
56,802
Victor M retweeted
Microsoft just released a new dataset for robot learning on Hugging Face A valuable resource for the robotics and embodied AI community. huggingface.co/datasets/micr…
6
29
2,469
Victor M retweeted
Agent Traces on the @huggingface hub now come with a receipt 🧾 Every trace shows tokens, cache hit rate, and cost for every run!
5
10
42
26,286
really cool model from Apple
Hugging Apps
Alert: Apple just dropped a new model on Hugging Face. It's a Qwen3.5-9B finetune that turns long documents into small page images to save tokens, then pulls up the full text of only the pages relevant to your question 💡 huggingface.co/apple/LensVLM…
4
10
71
6,963
Victor M retweeted
Speech recognition is easy—until you ask it to listen forever. Today we’re open-sourcing Audio8 ASR Infinite: Ultra-low latency, unlimited audio, 24/7 transcription, no drift. Built-in semantic turn detection keeps it listening like a human ear. New SOTA for streaming ASR.
6
29
338
29,658
Victor M retweeted
Today we release FLUX 3 Action - an open-weight 7B world-action model, achieving SOTA performance and efficiency on various leaderboards such as RoboLab. I am incredible proud of the team - and both backbone and embodiment-specific finetunes are open weight, alongside the training recipes!
Introducing FLUX 3 Action. An open weights 7B World Action Model that achieves first place on the RoboLab benchmark. It outperforms the previous best open model by 6.1 percentage points while using 56% fewer parameters and running up to 3.95x faster.⁠⁠ FLUX 3 Action removes the usual trade-off between world action model performance and VLA speed: it still predicts video and actions together, but plans more than twice as far ahead and runs faster per second of robot motion than the strongest open VLA. Teams can fine-tune FLUX 3 Action on their own demonstrations to create policies for a particular robot and task. Together with @nvidia, we also integrated FLUX 3 Action natively into @huggingface's LeRobot, with fine-tuning recipes included and edge deployment on NVIDIA Jetson. Beyond robotics, we’re also seeing promising results training task-specific policies for acting in simulated environments like gaming, controlling a vehicle, computer use, and wherever else a model needs to understand a visual environment and then choose what to do next. FLUX 3 Action builds on the same image, video, and audio pretraining as FLUX 3, but uses a smaller architecture designed for practical deployment. In midtraining, we trained the model to predict actions and future frames together. We’re releasing the weights, code, fine-tuning recipe, benchmarks, and reproducible examples so researchers and developers can build on the model with their own robots, environments, and tasks (see below).
23
32
346
24,428
Victor M retweeted
1/ Today we're introducing Cua-S1-4B-0.2, the first multimodal decision model trained with RLOO on live computer-use tasks, using task-completion rewards. Text and multimodal adapters are available under Apache-2.0: github.com/trycua/cua huggingface.co/cua-ai/cua-s1…
88
118
1,452
152,730
Alert: Apple just dropped a new model on Hugging Face. It's a Qwen3.5-9B finetune that turns long documents into small page images to save tokens, then pulls up the full text of only the pages relevant to your question 💡 huggingface.co/apple/LensVLM…
74
276
3,451
274,768
Victor M retweeted
Introducing FLUX 3 Action. An open weights 7B World Action Model that achieves first place on the RoboLab benchmark. It outperforms the previous best open model by 6.1 percentage points while using 56% fewer parameters and running up to 3.95x faster.⁠⁠ FLUX 3 Action removes the usual trade-off between world action model performance and VLA speed: it still predicts video and actions together, but plans more than twice as far ahead and runs faster per second of robot motion than the strongest open VLA. Teams can fine-tune FLUX 3 Action on their own demonstrations to create policies for a particular robot and task. Together with @nvidia, we also integrated FLUX 3 Action natively into @huggingface's LeRobot, with fine-tuning recipes included and edge deployment on NVIDIA Jetson. Beyond robotics, we’re also seeing promising results training task-specific policies for acting in simulated environments like gaming, controlling a vehicle, computer use, and wherever else a model needs to understand a visual environment and then choose what to do next. FLUX 3 Action builds on the same image, video, and audio pretraining as FLUX 3, but uses a smaller architecture designed for practical deployment. In midtraining, we trained the model to predict actions and future frames together. We’re releasing the weights, code, fine-tuning recipe, benchmarks, and reproducible examples so researchers and developers can build on the model with their own robots, environments, and tasks (see below).
75
234
1,785
163,911
Victor M retweeted
Unsloth has surpassed 500M model downloads on Hugging Face! 🦥🤗 Qwen3.8-27B GGUF is already Unsloth’s #1 most-downloaded model ever. Thanks for all your support!
51
68
1,059
47,399
Victor M retweeted
When several people talk at once, a transcript can get messy fast. Our new Nemotron 3 Diarization model tracks who spoke when, even when voices overlap. It handles up to eight speakers, has 100M parameters, and is now available on @huggingface 🤗
189
855
9,629
1,000,847
Trained a Jev-style classifier on @huggingface Jobs for ~$1.50. It's a 194M GLiNER2 model that suggests task tags for any Hub dataset from its column names and first row, and returns a label with a probability. Zero-shot, GLiNER2's first suggestion matched an owner's tag 10% of the time. After 17 minutes of fine-tuning: 69%. The fine-tuned model runs on a free CPU in about a second. Owners' tags are noisy, so some "wrong" answers are tags the owner left out. The recipe is open: one hf jobs command trains the same kind of model on your own labels. The README example (book titles) runs in ~2 minutes for about $0.02. Demo: huggingface.co/spaces/davans… Recipe: huggingface.co/datasets/uv-s…
10
19
139
21,037
New JS package: @huggingface/lerobot 🤖 Read LeRobot datasets on the Hub straight from the browser. No download. Point your coding agent at it and build custom viewers fast. Example 3D camera-frustum replay streams 1.1 MB of a 793 MB dataset: hf.co/spaces/mishig/lerobot-…
3
19
127
8,928
Opus 5.5 made this galloping horse (entirely in code every pixel drawn procedurally). One self-contained HTML file. Vanilla JS + Canvas 2D. No images or libraries. 128×96 pixels, articulated legs driven by inverse kinematics, 12-pose gallop... Something is happening...
55
49
1,058
139,594