Computer Vision • Multimodal AI • @huggingface Fellow ML🤗 • Computational Intelligence • Diffusion-Driven Adapters • hf.co/prithivMLmods

India
Qwen3-VL-Video-Grounding Demo. Perform point tracking, text-guided detection, and video question answering, all powered by the Qwen3-VL-4B vision-language model with real-time bounding box detection and cross-frame object matching. 🤗 @huggingface Demo in 🧵
4
64
472
50,988
Prithiv Sakthi retweeted
Introducing Naive-N0.5-Flash: Building Frontier AI with AI 🔹 309B MoE, 15.5B active: top-tier in coding, leading in AI R&D. 🔹 Native 1M context, no full-attention layers (hybrid SWA + DSA). 🔹 Inference runtime built by AI: up to 2,000 tok/s in Ultrafast mode. Weights are open today under MIT license. 🔗 Tech blog: naive.ai/en/research/ 🤗 Hugging Face: huggingface.co/NaiveAI/Naive… 💻 GitHub: github.com/NaiveAI-Labs/Naiv… 🌐 naive.ai/en/
113
102
760
109,598
Prithiv Sakthi retweeted
box it, and it's gone 🗑️ an object-remover LoRA by @prithivMLmods for Qwen-Image-2.1, stacked with Viggle's turbo: draw a box over the tourist, the bike or the boat, and it's erased in 6 steps ▶️ on Spaces hf.co/spaces/hugging-apps/qw…
5
10
75
3,770
This tweet is bait!
Local models are useless
1
117
Prithiv Sakthi retweeted
Qwen-Image-2.1 LoRA PnP: Plug and play any Qwen-Image-2.1 LoRA adapter in the app. It supports standard inference, 4-step Turbo inference, custom LoRA lazy repacks, and more, all in one setting! 🤗Space: huggingface.co/spaces/prithi…
1
2
229
Prithiv Sakthi retweeted
Unsloth has surpassed 500M model downloads on Hugging Face! 🦥🤗 Qwen3.8-27B GGUF is already Unsloth’s #1 most-downloaded model ever. Thanks for all your support!
56
68
1,065
48,649
Prithiv Sakthi retweeted
Congratulations to Unsloth on this tremendous milestone! 🎉 📦Unsloth’s most downloaded model is Qwen3-Coder-30B-A3B-Instruct-GGUF, with 28.4M+ downloads! huggingface.co/spaces/strang…
Unsloth has surpassed 500M model downloads on Hugging Face! 🦥🤗 Qwen3.8-27B GGUF is already Unsloth’s #1 most-downloaded model ever. Thanks for all your support!
1
5
204
Prithiv Sakthi retweeted
Transformers has supported loading GGUF files for a few years now, by unquantizing them. Thanks to @_marcsun, we're now using GGML kernels through the `kernels` library to run at the same performance as llama.cpp Huge kudos to the entire @ggml_org for making these kernels!
8
16
54
9,734
Prithiv Sakthi retweeted
MiMo-V2.6: The Hard Road to Scaling Up RL MiMo-V2.6 is very likely one of the largest single RL runs, by compute, that any open-source model team has undertaken to date. In an era when compute is brutally scarce, we still chose to dedicate a team of several dozen people to one goal over an extended period: scaling up RL. That takes more than research conviction. It takes a vision for AGI, respect for the unknown, and the nerve to walk straight into the hardest problems. The result is a model whose potential was built through mid-training and unlocked through heavy RL. Today, it is the number one open-source model. I strongly recommend reading the technical report. I believe it will become one of those papers that Agent RL practitioners keep reopening and discovering something new in each time. In my view, the research innovations and engineering challenges behind it surpass those of DeepSeek R1, which I was partly involved in. Some will ask: why MixRL instead of MOPD? First, they are not competing choices. We ran MixRL on verifiable tasks of moderate difficulty, including code and related agentic tasks, and found that the resulting models generalize remarkably well. Second, tasks that are difficult to verify, extremely long-horizon, or simply too challenging to include in a joint RL run are trained separately. Including them would substantially reduce rollout efficiency or introduce significant rollout staleness. We then merge the resulting capabilities through MOPD. Games, 3D tasks, and tasks with subjective evaluation signals all fall into this category. There is also a third, slightly cheeky answer. Our team is flat enough and free enough of organizational silos that MixRL simply is not difficult for us. More importantly, everyone enjoys working this way. People from different domains come together every day, driven by the pursuit of AGI and intelligence that can continuously improve itself, to confront and resolve the RL bottlenecks in each field. I will always remember the RL daily update meetings from this period. They were intense and dense, with intelligence emerging in real time. To help the open-source community focus on solving real Agentic RL problems, we have released a Qwen model distilled from MiMo RL trajectories as a stronger starting point for RL, along with 7K diverse environments and a complete RL training framework. We hope these resources will help move Agentic RL research forward. MiMo-V2.6 is only the beginning. In an era when intelligence is easy to replicate, we still choose the hard road toward self-improvement and AGI. Much of what lies ahead remains unknown. But we are willing to keep investing the time, compute, and passion required to take on one hard problem after another and work each of them all the way through, until intelligence crosses into a new regime.
235
357
4,318
652,060
Prithiv Sakthi retweeted
We’re open-sourcing the Ming-Image-0.1-Design family: • Ming-Image-0.1-Design, 6B • Ming-Image-0.1-Design-Layer, 6B • Two open-source Agent Skills: the Ling UI Design Skill and the Image-to-Editable-PPT Skill Ming-Image-0.1-Design ranks #1 among open-weight models on Artificial Analysis’s UI/UX Design leaderboard. 🧵
39
155
1,289
343,285
Prithiv Sakthi retweeted
OmniEdu: open 4B/9B/27B models for K-12 learning and teaching A family of foundation models that goes beyond problem solving to curriculum grounding, diagnostic reasoning, and pedagogical tutoring. Trained on 69,999 examples and 15.96M supervised tokens. From Solver to Tutor.
3
13
65
3,193
Prithiv Sakthi retweeted
Qwen-Image-2.1 works great in Unsloth Desktop via INT8 / FP8 and GGUFs! Pinned offloading to RAM also allows INT8 / FP8 to fit in under 6-8GB of VRAM, and is still relatively fast! We also made some dynamic GGUFs for it as well!
Qwen-Image-2.1 can now run locally on 12GB VRAM with Unsloth GGUFs! 🖼️ The 7B model performs on par with Nano Banana 2.0. For higher quality, you can also run Dynamic FP8 on just 6GB of VRAM via offloading. GGUF: huggingface.co/unsloth/Qwen-… Guide: unsloth.ai/docs/models/qwen-…
2
7
50
4,117
Prithiv Sakthi retweeted
MiMo-V2.6-Pro debuts as the top open weights model on the Artificial Analysis Intelligence Index (46). At $0.13 per Intelligence Index task, it lands on the Intelligence vs. Cost per Task Pareto frontier @Xiaomi has just released MiMo-V2.6-Pro, an open weights model with major advances in intelligence over its predecessor, MiMo-V2.5-Pro (Intelligence Index: 26). Despite the improvement, it retains the same attractive pricing at $0.435 per 1M input tokens (with a 99% cache-hit discount) and $0.87 per 1M output tokens. This makes MiMo-V2.6-Pro one of the most cost-efficient models to deploy. MiMo-V2.6-Pro is an MoE model with 1.02T total parameters and 42B active parameters. Stay tuned for additional analysis of the model. Check out MiMo-V2.6-Pro full benchmarking breakdown here: artificialanalysis.ai
175
391
4,142
895,303
Prithiv Sakthi retweeted
MiMo-V2.6 is out and the results seem to put it right after GPT-5.6 Sol 🥹 same arch as V2.5, comes with 1M context window and MTP drafter waiting for AA results 👀 huggingface.co/XiaomiMiMo/Mi…
6
8
133
12,016
Prithiv Sakthi retweeted
MiMo-V2.6-Pro: 1.02T-42B active MiMo-V2.6-Flash: 310B-15B active Top open model on Artificial Analysis + RL dashboard + tech report. Curious how adoption lands!
31
50
622
37,282
Prithiv Sakthi retweeted
Alert: Xiaomi just open-sourced MiMo-V2.6. Two natively omnimodal models, Pro and Flash, under MIT license 🔥. Pro scores 46.32 on the Artificial Analysis Intelligence Index, the highest of any open model to date. weights: huggingface.co/collections/X…
6
41
297
16,852
Prithiv Sakthi retweeted
Introducing Limite 1B - Violetto. A model for high-frequency mathematical intelligence.
64
226
1,818
283,441
Prithiv Sakthi retweeted
Today, we’re releasing Tinfield 1, an open-weight model for terminal work and long-horizon software engineering. Tinfield scores 33.0 on Terminal-Bench 4.0 and 62.0 on DeepSWE v1.1, exceeding Claude Opus 4.8’s published scores on both. 177B total parameters. 6.6B active per token. 256K context. Available today in three builds: Base, Compact and Mini.
31
38
429
50,128