open weight is the only way Enthusiastic about AI! Whee! leaky faucet message me! prev @MetaAI, @AMD, @ServiceNow

/mat/mul
shaun retweeted
not bad for our little 7B
Qwen-Image-2.1 by @Alibaba_Qwen just landed as the #1 open source model in the Image Edit Arena and Text-to-Image Arena! With 1367 pts in the Image Edit Arena, Qwen-Image-2.1 took the #1 spot among open. It landed #16 overall, just 3 pts from GPT-Image-1.5-high-fidelity at #15. See the leaderboard for the Text-to-Image arena below. Congrats to the @Alibaba_Qwen team on this contribution to the open source ecosystem!
40
65
1,774
76,089
shaun retweeted
Introducing Qwen Intelligence, bringing personal intelligence within everyone's reach. 📱✨ It launches with three SOTA agents: 🥳 - Mobile Planner Agent: plans, decomposes & orchestrates complex tasks. #1 on MobilePA-Bench, MobilePA-Bench Business & Memory. - Mobile-Use Agent: gets things done, API-first with GUI fallback. MobileWorld 82.1, MobileWorld-Real 92.2, AndroidDaily 97.2, 90% end-to-end success rate. - Mobile Creative Agent: turns one sentence into ready-to-use creations. Image generated in 3s, about 2x faster than leading peers. We're also opening up our benchmark suite: MobilePA-Bench, MobileWorld, MobileWorld-Real, and MobileWorld-Safety, covering planning, cross-app execution, real-device performance and safety. 🔗 Learn more about the agents: - Qwen Intelligence official website: qwenintelligence.com - Mobile Planner Agent: github.com/Tongyi-MAI/Qwen-P… - Mobile-Use Agent: tongyi-mai.github.io/Qwen-UI… - Mobile Creative Agent: arxiv.org/abs/2608.16887 🔗 Explore our open benchmark suite: - MobilePA-Bench: tongyi-mai.github.io/MobileP… - MobileWorld (GitHub): github.com/Tongyi-MAI/Mobile… - Leaderboard: tongyi-mai.github.io/MobileW…
130
297
2,886
213,300
shaun retweeted
⚡ Meet Qwen-Audio-3.1! ASR, TTS & Realtime are fully upgraded, joined by two new models: TTS-Next for audio creation and ASR-Next for audio understanding. Five models, one complete audio stack: understanding, generation, interaction & creation. Plus big price cuts across the lineup: TTS ~70% off, Realtime ~85% off, and ASR up to 95% off. Highlights: 🥳 - ASR: stronger multilingual & dialect recognition, plus native polishing that auto-removes fillers & repetitions for cleaner, more logical transcripts. - ASR-Next: supports multi-speaker ASR with speaker labels, timestamps & aligned transcripts, and understands emotions, ambient & machine sounds for sound captioning, event localization, audio QA & reasoning. - TTS: multilingual & dialect synthesis with natural cross-lingual voice transfer; control emotion, speed & style via simple instructions. - TTS-Next: unified LM + diffusion framework generating voice, sound effects & background audio in one pass for audiobooks, podcasts, games & ads. - Realtime: speak & listen at once with anytime interruption, just like a real call; it even slows down and responds empathetically when it senses a low mood. Unlock the full potential of Qwen-Audio-3.1! 👇 - Blog: fun-resource-shanghai.oss-cn… - Qwen-Audio-3.1-ASR: qwencloud.com/models/qwen-au… - Qwen-Audio-3.1-Realtime: qwencloud.com/models/qwen-au… - More APIs: coming soon @qwen_cloud
124
357
3,988
239,803
a week ago or w/e when I mentioned that alibaba were targetting 10T params, I didn't exactly expect it would happen so soon (or even be announced, tbh). if we truly see even 5T with qwen 4.5, that'd be something else, especially running locally 😋
2
10
905
and btw, the qwen4 release is still aiming for the end of this month or the beginning of october, depending on how training goes
2
1
60
2,464
GLM-5.5 isn't too far away now either. we're getting spoiled for open weight releases. both GLM & Qwen4 are being trained on a mix of Huawei Ascend + NVIDIA clusters Alibaba have "significantly overhauled" post-training and RL pipelines to help improve scalability on same base
2
25
1,897
Should see some llama.cpp PRs for full Qwen4 support within the next few days!
3
1
46
2,078
ask your agent to monitor the llama.cpp github for it if you're bored lol
2
370
shaun retweeted
Hong Kong Airport
12
18
428
15,563
Xiaomi have distilled MiMo V2.6 into Qwen3.5-9B and have introduced an MTP head into it along with support for preserve_thinking. might be worth checking out if you've been looking for a good, smaller dense model. huggingface.co/XiaomiMiMo/Mi… ty @XiaomiTech_
3
1
37
2,615
for those curious, my own testing (pure vibes, haven't done any actual evals) so far shows it to a pretty capable agentic model. does great with tool work. coding wise, I haven't had it dig into any real codebase yet or anything. just merely web dev stuff, and it's alright there
2
4
372
shaun retweeted
Please welcome GPT-6 Sol and GPT-6 Luna to the GPT-6 universe. GPT-6 Sol and Luna build on the advances behind GPT-6 Astra, bringing much of its strengths into faster and more affordable models to support work at scale. We’ve also made caching and inference more efficient, and we’re passing the savings directly to you: 50% lower API prices for Sol and Luna compared with GPT‑5.6 promotional pricing.
2,205
5,271
53,277
9,633,027
shaun retweeted
Qwen-Image-2.1 by @Alibaba_Qwen just landed as the #1 open source model in the Image Edit Arena and Text-to-Image Arena! With 1367 pts in the Image Edit Arena, Qwen-Image-2.1 took the #1 spot among open. It landed #16 overall, just 3 pts from GPT-Image-1.5-high-fidelity at #15. See the leaderboard for the Text-to-Image arena below. Congrats to the @Alibaba_Qwen team on this contribution to the open source ecosystem!
Meet Qwen-Image-2.1, the most balanced and cost-effective image generation model in the Qwen-Image series! Now open weights! 🎨 A unified model for both generation and editing, delivering top-tier quality in a lightweight package. Highlights: 👀 - Compact & exceptionally fast: A lightweight 7B architecture that outperforms most closed-source models, with drastically accelerated inference for multi-image inputs. - Native transparency: Natively generates and edits RGBA layers, enabling seamless compositing and text editing within transparent images. - Versatile, high-fidelity editing: Supports up to 10 reference images and precise local control while preserving strict fidelity for portraits and products. - Broad coverage & stunning aesthetics: Excels at panoramas, infographics, and virtual try-ons, delivering realistic textures and elegant typography. Start to create your next masterpiece with Qwen-Image-2.1! 🖼️ - Blog: qwen.ai/blog?id=qwen-image-2… - GitHub: github.com/QwenLM/Qwen-Image… - Model Scope: modelscope.cn/models/Qwen/Qw… - Hugging Face: huggingface.co/Qwen/Qwen-Ima…
32
31
584
220,320
While I don't want to say too much, hold out hope for a Qwen4-35B-A3B. It may not have been announced, but I know one is being tested.
16
10
212
28,821
I've been told that the reason there wasn't a 35B-A3B 3.8 is because it was difficult to scale the performance relative to the 27B dense. The bottleneck is the active params, but they want to keep 3.3B. So right now they're experimenting.
1
39
2,084
Realize this is slightly unclear, so to make it a little better: Qwen 3.8 35B-A3B compared to the 27B had a much larger perf delta than 3.5/3.6. So it never saw the light of day. Hopefully, though, the new architecture allows for those 3.3B params to hold up how they want.
2
36
1,947
Is there any good podcast out there about AI where the hosts actually know how running models works, etc? I tried Nerd Snipe, but it's basically just toddlers coming up with authoritative sounding stuff on the spot
12
16
1,427
Announced today :)
Qwen4 全家桶即将到来! Qwen4-Max Qwen4-Flash Qwen4-Plus 还有炙手可热的 Qwen4-27B!!! 现场说未来Qwen 的模型规格会达到 5-10T! Qwen 加油!干死达里奥!!
3
1
36
2,427
told ya ;)
1
6
506
shaun retweeted
🚀 Meet Qwen3.8-Omni-Flash, Qwen's first omni-modal model built around agentic capabilities! Native audio-video understanding, reasoning, and tool use come together in one model: understand the content, plan the task, execute with tools, and deliver the result. Highlights: 🥳 - Audio-video intelligence that gets things done: jointly reason over what's seen and heard, and orchestrate tools across long workflows to auto-edit vlogs, translate short videos, and turn movies into recaps. - A major leap: approaching Gemini 3.8 Flash in audio-video capabilities; +19.5 points on average in agent performance across WildClawBench-MM & UniClawBench. - 1M-token context with agentic perception: actively explore long videos and locate key moments with higher accuracy, using 51.8% fewer tokens than static understanding on OmniVideoBench. Video input costs are reduced by about 89% compared with Qwen3.5-Omni-Plus, making long-form audio-video understanding and agentic workflows more affordable than ever. To help you build apps around Omni, we're also open-sourcing Qwen-MM-Plugins and Qwen-Live Harness! 🛠️ We can't wait to see what you build with Qwen3.8-Omni-Flash! 👀 - Blog: qwen.ai/blog?id=qwen3.8-omni… - Qwencloud: qwencloud.com/models/qwen3.8… - Qwen Studio: chat.qwen.ai/ - API: alibabacloud.com/help/en/mod… - Qwen-MM-Plugins: github.com/QwenLM/Qwen-MM-Pl… - Qwen-Live Harness: coming soon github.com/QwenLM/Qwen-Live-…
177
362
3,609
282,146