Qwen3.8-Flash-Next is here! An open-weight multimodal MoE built on a brand new architecture, with native 256K context extendable to 1M via YaRN. 🤖modelscope.ai/collections/Qw…
🏆 Leads every compared model on SWE-bench Pro (62.5 vs 53.4 for Claude-Opus-4.6 Max), SWE-bench Multilingual, CoWorkBench, JobBench, and Toolathlon Verified.
⚡ 125B params plus 51B n-gram embeddings, only 6B active per token. Stronger than Qwen3.7-Plus on coding and office tasks at roughly 1/9 the training cost. 📚 At 1M context, attention kernels run up to 7.6× faster on prefill and 4.9× on decode.
59
126
1,061
330,958
Just tried out @Alibaba_Qwen 3.8 Flash Next and its 120TPS in Apple M5 Max machine !
youtu.be/RTydhfQNBak
#executeautomation #qwen3flashnext
Aug 26, 2026 · 8:16 PM UTC
19


