🚀 Meet Qwen3.8-Omni-Flash, Qwen's first omni-modal model built around agentic capabilities!
Native audio-video understanding, reasoning, and tool use come together in one model: understand the content, plan the task, execute with tools, and deliver the result.
Highlights: 🥳
- Audio-video intelligence that gets things done: jointly reason over what's seen and heard, and orchestrate tools across long workflows to auto-edit vlogs, translate short videos, and turn movies into recaps.
- A major leap: approaching Gemini 3.8 Flash in audio-video capabilities; +19.5 points on average in agent performance across WildClawBench-MM & UniClawBench.
- 1M-token context with agentic perception: actively explore long videos and locate key moments with higher accuracy, using 51.8% fewer tokens than static understanding on OmniVideoBench.
Video input costs are reduced by about 89% compared with Qwen3.5-Omni-Plus, making long-form audio-video understanding and agentic workflows more affordable than ever.
To help you build apps around Omni, we're also open-sourcing Qwen-MM-Plugins and Qwen-Live Harness! 🛠️
We can't wait to see what you build with Qwen3.8-Omni-Flash! đź‘€
- Blog:
qwen.ai/blog?id=qwen3.8-omni…
- Qwencloud:
qwencloud.com/models/qwen3.8…
- Qwen Studio:
chat.qwen.ai/
- API:
alibabacloud.com/help/en/mod…
- Qwen-MM-Plugins:
github.com/QwenLM/Qwen-MM-Pl…
- Qwen-Live Harness: coming soon
github.com/QwenLM/Qwen-Live-…