Qwen3.8-Omni-Flash
AvailableQwen3.8-Omni-Flash is Alibaba Qwen's first omni-modal model built around agentic capabilities, released 2026-09-18. It natively accepts text, image, audio and video input and returns text-only output (no speech synthesis). It offers a 1M-token context window (max output 131,072 tokens; a "thinking" mode adds up to ~262K reasoning tokens) and is served as a proprietary hosted model on Alibaba Cloud (Model Studio / DashScope, with OpenAI-compatible and Chat Completions / Responses APIs), priced at $0.15 / 1M input and $0.47 / 1M output (cache read $0.016). It targets agentic audio-video understanding — proactively locating key segments in long videos, meeting summaries, video research, vlog auto-editing, short-video translation and movie recaps — and supports function calling, web search, structured outputs, prompt/context caching and batch calls, plus two-/four-channel spatial audio input. Alibaba reports an average improvement of more than 26% across 30 evaluations versus Qwen3.5-Omni-Plus, about 51.8% fewer tokens on OmniVideoBench in agent-perception mode, and roughly 89% lower video input cost. Companion open-source Qwen-MM-Plugins shipped alongside it (a Qwen-Live Harness was announced as coming soon). Benchmark and efficiency figures are vendor / self-reported.
Specifications
- License
- Proprietary · Proprietary (Alibaba Cloud)
- Weights
- Not released
- Architecture
- unknown
- Parameters
- Undisclosed
- Context window
- 1M tokens
- Max output
- 131K tokens
- Knowledge cutoff
- —
- Price (in / out, $/M)
- $0.15 / $0.47
- Modalities
- TextVisionVideoAudioCode
Benchmarks
No benchmark scores recorded yet. Spotted some? Submit a correction.
Vendor-reported figures are claims until independently verified. See methodology.