LLM Releases
← Catalog

Qwen3.8-Omni-Flash

Available
Alibaba (Qwen)Proprietary

Qwen3.8-Omni-Flash is Alibaba Qwen's first omni-modal model built around agentic capabilities, released 2026-09-18. It natively accepts text, image, audio and video input and returns text-only output (no speech synthesis). It offers a 1M-token context window (max output 131,072 tokens; a "thinking" mode adds up to ~262K reasoning tokens) and is served as a proprietary hosted model on Alibaba Cloud (Model Studio / DashScope, with OpenAI-compatible and Chat Completions / Responses APIs), priced at $0.15 / 1M input and $0.47 / 1M output (cache read $0.016). It targets agentic audio-video understanding — proactively locating key segments in long videos, meeting summaries, video research, vlog auto-editing, short-video translation and movie recaps — and supports function calling, web search, structured outputs, prompt/context caching and batch calls, plus two-/four-channel spatial audio input. Alibaba reports an average improvement of more than 26% across 30 evaluations versus Qwen3.5-Omni-Plus, about 51.8% fewer tokens on OmniVideoBench in agent-perception mode, and roughly 89% lower video input cost. Companion open-source Qwen-MM-Plugins shipped alongside it (a Qwen-Live Harness was announced as coming soon). Benchmark and efficiency figures are vendor / self-reported.

Specifications

License
Proprietary · Proprietary (Alibaba Cloud)
Weights
Not released
Architecture
unknown
Parameters
Undisclosed
Context window
1M tokens
Max output
131K tokens
Knowledge cutoff
Price (in / out, $/M)
$0.15 / $0.47
Modalities
TextVisionVideoAudioCode

Benchmarks

No benchmark scores recorded yet. Spotted some? Submit a correction.

Vendor-reported figures are claims until independently verified. See methodology.