LLM Releases
← Catalog

Qwen3.8-Flash

Available
Alibaba (Qwen)Proprietary

The productionized, managed Qwen Cloud API model announced Aug 26 2026 alongside the open-weight Qwen3.8-Flash-Next, with official QwenCloud pricing confirmed Aug 27-28. It runs the same Qwen4-preview architecture as Flash-Next — a 125B-total sparse Mixture-of-Experts that activates ~6B parameters per token (~180B stored once a 51B n-gram embedding table and a multi-token-prediction module are counted; roughly 95% sparsity), built on Gated-DeltaNet + Qwen Sparse Attention with a gated residual stream and Muon-trained large linear layers — but ships as the hosted service rather than the self-host weights. On QwenCloud it defaults to a 1M-token context with built-in tools and accepts text/image/video in, returning text out. List pricing is $0.15 input / $0.47 output / $0.016 cache-hit per Mtok (domestic China Y0.8/Y2.7/Y0.1), roughly a third of DeepSeek-V4-Flash and about one-thirteenth of the Qwen3.8-Max flagship. Also reachable through OpenCode Go's flat-rate subscription. Distinct catalog row from the open-weight Qwen3.8-Flash-Next (self-host, Qwen Community License 1.0); this managed API is recorded proprietary/API-only. Self-reported benchmarks carry over from Flash-Next (SWE-bench Pro 62.5, DeepSWE 58.7) — all vendor numbers, unverified by independent labs at launch. Thinking on by default.

Specifications

License
Proprietary · Proprietary managed API (underlying architecture open-weight as Qwen3.8-Flash-Next, Qwen Community License 1.0)
Weights
Not released
Architecture
Mixture-of-Experts
Parameters
125B · 6B active
Context window
1M tokens
Max output
131K tokens
Knowledge cutoff
Price (in / out, $/M)
$0.15 / $0.47
Modalities
TextVisionCode

Benchmarks

No benchmark scores recorded yet. Spotted some? Submit a correction.

Vendor-reported figures are claims until independently verified. See methodology.