Qwen3.8-Flash
AvailableThe productionized, managed Qwen Cloud API model announced Aug 26 2026 alongside the open-weight Qwen3.8-Flash-Next, with official QwenCloud pricing confirmed Aug 27-28. It runs the same Qwen4-preview architecture as Flash-Next — a 125B-total sparse Mixture-of-Experts that activates ~6B parameters per token (~180B stored once a 51B n-gram embedding table and a multi-token-prediction module are counted; roughly 95% sparsity), built on Gated-DeltaNet + Qwen Sparse Attention with a gated residual stream and Muon-trained large linear layers — but ships as the hosted service rather than the self-host weights. On QwenCloud it defaults to a 1M-token context with built-in tools and accepts text/image/video in, returning text out. List pricing is $0.15 input / $0.47 output / $0.016 cache-hit per Mtok (domestic China Y0.8/Y2.7/Y0.1), roughly a third of DeepSeek-V4-Flash and about one-thirteenth of the Qwen3.8-Max flagship. Also reachable through OpenCode Go's flat-rate subscription. Distinct catalog row from the open-weight Qwen3.8-Flash-Next (self-host, Qwen Community License 1.0); this managed API is recorded proprietary/API-only. Self-reported benchmarks carry over from Flash-Next (SWE-bench Pro 62.5, DeepSWE 58.7) — all vendor numbers, unverified by independent labs at launch. Thinking on by default.
Specifications
- License
- Proprietary · Proprietary managed API (underlying architecture open-weight as Qwen3.8-Flash-Next, Qwen Community License 1.0)
- Weights
- Not released
- Architecture
- Mixture-of-Experts
- Parameters
- 125B · 6B active
- Context window
- 1M tokens
- Max output
- 131K tokens
- Knowledge cutoff
- —
- Price (in / out, $/M)
- $0.15 / $0.47
- Modalities
- TextVisionCode
Benchmarks
No benchmark scores recorded yet. Spotted some? Submit a correction.
Vendor-reported figures are claims until independently verified. See methodology.