LLM Releases
← Catalog

Qwen3.8-Flash-Next

Preview
Alibaba (Qwen)Open weights

An open-weight, experimental preview of the architecture that will underpin Qwen4, released Aug 26 2026 (Qwen/Qwen3.8-Flash-Next). A sparse MoE with ~6B active parameters (headline 125B-with-6B-activated; ~180B stored once a 51B n-gram embedding table and 4B multi-token-prediction module are counted), 512 experts (10 routed + 1 shared), and a hybrid Gated-DeltaNet + Qwen Sparse Attention design. Native 262,144-token context, extensible to 1M via YaRN. Accepts text, image, and video in and returns text out. Distinct from the managed Qwen Cloud 'Qwen3.8-Flash' API (which defaults to 1M context and bundled tools); this Next build is catalog/self-host only with no hosted list price at launch, served via Transformers, vLLM, SGLang, and TokenSpeed. Weights under the Qwen Community License 1.0. Self-reported vs DeepSeek-V4-Flash-0731: DeepSWE 58.7 vs 54.4, SWE-bench Pro 62.5 vs 56.0, LiveCodeBench v6 91.9, GPQA Diamond 91.7, though NL2Repo 48.1 vs 54.2 is a regression; vision self-reports include AndroidWorld 84.5 and RealWorldQA 88.5 — all vendor numbers, unverified at launch. Thinking on by default.

Specifications

License
Open weights · Qwen Community License 1.0
Weights
Downloadable
Architecture
Mixture-of-Experts
Parameters
125B · 6B active
Context window
262K tokens
Max output
Knowledge cutoff
Price (in / out, $/M)
Modalities
TextVisionCode

Benchmarks

No benchmark scores recorded yet. Spotted some? Submit a correction.

Vendor-reported figures are claims until independently verified. See methodology.