DeepSeek-V4.1-Flash
AvailableDeepSeek's efficient, very-low-cost flagship released September 10 2026, retiring V4-Flash and taking over V4-Pro API traffic on September 14 (DeepSeek reports it beats V4-Pro on performance, cost, speed, and total time). A 552B-total-parameter multimodal sparse Mixture-of-Experts model built on a new Causal Encoder-Decoder architecture that activates only ~8B parameters per token on input and ~16B on output for cheaper long-context prefill, with a 1M-token context, up to 384K output tokens, and native image understanding (vision in, text out). The other headline change over July's V4-Flash is memory: FP4 quantization plus "pure CSA2" cross-layer attention reuse compress the KV cache to roughly 890 bytes per token — about an 8x reduction, cutting HBM to a quarter and SSD to an eighth for an equivalent conversation state, which is what makes million-token agentic runs practical on a single node. Open-weight under the MIT license, downloadable and self-hostable, and served across many inference providers. API pricing is $0.30/$1.20 per Mtok input/output at peak (01:00-04:00 and 06:00-10:00 UTC weekdays) and half that off-peak ($0.15/$0.60), with cache hits around $0.006/$0.003 per Mtok. DeepSeek reports it narrowly edges Claude Opus 5 and GPT-5.6 Sol on DeepSWE, but there is no independent benchmark table at launch and vendor performance claims are unverified.
Specifications
- License
- Open source · MIT
- Weights
- Downloadable
- Architecture
- Mixture-of-Experts
- Parameters
- 552B · 16B active
- Context window
- 1M tokens
- Max output
- 384K tokens
- Knowledge cutoff
- —
- Price (in / out, $/M)
- $0.3 / $1.2
- Modalities
- TextVisionCode
Benchmarks
No benchmark scores recorded yet. Spotted some? Submit a correction.
Vendor-reported figures are claims until independently verified. See methodology.