Lab release history
Last updated Sep 14, 2026
DeepSeek model releases
Chinese lab shipping permissively licensed frontier-class models. This page collects the lab's model releases, lifecycle events, source links, and model metadata in one crawlable record.
23 models
DeepSeek-V4.1-Flash
AvailableDeepSeek's efficient, very-low-cost flagship released September 10 2026, retiring V4-Flash and taking over V4-Pro API traffic on September 14 (DeepSeek reports it beats V4-Pro on performance, cost, speed, and total time). A 552B-total-parameter multimodal sparse Mixture-of-Experts model built on a new Causal Encoder-Decoder architecture that activates only ~8B parameters per token on input and ~16B on output for cheaper long-context prefill, with a 1M-token context, up to 384K output tokens, and native image understanding (vision in, text out). The other headline change over July's V4-Flash is memory: FP4 quantization plus "pure CSA2" cross-layer attention reuse compress the KV cache to roughly 890 bytes per token — about an 8x reduction, cutting HBM to a quarter and SSD to an eighth for an equivalent conversation state, which is what makes million-token agentic runs practical on a single node. Open-weight under the MIT license, downloadable and self-hostable, and served across many inference providers. API pricing is $0.30/$1.20 per Mtok input/output at peak (01:00-04:00 and 06:00-10:00 UTC weekdays) and half that off-peak ($0.15/$0.60), with cache hits around $0.006/$0.003 per Mtok. DeepSeek reports it narrowly edges Claude Opus 5 and GPT-5.6 Sol on DeepSWE, but there is no independent benchmark table at launch and vendor performance claims are unverified.
DeepSeek-V4-Flash-Vision-Exp
RetiredDeepSeek's first multimodal V4 model — an experimental vision-understanding checkpoint that went live on the DeepSeek API (model='deepseek-v4-flash-vision-exp') on Aug 21, 2026. It extends DeepSeek-V4-Flash with image understanding while keeping its full text capabilities (agents, reasoning, coding, and world knowledge), matching V4-Flash on text benchmarks. DeepSeek reports a major jump on multimodal agent benchmarks over V4-Flash, bringing multimodal-agent performance close to Opus-4.8 — a vendor-reported result, unverified by an independent harness at launch. Keeps V4-Flash's 284B-total / 13B-active sparse MoE architecture and 1M-token context, with up to ~393K output tokens; accepts text plus up to 600 images per request (8,192px per side, 64 MiB payload) and returns text only, with images billed at up to 384 tokens each. API pricing held at V4-Flash rates: $0.22 / $0.66 per 1M input/output tokens, with a $0.007 per 1M cached-input rate. API-only and experimental at launch — weights were not published, so treated as proprietary. RETIRED 2026-09-10: superseded by DeepSeek-V4.1-Flash, whose native multimodal support absorbs this experiment. For compatibility the `deepseek-v4-flash-vision-exp` API id temporarily routes to V4.1-Flash.
DeepSeek-V4-Pro-0813
DeprecatedThe dated GA build behind DeepSeek's 'deepseek-v4-pro' API id, pinned on the official pricing table with an OpenRouter listing dated Aug 12 2026 — the Pro-tier counterpart to the way V4-Flash graduated as DeepSeek-V4-Flash-0731. It is a quiet version pin (no separate launch post or benchmark card), keeping the V4-Pro architecture and the 1M-token context / ~384K max-output window. List pricing is cache-heavy: $0.435 input cache-miss / $0.003625 cache-hit / $0.87 output per Mtok, with concurrency 500 (vs Flash's 2500); DeepSeek warns a significant, undated API price increase is coming. Thinking is on by default at effort 'high' (requested medium/xhigh both collapse to high). Architecture figures (1.6T total / 49B active MoE, hybrid long-context attention) are carried over from the April 2026 V4-Pro preview and are not independently reconfirmed for the 0813 build; Hugging Face still hosts only the April preview weights (MIT), with no confirmed separate 0813 open-weight repo, so this row is recorded as proprietary/API-only. DEPRECATED 2026-09-14: the `deepseek-v4-pro` API id this dated build served now routes to DeepSeek-V4.1-Flash at Flash rates, pending the V4.1-Pro launch.
DeepSeek-V4-Flash-0731
RetiredThe production release of DeepSeek's V4-Flash tier — the April V4-Flash preview retrained on a substantially improved post-training pipeline targeting coding, agents, reasoning, and tool use, with no change to the base architecture. Retains 284B total / 13B active parameters (MoE) and the 1M-token context window. DeepSeek reports the 0731 build scoring higher than its own larger V4-Pro-Preview on all nine agent and coding benchmarks it published — a vendor-reported result, with independent replication still limited at launch. Weights released on Hugging Face under the MIT license; API pricing held at $0.14 / $0.28 per Mtok. The upgrade is silent for existing callers: same endpoint, same key, same deepseek-v4-flash model name, zero migration cost. RETIRED 2026-09-10: the `deepseek-v4-flash` API id this dated build served was retired in favour of DeepSeek-V4.1-Flash and now routes there.
DeepSeek V4-Flash
RetiredEfficient V4 companion model with 284B total / 13B active parameters and the same one-million-token context window. RETIRED 2026-09-10: superseded by DeepSeek-V4.1-Flash. For compatibility the `deepseek-v4-flash` API id temporarily routes to V4.1-Flash.
DeepSeek V4-Pro
DeprecatedPreview-series sparse MoE flagship with a one-million-token context window and 1.6T total / 49B active parameters. DEPRECATED 2026-09-14: from 04:00 UTC all `deepseek-v4-pro` requests route to DeepSeek-V4.1-Flash, billed at V4.1-Flash rates, and will continue to until V4.1-Pro launches. The id still answers, so this is a redirect rather than a retirement.
DeepSeek-V3.2
AvailableReasoning-first agent model that adds DeepSeek Sparse Attention and thinking directly inside tool-use workflows.
DeepSeek-V3.2-Speciale
AvailableHigh-compute reasoning variant of V3.2, positioned for olympiad-level math, programming, and other deep reasoning tasks.
DeepSeek-V3.2-Exp
PreviewExperimental checkpoint that introduced DeepSeek Sparse Attention as an efficiency bridge between V3.1-Terminus and V3.2.
DeepSeek-V3.1-Terminus
AvailableStability update to V3.1 focused on language consistency, code-agent reliability, and search-agent behavior.
DeepSeek-V3.1
AvailableHybrid thinking/non-thinking release that upgraded tool calling, long-context training, and agent task performance.
DeepSeek R2
RumoredRumored successor to DeepSeek R1. Reports say development and launch timing were affected by hardware constraints around Huawei Ascend training and Nvidia availability; final specs, license, and release date remain unconfirmed.
DeepSeek-R1-0528
AvailableMajor R1 reasoning update with stronger math, programming, general logic, function calling, and reduced hallucinations.
DeepSeek-V3-0324
AvailablePost-R1 V3 update with improved reasoning, front-end coding, Chinese writing, search, and function calling.
DeepSeek-R1
AvailableBreakout open reasoning model trained with large-scale reinforcement learning and released with weights under MIT.
DeepSeek-V3
AvailableThe 671B/37B-active MoE release that made DeepSeek a central open-model lab before the R1 breakthrough.
DeepSeek-R1-Lite-Preview
RetiredReasoning-preview model exposed in DeepSeek Chat ahead of the open DeepSeek-R1 release.
DeepSeek-V2.5
AvailableUnified DeepSeek V2 generation combining general-chat and coding strengths before the V3 series.
DeepSeek-Coder-V2
AvailableOpen code-focused MoE built from DeepSeek-V2, expanding programming-language coverage and coding benchmark performance.
DeepSeek-V2
AvailableDeepSeek's first major MoE general model with Multi-head Latent Attention and low-cost API positioning.
DeepSeekMoE 16B
AvailableEarly DeepSeek sparse MoE research model that foreshadowed the later V2/V3 architecture direction.
DeepSeek LLM 67B
AvailableFirst general DeepSeek language model family, with 7B and 67B base/chat checkpoints.
DeepSeek Coder 33B
AvailableDeepSeek's first public code-model family, released before the general DeepSeek LLM line.