LLM Releases
← Catalog

DeepSeek-V4-Flash-Vision-Exp

Preview
DeepSeekProprietary

DeepSeek's first multimodal V4 model — an experimental vision-understanding checkpoint that went live on the DeepSeek API (model='deepseek-v4-flash-vision-exp') on Aug 21, 2026. It extends DeepSeek-V4-Flash with image understanding while keeping its full text capabilities (agents, reasoning, coding, and world knowledge), matching V4-Flash on text benchmarks. DeepSeek reports a major jump on multimodal agent benchmarks over V4-Flash, bringing multimodal-agent performance close to Opus-4.8 — a vendor-reported result, unverified by an independent harness at launch. Keeps V4-Flash's 284B-total / 13B-active sparse MoE architecture and 1M-token context, with up to ~393K output tokens; accepts text plus up to 600 images per request (8,192px per side, 64 MiB payload) and returns text only, with images billed at up to 384 tokens each. API pricing held at V4-Flash rates: $0.22 / $0.66 per 1M input/output tokens, with a $0.007 per 1M cached-input rate. API-only and experimental at launch — weights were not published, so treated as proprietary.

Specifications

License
Proprietary
Weights
Not released
Architecture
Mixture-of-Experts
Parameters
284B · 13B active
Context window
1M tokens
Max output
393K tokens
Knowledge cutoff
Price (in / out, $/M)
$0.22 / $0.66
Modalities
TextVisionCode

Benchmarks

No benchmark scores recorded yet. Spotted some? Submit a correction.

Vendor-reported figures are claims until independently verified. See methodology.