LLM Releases
← Catalog

Gemini 3.8 Flash

Available
Google DeepMindProprietary

Google DeepMind's fast, cost-efficient Flash model released Sep 2, 2026 — its fourth Flash model in under four months and the successor to Gemini 3.7 Flash. At high reasoning it scores 59 on the Artificial Analysis Intelligence Index (up 3 points from Gemini 3.7 Flash and level with sub-maximum efforts of GPT-5.6 Sol and Grok 4.6), 57 at medium (matching GPT-5.6 Terra and Muse Spark 1.2), and 52 at low (matching Gemini 3.6 Flash at ~30% lower cost per task). The improvement is driven mainly by agentic evaluations — t^3-Banking tool use (+12 points to 45%), Terminal-Bench v2.1 coding, and GDPval-AA v2 real-world tasks. It keeps a 1M-token context window and multimodal input (text, image, video, speech) with text output. Pricing matches Gemini 3.7 Flash's current discounted rate of $0.75/$3.75 per Mtok input/output through the end of 2026 ($1.50/$7.50 at standard pricing), with cached input keeping a 90% discount; a ~30% rise in average output tokens per task (to ~48k) lifts cost per task to ~$0.58 at high reasoning despite unchanged per-token pricing. Available in the Gemini app for AI Pro and Ultra subscribers, AI Mode, and Gemini in Google Sheets, and for developers via Google Antigravity, AI Studio, and the Gemini API. Benchmark figures are vendor/third-party-reported.

Specifications

License
Proprietary
Weights
Not released
Architecture
unknown
Parameters
Undisclosed
Context window
1.0M tokens
Max output
66K tokens
Knowledge cutoff
Price (in / out, $/M)
$0.75 / $3.75
Modalities
TextVisionAudioVideoCode

Benchmarks

No benchmark scores recorded yet. Spotted some? Submit a correction.

Vendor-reported figures are claims until independently verified. See methodology.