Gemini 3.8 Flash
AvailableGoogle DeepMind's fast, cost-efficient Flash model released Sep 2, 2026 — its fourth Flash model in under four months and the successor to Gemini 3.7 Flash. At high reasoning it scores 59 on the Artificial Analysis Intelligence Index (up 3 points from Gemini 3.7 Flash and level with sub-maximum efforts of GPT-5.6 Sol and Grok 4.6), 57 at medium (matching GPT-5.6 Terra and Muse Spark 1.2), and 52 at low (matching Gemini 3.6 Flash at ~30% lower cost per task). The improvement is driven mainly by agentic evaluations — t^3-Banking tool use (+12 points to 45%), Terminal-Bench v2.1 coding, and GDPval-AA v2 real-world tasks. It keeps a 1M-token context window and multimodal input (text, image, video, speech) with text output. Pricing matches Gemini 3.7 Flash's current discounted rate of $0.75/$3.75 per Mtok input/output through the end of 2026 ($1.50/$7.50 at standard pricing), with cached input keeping a 90% discount; a ~30% rise in average output tokens per task (to ~48k) lifts cost per task to ~$0.58 at high reasoning despite unchanged per-token pricing. Available in the Gemini app for AI Pro and Ultra subscribers, AI Mode, and Gemini in Google Sheets, and for developers via Google Antigravity, AI Studio, and the Gemini API. Benchmark figures are vendor/third-party-reported.
Specifications
- License
- Proprietary
- Weights
- Not released
- Architecture
- unknown
- Parameters
- Undisclosed
- Context window
- 1.0M tokens
- Max output
- 66K tokens
- Knowledge cutoff
- —
- Price (in / out, $/M)
- $0.75 / $3.75
- Modalities
- TextVisionAudioVideoCode
Benchmarks
No benchmark scores recorded yet. Spotted some? Submit a correction.
Vendor-reported figures are claims until independently verified. See methodology.