Gemini 3.5 Flash-Lite
AvailableGoogle DeepMind's fastest and most cost-effective 3.5-class model, released July 21, 2026 for low-latency and high-throughput agentic workloads like agentic search and document processing. Runs at ~350 output tokens/s (Artificial Analysis) with configurable thinking levels and built-in computer use, priced at $0.30 / $2.50 per 1M input/output tokens. Multimodal over a 1M-token context and a large step up on 3.1 Flash-Lite: Terminal-Bench 2.1 54% (vs 31%), GDM-MRCR v2 72.2% (vs 60.1%), GDPval-AA v2 1140 (vs 642); on several agentic and coding evals it even surpasses 3 Flash (SWE-Bench Pro 54.2% vs 49.6%, OSWorld-Verified 74.0% vs 65.1%). Available in the Gemini API (AI Studio, Android Studio), Gemini Enterprise, the Gemini app, and rolling out in Google Search.
Specifications
- License
- Proprietary
- Weights
- Not released
- Architecture
- unknown
- Parameters
- Undisclosed
- Context window
- 1M tokens
- Max output
- 66K tokens
- Knowledge cutoff
- β
- Price (in / out, $/M)
- $0.3 / $2.5
- Modalities
- TextVisionAudioVideoCode
Benchmarks
No benchmark scores recorded yet. Spotted some? Submit a correction.
Vendor-reported figures are claims until independently verified. See methodology.