LLM Releases

Lab release history

Last updated Jul 21, 2026

Google DeepMind model releases

Google's combined AI research organization; builds the Gemini family. This page collects the lab's model releases, lifecycle events, source links, and model metadata in one crawlable record.

24
Models
1
Labs
9
Open
4
Recent

24 models

Gemini 3.5 Flash Cyber

Preview
Google DeepMindProprietary

A specialized, highly efficient cybersecurity model built on Gemini 3.5 Flash and fine-tuned to find and fix software vulnerabilities at a lower price per token than larger models. Deployed inside Google's CodeMender agent, where multiple 3.5 Flash Cyber agents collaborate to produce a single combined report, reaching competitive frontier performance on the CyberGym benchmark. Given its dual-use nature, it is not generally available: access is limited to governments and trusted partners via CodeMender as part of a limited-access pilot program.

Undisc.1M ctxJul 21, 2026

Gemini 3.5 Flash-Lite

Available
Google DeepMindProprietary

Google DeepMind's fastest and most cost-effective 3.5-class model, released July 21, 2026 for low-latency and high-throughput agentic workloads like agentic search and document processing. Runs at ~350 output tokens/s (Artificial Analysis) with configurable thinking levels and built-in computer use, priced at $0.30 / $2.50 per 1M input/output tokens. Multimodal over a 1M-token context and a large step up on 3.1 Flash-Lite: Terminal-Bench 2.1 54% (vs 31%), GDM-MRCR v2 72.2% (vs 60.1%), GDPval-AA v2 1140 (vs 642); on several agentic and coding evals it even surpasses 3 Flash (SWE-Bench Pro 54.2% vs 49.6%, OSWorld-Verified 74.0% vs 65.1%). Available in the Gemini API (AI Studio, Android Studio), Gemini Enterprise, the Gemini app, and rolling out in Google Search.

Undisc.1M ctxJul 21, 2026

Gemini 3.6 Flash

Available
Google DeepMindProprietary

Google DeepMind's July 2026 workhorse Flash model, built for scaling agentic workflows. Multimodal over a 1M-token context, it improves on Gemini 3.5 Flash in coding, knowledge work, and computer use while cutting output-token usage ~17% (up to 65% on some benchmarks like DeepSWE) and taking fewer reasoning steps and tool calls. Ships at a lower price than 3.5 Flash ($1.50 / $7.50 per 1M input/output tokens). Google-reported gains: DeepSWE 49% (vs 37%), MLE-Bench 63.9% (vs 49.7%), OSWorld-Verified 83.0% (vs 78.4%), GDPval-AA v2 1421 (vs 1349); knowledge cutoff advances to March 2026. Computer use is a built-in client-side tool. Available in the Gemini API (AI Studio, Android Studio, Antigravity), Gemini Enterprise, and the Gemini app.

Undisc.1M ctxJul 21, 2026

DiffusionGemma 26B-A4B

Available
Google DeepMindOpen source

An open-weight text-diffusion model built on the Gemma 4 26B-A4B MoE backbone (25.2B total / 3.8B active). Denoises text in parallel 256-token blocks for up to ~4x faster generation (1,000+ tok/s on an H100), with a 256K context and text, image, and video input. Apache-2.0.

MoE25.2B256K ctxJun 10, 2026

Gemma 4 12B

Available
Google DeepMindOpen source

A dense 12B member of the Gemma 4 family with a unified, encoder-free multimodal architecture: vision and audio are projected straight into the LLM backbone. First medium-size Gemma to natively ingest audio; runs on a 16GB laptop. 256K context, Apache-2.0.

Dense12B256K ctxJun 3, 2026

Gemini 3.5 Pro

Preview
Google DeepMindFrontierProprietary

Announced at Google I/O 2026; emphasizes deep multimodal reasoning over a 2M-token context. Recent reporting says the broad launch slipped from June toward July while testers continue using it in Google Antigravity and LMArena.

MoEUndisc.2M ctxMay 19, 2026

Gemini 3.5 Flash

Available
Google DeepMindProprietary

Google's fast, cost-efficient Gemini 3.5 tier, unveiled at I/O 2026. Multimodal over a 1M-token context and tuned for agentic and coding workflows; Google says it beats Gemini 3.1 Pro on coding and tool-use while running ~4x faster.

Undisc.1M ctxMay 19, 2026

Gemma 4 31B

Available
Google DeepMindOpen source

Google DeepMind's Gemma 4 advanced-reasoning open model for personal computers, part of the April 2026 Gemma 4 family.

Dense31B ctxApr 2, 2026

Gemini 3.1 Pro

Available
Google DeepMindFrontierProprietary

Generally available multimodal flagship with native tool use and a 2M-token context.

MoEUndisc.2M ctxFeb 19, 2026

Gemma 3 27B

Available
Google DeepMindOpen weights

Google's open multimodal model: 128k context, 140+ languages, runs on a single GPU.

Dense27B128K ctxSep 4, 2025

Gemini 2.5 Deep Think

Available
Google DeepMindFrontierProprietary

Google's enhanced Gemini 2.5 reasoning mode for harder math, science, coding, and multimodal analysis. Previewed at Google I/O 2025 and later made available to Gemini app subscribers, Deep Think uses more deliberative reasoning for complex prompts.

Undisc.1M ctxAug 1, 2025

Gemini 2.5 Flash-Lite

Available
Google DeepMindProprietary

Google's lowest-latency, lowest-cost Gemini 2.5 tier, designed for summarization, classification, extraction, routing, and other high-volume production tasks. Proprietary API model with a 1M-token context and multimodal support.

Undisc.1M ctxJul 22, 2025

Gemini 2.5 Flash

Available
Google DeepMindProprietary

Google's faster, lower-cost Gemini 2.5 model for high-throughput multimodal and agentic workloads. It brought Gemini 2.5's reasoning improvements to a production Flash tier with a 1M-token context and broad text, image, audio, video, and coding support.

Undisc.1M ctxJun 17, 2025

Gemma 3n E4B

Available
Google DeepMindOpen weights

Google's mobile-first Gemma 3n model variant, built with a MatFormer-style architecture for efficient on-device multimodal inference. The E4B variant has roughly 4B effective parameters, supports text, vision, audio, and video-oriented use cases, and is released under Gemma terms.

Hybrid8B32K ctxMay 20, 2025

Gemini 2.5 Pro

Deprecated
Google DeepMindProprietary

Reasoning-focused Gemini 2.5 model that made thinking a core part of Google's flagship model line.

Undisc.1M ctxMar 25, 2025

Gemini 2.0 Flash

Deprecated
Google DeepMindProprietary

First Gemini 2.0 release, built for native multimodal input/output, tool use, and agentic product integrations.

Undisc.1M ctxDec 11, 2024

Gemma 2 27B

Available
Google DeepMindOpen weights

Second-generation Gemma model, improving open-weight quality and efficiency at 9B and 27B sizes.

Dense27B8K ctxJun 27, 2024

CodeGemma 7B

Available
Google DeepMindOpen weights

Open code-specialized Gemma model for local code completion, generation, and instruction-following.

Dense7B8K ctxApr 9, 2024

Gemma 7B

Available
Google DeepMindOpen weights

First Gemma open-weight text model family, derived from the same research lineage as Gemini.

Dense7B8K ctxFeb 21, 2024

Gemini 1.5 Pro

Deprecated
Google DeepMindProprietary

Gemini generation that introduced production-scale long context, eventually expanding to a two-million-token window.

MoEUndisc.2M ctxFeb 15, 2024

Gemini 1.0 Ultra

Deprecated
Google DeepMindProprietary

Google's first natively multimodal Gemini flagship, since superseded by the 1.5/2/3 lines.

Undisc.32K ctxDec 6, 2023

PaLM 2

Retired
Google DeepMindProprietary

Google's improved multilingual, reasoning, and coding foundation model family introduced at I/O 2023.

DenseUndisc. ctxMay 10, 2023

PaLM

Retired
Google DeepMindProprietary

Google's 540B Pathways model; the API was later deprecated in favor of Gemini.

Dense540B ctxApr 4, 2022

BERT

Available
Google DeepMindOpen source

The bidirectional encoder that reshaped NLP and seeded the transformer era.

Dense0.34B512 ctxOct 11, 2018