LLM Releases

Vision, audio, video

Last updated Jul 24, 2026

Multimodal LLM releases

Large language model releases with multimodal capabilities, including vision-language, audio, video, image-generation, and document-understanding models.

103
Models
21
Labs
39
Open
4
Recent

103 models

Claude Opus 5

Available
AnthropicFrontierProprietary

Anthropic's flagship Opus model, released July 24, 2026 — positioned as the go-to model for most knowledge work and automation, approaching the capability of Claude Fable 5 in many categories at roughly half the price. Built for demanding reasoning, autonomous coding, software development, and long-horizon agentic work. Introduces a five-level 'effort' dial exposed to developers on the Claude API and Platform, letting them trade compute and tokens for capability — at lower effort it preserves much of its performance while using fewer tokens and costing less to run. 1M-token context window (available at standard token pricing, not a separate long-context surcharge) with up to 128K output tokens; text, vision, and code. Standard API pricing $5/$25 per Mtok (the same as its predecessor Opus 4.8), plus a Fast mode at $10/$50. Anthropic describes it as its most aligned Opus model and the least susceptible to being tricked into misuse. Becomes the default model for Claude Max subscribers and is available across Anthropic's paid plans; scores 61 on the Artificial Analysis Intelligence Index. Closed weights; architecture, parameter count, and training compute undisclosed.

Undisc.1M ctxJul 24, 2026

Gemini 3.5 Flash-Lite

Available
Google DeepMindProprietary

Google DeepMind's fastest and most cost-effective 3.5-class model, released July 21, 2026 for low-latency and high-throughput agentic workloads like agentic search and document processing. Runs at ~350 output tokens/s (Artificial Analysis) with configurable thinking levels and built-in computer use, priced at $0.30 / $2.50 per 1M input/output tokens. Multimodal over a 1M-token context and a large step up on 3.1 Flash-Lite: Terminal-Bench 2.1 54% (vs 31%), GDM-MRCR v2 72.2% (vs 60.1%), GDPval-AA v2 1140 (vs 642); on several agentic and coding evals it even surpasses 3 Flash (SWE-Bench Pro 54.2% vs 49.6%, OSWorld-Verified 74.0% vs 65.1%). Available in the Gemini API (AI Studio, Android Studio), Gemini Enterprise, the Gemini app, and rolling out in Google Search.

Undisc.1M ctxJul 21, 2026

Gemini 3.6 Flash

Available
Google DeepMindProprietary

Google DeepMind's July 2026 workhorse Flash model, built for scaling agentic workflows. Multimodal over a 1M-token context, it improves on Gemini 3.5 Flash in coding, knowledge work, and computer use while cutting output-token usage ~17% (up to 65% on some benchmarks like DeepSWE) and taking fewer reasoning steps and tool calls. Ships at a lower price than 3.5 Flash ($1.50 / $7.50 per 1M input/output tokens). Google-reported gains: DeepSWE 49% (vs 37%), MLE-Bench 63.9% (vs 49.7%), OSWorld-Verified 83.0% (vs 78.4%), GDPval-AA v2 1421 (vs 1349); knowledge cutoff advances to March 2026. Computer use is a built-in client-side tool. Available in the Gemini API (AI Studio, Android Studio, Antigravity), Gemini Enterprise, and the Gemini app.

Undisc.1M ctxJul 21, 2026

Qwen3.8-Max-Preview

Preview
Alibaba (Qwen)FrontierProprietary

Alibaba's largest model to date and the flagship of the new Qwen3.8 line — a 2.4-trillion-parameter, fully multimodal model that Alibaba positions just behind Anthropic's Fable 5 on overall performance (vendor internal evals; no independent third-party benchmarks yet). Launched Jul 19, 2026 as Qwen3.8-Max-Preview, available to developers via Alibaba's Token Plan subscription and the Qoder / QoderWork coding platforms. Architecture is presumed sparse-MoE; active-parameter count, context window, and pricing are undisclosed. Breaking from the API-only pattern of earlier Max models, the Qwen team says it will release open weights "soon," though no timeline, license, or full specs have been announced.

MoE2.4T ctxJul 19, 2026

Kimi K3

Available
Moonshot AIFrontierOpen weights

Moonshot's flagship open-weight agentic model and the largest open model released to date: a 2.8T-parameter MoE (896 experts, 16 active per token) using Kimi Delta Attention and Attention Residuals, with native multimodal input and a 1M-token context. Launched via API on Jul 16, 2026 at $3/$15 per Mtok (cached input $0.30); full open weights published to Hugging Face on Jul 26, 2026 — a day ahead of the announced Jul 27 target — under a Modified MIT license, making it freely downloadable and self-hostable.

MoE2.8T1.0M ctxJul 16, 2026

Inkling

Available
Thinking Machines LabFrontierOpen source

Thinking Machines Lab's first model and the leading U.S. open-weights release: a natively multimodal Mixture-of-Experts with 975B total / 41B active parameters that reasons across text, image, and audio inputs and emits text. Pretrained on ~45T tokens; served with a 1M-token context from the Hugging Face weights (256K on the hosted Tinker API). Apache-2.0 licensed (BF16 + NVFP4 checkpoints on Hugging Face), built for developers fine-tuning on proprietary data — coding assistants, agents/tool use, chatbots, and RAG — with an explicit low-cost and censorship-resistance focus. Debuted at 41 on the Artificial Analysis Intelligence Index. Hosted pricing (256K) $3.74/$9.36 per Mtok reflects a limited-time 50% launch discount.

MoE975B1.0M ctxJul 15, 2026

GPT-5.6

Available
OpenAIFrontierProprietary

OpenAI's GPT-5.6 series umbrella row. Officially previewed June 26, 2026 as three durable capability tiers — Sol (flagship), Terra (balanced, for everyday work), and Luna (fast and affordable) — introduced with a new `max` reasoning effort for deeper reasoning and an `ultra` mode that leverages subagents to accelerate complex work. In the GPT-5.6 naming system the number marks the generation while Sol/Terra/Luna are tiers that can advance on their own cadence. Initially a limited preview via the API and Codex for a small group of vetted partners (after U.S. government review), with general availability across ChatGPT, Codex, and the API planned in the following weeks.

MoEUndisc.1.5M ctxJul 9, 2026

GPT-5.6 Terra

Available
OpenAIFrontierProprietary

Balanced, everyday-work tier of OpenAI's GPT-5.6 series, officially previewed June 26, 2026. OpenAI positions Terra as competitive with GPT-5.5 while being roughly 2x cheaper. Shares the series' new `max` reasoning effort and `ultra` subagent mode and OpenAI's GPT-5.6 safety stack. Begins as a limited preview via the API and Codex for vetted partners after U.S. government review, with broad availability planned in the following weeks. Priced at $2.50 / $15 per million input/output tokens.

Undisc.1.5M ctxJul 9, 2026

GPT-5.6 Sol

Available
OpenAIFrontierProprietary

Flagship tier of OpenAI's GPT-5.6 series, officially previewed June 26, 2026 — OpenAI's strongest model to date. Adds a new `max` reasoning effort for the deepest reasoning and an `ultra` mode that uses subagents to accelerate complex work. Sets a new state of the art on Terminal-Bench 2.1 (command-line, agentic coding) and shows broad gains in long-horizon biology (GeneBench v1) and cybersecurity (ExploitBench, ExploitGym), paired with OpenAI's most robust safety stack and a phased release. Begins as a limited preview via the API and Codex for a small group of vetted partners after U.S. government review, with general availability planned in the following weeks; also launching on Cerebras at up to 750 tokens/sec in July. Priced at $5 / $30 per million input/output tokens.

Undisc.1.5M ctxJul 9, 2026

GPT-5.6 Luna

Available
OpenAIProprietary

Fast, low-cost tier of OpenAI's GPT-5.6 series, officially previewed June 26, 2026 — the most affordable model in the family, bringing strong capability at OpenAI's lowest cost. Shares the series' capability-tier naming and GPT-5.6 safety stack. Begins as a limited preview via the API and Codex for vetted partners after U.S. government review, with broader availability planned in the following weeks. Priced at $1 / $6 per million input/output tokens.

Undisc.1.5M ctxJul 9, 2026

Muse Spark 1.1

Preview
Meta AIFrontierProprietary

Meta Superintelligence Labs' first paid model, released July 9, 2026 in US public preview on the Meta Model API. A natively multimodal reasoning model (text, image, video, PDF, and audio input; text output) with explicit chain-of-thought reasoning and a 1M-token context window that the model actively compacts. Positioned for agentic work and coding — tool use, multi-step workflow coordination, and long-horizon autonomous tasks — and pitched by Meta as roughly a quarter of the price of comparable Anthropic and OpenAI models at $1.25 in / $4.25 out per Mtok (with $20 in free credits per new API account). Marks the first time Meta has charged businesses for one of its models, a departure from the open-weight Llama strategy. Closed weights, undisclosed size. Vendor and third-party benchmarks place it around the Opus 4.8 / GPT-5.5 tier — strongest as an agent/workflow model and in tool-augmented reasoning, competitive but not dominant on coding and multimodal tasks.

Undisc.1.0M ctxJul 9, 2026

Grok 4.5

Available
xAIFrontierProprietary

SpaceXAI's "Opus-class" agentic flagship, released July 8, 2026 — the first Grok model trained jointly with the coding startup Cursor (Anysphere), on trillions of tokens of real Cursor usage data plus STEM tasks, research papers, and other knowledge work. Targets software engineering, agentic tasks, and knowledge work, with explicit strength in legal and finance use cases (SpaceXAI claims the top spot on the Harvey Legal Agent Benchmark). A Mixture-of-Experts model reported to be built on a ~1.5-trillion-parameter "V9" foundation. Vendor benchmarks are mixed against Anthropic's Opus 4.8 — ahead on DeepSWE 1.0 and Terminal-Bench 2.1, behind on DeepSWE 1.1 and SWE-bench Pro — but markedly more token-efficient (~15,900 output tokens on SWE-bench Pro tasks, ~4.2x fewer than Opus 4.8) and served at ~80 tokens/sec. Base pricing $2/$6 per Mtok; Cursor lists a faster variant at $4/$18. Available in Grok Build (default model), Cursor (all plans), and the SpaceXAI console; initially unavailable in the EU (expected mid-July). xAI was absorbed by SpaceX in Feb 2026 and rebranded SpaceXAI.

MoEUndisc. ctxJul 8, 2026

Claude Sonnet 5

Available
AnthropicFrontierProprietary

Anthropic's most agentic Sonnet model yet, with performance approaching Claude Opus 4.8 at a lower price. Built for coding, tool use (browsers and terminals), and autonomous multi-step agentic work, with selectable effort levels up to xhigh. A substantial upgrade over its predecessor Sonnet 4.6 on reasoning, tool use, coding, and knowledge work, with gains shown on SWE-bench, OSWorld-Verified, BrowseComp, and Humanity's Last Exam. The default model on the Free and Pro plans and available to Max, Team, and Enterprise users; in Claude Code and via the Claude API as claude-sonnet-5. Uses an updated tokenizer (same approach as Opus 4.7). Ships with real-time cyber safeguards enabled by default. Introductory API pricing of $2/$10 per Mtok through Aug 31, 2026, then standard $3/$15.

Undisc.200K ctxJun 30, 2026

Seed 2.1 Turbo

Available
ByteDance SeedProprietary

The low-cost, low-latency tier of ByteDance's Seed 2.1 family (served as Doubao-Seed-2.1-turbo), built for large-scale production with full features and performance ByteDance positions as comparable to Seed 2.1 Pro. Shares the family's coding, long-chain agent, and multimodal-understanding focus and 256K-token context, priced for high-volume online calls. Proprietary; available via Doubao and Volcano Engine. List price ¥3 / ¥15 per million input/output tokens — roughly half the Pro tier.

Undisc.256K ctxJun 23, 2026

Seed 2.1 Pro

Available
ByteDance SeedFrontierProprietary

ByteDance's flagship next-generation agent model (served as Doubao-Seed-2.1-pro), built for the "coding and agent era." A deep-thinking model tuned for strong demand understanding, long-horizon planning, and continuous self-repair across complex coding, long-chain agents, and multi-step engineering delivery, with a 256K-token context. ByteDance reports its core coding, agent, and multimodal capabilities are comparable to GPT-5.5, with the highest score on GDPVal, top-tier results on the Agents' Last Exam, the highest score on MobileWorld, and SOTA results across several visual and video-understanding benchmarks (CharXiv-RQ, MeasureBench, TVBench, TOMATO). Proprietary; available via Doubao and Volcano Engine. List price ¥6 / ¥30 per million input/output tokens.

Undisc.256K ctxJun 23, 2026

Kimi K2.7 Code

Available
Moonshot AIFrontierOpen weights

Moonshot's open coding-focused agentic model built on K2.6, with native vision/video input, forced thinking mode, and stronger long-horizon software-engineering performance.

MoE1T262K ctxJun 18, 2026

MiniMax-M3

Available
MiniMaxFrontierOpen weights

Native multimodal MiniMax model with a one-million-token context, sparse attention, and agentic coding/cowork positioning.

MoE428B1M ctxJun 16, 2026

DiffusionGemma 26B-A4B

Available
Google DeepMindOpen source

An open-weight text-diffusion model built on the Gemma 4 26B-A4B MoE backbone (25.2B total / 3.8B active). Denoises text in parallel 256-token blocks for up to ~4x faster generation (1,000+ tok/s on an H100), with a 256K context and text, image, and video input. Apache-2.0.

MoE25.2B256K ctxJun 10, 2026

Claude Fable 5

Available
AnthropicFrontierProprietary

The public, guardrailed sibling of Claude Mythos 5 and Anthropic's most capable widely released model, built for long-horizon agentic work, coding, vision, and knowledge workflows. Launched June 9, 2026 across the Claude API, AWS, and Microsoft Foundry, then suspended three days later under a U.S. export-control directive. Anthropic says those controls were lifted June 30 and Fable 5 was restored globally on July 1 across Claude Platform, Claude.ai, Claude Code, and Claude Cowork, with cloud partner access being re-enabled. Its safeguards route flagged cybersecurity, biology/chemistry, and distillation requests to Claude Opus 4.8.

Undisc.1M ctxJun 9, 2026

Qwen3.7-Plus

Available
Alibaba (Qwen)Proprietary

Multimodal sibling of Qwen3.7-Max that adds vision input and GUI grounding for screen perception, browser automation, and hybrid GUI+CLI agent workflows. 1M-token context; closed-weights and API-only. Previewed at the May 2026 Alibaba Cloud Summit and reached general availability in June 2026 at a low price point ($0.40/$1.60 per 1M tokens).

MoEUndisc.1M ctxJun 3, 2026

Gemma 4 12B

Available
Google DeepMindOpen source

A dense 12B member of the Gemma 4 family with a unified, encoder-free multimodal architecture: vision and audio are projected straight into the LLM backbone. First medium-size Gemma to natively ingest audio; runs on a 16GB laptop. 256K context, Apache-2.0.

Dense12B256K ctxJun 3, 2026

Nex-N2-Pro

Available
Nex AGIOpen source

Nex AGI's open-weight agentic flagship, post-trained on Qwen3.5-397B-A17B (397B total / ~17B active MoE) by the Shanghai Innovation Institute-led Nex alliance. Built around an "Agentic Thinking" framework for long-horizon coding, deep research, tool calling, and terminal execution; accepts text and image input and emits text with explicit reasoning traces and function calling. Apache-2.0, ~262K context. Nex reports parity with GPT-5.5 and Claude Opus 4.7 on several agentic and coding evals. A smaller Nex-N2-mini (35B/3B-active) was announced but is not yet open-sourced.

MoE397B262K ctxJun 2, 2026

Step-3.7-Flash

Available
StepFunOpen source

StepFun's high-efficiency multimodal sparse-MoE successor to Step-3.5-Flash: a ~196B-total / ~11B-active vision-language model with native image and video understanding, a 256K context, and selectable reasoning tiers (high/medium/low). Tuned for coding agents and search workflows.

MoE196B256K ctxMay 29, 2026

Claude Opus 4.8

Available
AnthropicFrontierProprietary

Anthropic's most capable model, with strengthened agentic and long-running task performance.

Undisc.500K ctxMay 28, 2026

Gemini 3.5 Pro

Preview
Google DeepMindFrontierProprietary

Announced at Google I/O 2026; emphasizes deep multimodal reasoning over a 2M-token context. Recent reporting says the broad launch slipped from June toward July while testers continue using it in Google Antigravity and LMArena.

MoEUndisc.2M ctxMay 19, 2026

Gemini 3.5 Flash

Available
Google DeepMindProprietary

Google's fast, cost-efficient Gemini 3.5 tier, unveiled at I/O 2026. Multimodal over a 1M-token context and tuned for agentic and coding workflows; Google says it beats Gemini 3.1 Pro on coding and tool-use while running ~4x faster.

Undisc.1M ctxMay 19, 2026

GPT-5.5

Available
OpenAIFrontierProprietary

OpenAI's May 2026 GPT-5.5 release: a stronger frontier workhorse positioned for deep reasoning, coding, multimodal analysis, and long-context agent workflows. OpenAI lists an 800K-token input context and 128K-token output limit, with API pricing at $3 / $20 per million input/output tokens.

Undisc.800K ctxMay 7, 2026

Grok 4.3

Available
xAIFrontierProprietary

xAI's agentic flagship with a 1M-token context and aggressive API pricing.

MoEUndisc.1M ctxMay 6, 2026

Muse Spark

Available
Meta AIFrontierProprietary

Meta's new frontier model behind Meta AI for U.S. users, identified in reporting as the public release of the former Avocado effort and positioned to compete with Gemini, GPT, and Claude on multimodal assistant tasks.

Undisc. ctxApr 8, 2026

Gemma 4 31B

Available
Google DeepMindOpen source

Google DeepMind's Gemma 4 advanced-reasoning open model for personal computers, part of the April 2026 Gemma 4 family.

Dense31B ctxApr 2, 2026

GLM-5V-Turbo

Available
Z.ai (Zhipu AI)Proprietary

Z.ai's native-multimodal vision agent: the first GLM model designed from the start as a multimodal agent, taking image, video, and text input and producing agent-oriented output (tool calling, task decomposition, and GUI interaction). Served via API with a ~203K-token context.

Undisc.203K ctxApr 1, 2026

Kimi K2.6

Available
Moonshot AIFrontierOpen weights

Moonshot's open native multimodal agentic model for long-horizon coding, visual interface generation, and autonomous tool orchestration.

MoE1T256K ctxMar 30, 2026

Mistral Small 4

Available
Mistral AIOpen source

Mistral's March 2026 Small release: the first Mistral model to unify reasoning (Magistral), multimodal understanding (Pixtral), and agentic coding (Devstral) into one Apache 2.0 model. A 119B-total / ~6B-active Mixture-of-Experts (128 experts, 4 active per token) with native text+image input, a 256K context, and a configurable reasoning_effort toggle for fast or deep responses. API pricing is $0.15 / $0.60 per million input/output tokens.

MoE119B256K ctxMar 16, 2026

Meta Avocado

Retired
Meta AIFrontierProprietary

Historical rumor/codename row retained for provenance. Recent reporting identifies the public productized model as Muse Spark, now tracked separately in the catalog.

Undisc. ctxMar 13, 2026

GPT-5.4

Available
OpenAIFrontierProprietary

Workhorse GPT-5 release with a dedicated Thinking mode; widely deployed across ChatGPT and the API.

MoEUndisc.400K ctxMar 5, 2026

Qwen3.5-9B

Available
Alibaba (Qwen)Open source

The flagship of Alibaba's small dense Qwen3.5 models. Independent analysis (Artificial Analysis) rated it the most intelligent model under 10B parameters at launch — roughly double the score of the next-closest sub-10B models — and the most intelligent multimodal model under 15B, leading peers on MMMU-Pro (~69%). A dense 9B with native vision, a 262K-token context, and the Qwen3.5 family's unified hybrid thinking / non-thinking mode. Native weights are BF16; in 4-bit it needs ~6GB, within reach of consumer laptops. High intelligence comes with heavy reasoning token usage (~260M output tokens to run the Intelligence Index).

Dense9B262K ctxMar 2, 2026

Qwen3.5-4B

Available
Alibaba (Qwen)Open source

A dense 4B in Alibaba's small Qwen3.5 family, rated by Artificial Analysis as the most intelligent model under 5B parameters at launch — outscoring several 7B–9B peers despite roughly half the parameters. Native vision, a 262K-token context, and the family's hybrid thinking / non-thinking mode; Apache-2.0 licensed. Scores ~65% on MMMU-Pro multimodal reasoning and runs in ~3GB at 4-bit, suitable for lightweight on-device agents.

Dense4B262K ctxMar 2, 2026

Qwen3.5-2B

Available
Alibaba (Qwen)Open source

A dense 2B Qwen3.5 model built for high-throughput, low-latency edge and on-device use. Despite its size it matches a 7B-class peer on Artificial Analysis's Intelligence Index. Apache-2.0, with native vision, a 262K-token context, and the family's hybrid thinking / non-thinking mode; runs in under 2GB at 4-bit, fitting laptops and smartphones.

Dense2B262K ctxMar 2, 2026

Qwen3.5-0.8B

Available
Alibaba (Qwen)Open source

The smallest Qwen3.5 model — a dense 0.8B designed for the most constrained on-device deployments, operating in non-thinking (instruct) mode by default. Apache-2.0, with native vision, a 262K-token context, and the family's hybrid thinking / non-thinking mode; needs roughly 2GB of VRAM and runs under 2GB at 4-bit, targeting smartphones and embedded hardware. Notable for a sub-1B model, it still scores ~26% on MMMU-Pro multimodal reasoning.

Dense0.8B262K ctxMar 2, 2026

Qwen3.5-397B

Available
Alibaba (Qwen)FrontierOpen source

Native vision-language MoE supporting 201 languages with a 1M-token context.

MoE397B1M ctxFeb 20, 2026

Gemini 3.1 Pro

Available
Google DeepMindFrontierProprietary

Generally available multimodal flagship with native tool use and a 2M-token context.

MoEUndisc.2M ctxFeb 19, 2026

Claude Opus 4.6

Available
AnthropicFrontierProprietary

Introduced genuinely autonomous multi-file coding and stronger computer use.

Undisc.200K ctxFeb 5, 2026

GPT-5.3-Codex

Available
OpenAIProprietary

OpenAI's February 2026 Codex update, optimized for agentic software engineering in ChatGPT, Codex, and the API. GPT-5.3-Codex improves code quality, patch reliability, repository-scale reasoning, and long-running autonomous coding workflows while keeping the 400K-token input context and 128K-token output limit of the Codex line.

Undisc.400K ctxFeb 5, 2026

Kimi K2.5

Available
Moonshot AIFrontierOpen weights

Open multimodal Kimi model that adds native visual agentic intelligence, instant and thinking modes, and agent-swarm workflows on top of the K2 base.

MoE1T256K ctxJan 27, 2026

GPT-5.2-Codex

Available
OpenAIProprietary

OpenAI's December 2025 Codex model for agentic coding, released after GPT-5.2 with stronger repository understanding, code generation, and tool-use behavior for software engineering agents. The API model is listed with a 400K-token input context, 128K-token output limit, and $1.25 / $10 per million input/output tokens.

Undisc.400K ctxDec 18, 2025

GPT-5.2

Available
OpenAIFrontierProprietary

OpenAI's December 2025 GPT-5.2 general model release, positioned as a stronger default for reasoning, coding, vision, instruction following, and long-context analysis. OpenAI lists a 400K-token input context, 128K-token output limit, and $2 / $12 per million input/output tokens.

Undisc.400K ctxDec 11, 2025

GLM-4.6V

Available
Z.ai (Zhipu AI)Open source

Open 106B-class vision-language model with native multimodal function calling for visual agents.

MoE106B128K ctxDec 8, 2025

Mistral Large 3

Available
Mistral AIFrontierOpen weights

Mistral's largest open-weight MoE, aimed at frontier reasoning while remaining self-hostable.

MoE675B256K ctxDec 2, 2025

Claude Haiku 4.5

Available
AnthropicProprietary

Anthropic's fast, low-cost Claude 4.5 model, released in October 2025 for latency-sensitive coding, tool-use, and customer-facing agent workloads. Anthropic positions it as bringing near-Sonnet capability to the Haiku tier at substantially lower cost and higher speed.

Undisc.200K ctxOct 15, 2025

Claude Sonnet 4.5

Available
AnthropicFrontierProprietary

Anthropic's September 2025 Sonnet release, positioned as its strongest model for coding, agents, and computer-use workflows at launch. Proprietary API model with text, vision, and code capabilities, 200K context, and Sonnet-tier list pricing.

Undisc.200K ctxSep 29, 2025

Gemma 3 27B

Available
Google DeepMindOpen weights

Google's open multimodal model: 128k context, 140+ languages, runs on a single GPU.

Dense27B128K ctxSep 4, 2025

GLM-4.5V

Available
Z.ai (Zhipu AI)Open source

Vision-language GLM based on GLM-4.5-Air, covering image, video, document, grounding, and GUI-agent tasks.

MoE106B ctxAug 11, 2025

Grok 5

Rumored
xAIFrontierProprietary

Rumored next major Grok model. Elon Musk said after GPT-5's launch that Grok 5 would arrive before the end of 2025, but no broad public Grok 5 release is logged in this catalog yet.

Undisc. ctxAug 8, 2025

Claude Opus 4.1

Available
AnthropicFrontierProprietary

Anthropic's August 2025 Opus point release, focused on stronger coding, reasoning, and agentic reliability over Claude Opus 4. Proprietary API model with text, vision, and code capabilities.

Undisc.200K ctxAug 5, 2025

Gemini 2.5 Deep Think

Available
Google DeepMindFrontierProprietary

Google's enhanced Gemini 2.5 reasoning mode for harder math, science, coding, and multimodal analysis. Previewed at Google I/O 2025 and later made available to Gemini app subscribers, Deep Think uses more deliberative reasoning for complex prompts.

Undisc.1M ctxAug 1, 2025

Gemini 2.5 Flash-Lite

Available
Google DeepMindProprietary

Google's lowest-latency, lowest-cost Gemini 2.5 tier, designed for summarization, classification, extraction, routing, and other high-volume production tasks. Proprietary API model with a 1M-token context and multimodal support.

Undisc.1M ctxJul 22, 2025

Grok 4

Deprecated
xAIProprietary

xAI's fourth-generation Grok line, preceding the later 4.x API updates already tracked in the catalog.

Undisc. ctxJul 9, 2025

ERNIE-4.5-VL-424B-A47B

Available
BaiduOpen source

Baidu's largest ERNIE 4.5 vision-language MoE, supporting text, image, and video inputs with thinking and non-thinking modes.

MoE424B128K ctxJun 30, 2025

Kimi-VL-A3B-Thinking-2506

Available
Moonshot AIOpen source

Updated MIT-licensed Kimi-VL reasoning model with better multimodal reasoning, video understanding, high-resolution perception, and lower thinking-token use.

MoE16B128K ctxJun 21, 2025

Gemini 2.5 Flash

Available
Google DeepMindProprietary

Google's faster, lower-cost Gemini 2.5 model for high-throughput multimodal and agentic workloads. It brought Gemini 2.5's reasoning improvements to a production Flash tier with a 1M-token context and broad text, image, audio, video, and coding support.

Undisc.1M ctxJun 17, 2025

Claude Opus 4

Deprecated
AnthropicProprietary

First Claude 4 Opus model, positioned for long-running agentic and coding work before the 4.x point releases.

Undisc.200K ctxMay 22, 2025

Gemma 3n E4B

Available
Google DeepMindOpen weights

Google's mobile-first Gemma 3n model variant, built with a MatFormer-style architecture for efficient on-device multimodal inference. The E4B variant has roughly 4B effective parameters, supports text, vision, audio, and video-oriented use cases, and is released under Gemma terms.

Hybrid8B32K ctxMay 20, 2025

Kimi-Audio-7B-Instruct

Available
Moonshot AIOpen source

Open audio foundation model for audio understanding, generation, speech recognition, audio QA, captioning, and speech conversation.

Hybrid10B ctxApr 25, 2025

Kimi-VL-A3B-Instruct

Available
Moonshot AIOpen source

Efficient MIT-licensed vision-language MoE for OCR, image/video understanding, long documents, and OS-style agent tasks.

MoE16B128K ctxApr 17, 2025

OpenAI o3

Available
OpenAIProprietary

Reasoning model released alongside o4-mini with tool use, image reasoning, and stronger agentic problem solving.

Undisc. ctxApr 16, 2025

GPT-4.1

Deprecated
OpenAIProprietary

API model family focused on coding, instruction following, and one-million-token long-context work.

Undisc.1M ctxApr 14, 2025

Llama 4 Maverick

Available
Meta AIFrontierOpen weights

Meta's flagship open-weight MoE; highest MMLU among open models at release.

MoE400B1M ctxApr 5, 2025

Llama 4 Scout

Available
Meta AIOpen weights

Efficient open-weight MoE designed for very long context on modest hardware.

MoE109B10M ctxApr 5, 2025

Llama 4 Behemoth

Announced
Meta AIFrontierOpen weights

Meta's announced but unreleased Llama 4 teacher model: a multimodal MoE with 288B active parameters and nearly 2T total parameters. Meta says it was still training when Scout and Maverick shipped and that those released models were distilled from Behemoth.

MoE2T ctxApr 5, 2025

Qwen2.5-Omni-7B

Available
Alibaba (Qwen)Open weights

Local omni-modal Qwen model that supports text, image, audio, video, and speech generation in a 7B package.

Dense7B ctxMar 26, 2025

Gemini 2.5 Pro

Deprecated
Google DeepMindProprietary

Reasoning-focused Gemini 2.5 model that made thinking a core part of Google's flagship model line.

Undisc.1M ctxMar 25, 2025

Mistral Small 3.1

Available
Mistral AIOpen source

Apache-licensed Small update adding vision and a 128K context window to the efficient 24B line.

Dense24B128K ctxMar 17, 2025

Claude 3.7 Sonnet

Retired
AnthropicProprietary

Anthropic's first hybrid-reasoning Sonnet. Shut down May 11, 2026 as the 4.x line matured.

Undisc.200K ctxFeb 24, 2025

Grok 3

Deprecated
xAIProprietary

xAI's third-generation model family, introduced with stronger reasoning, search, and coding modes.

Undisc. ctxFeb 17, 2025

Qwen2.5-VL-72B

Available
Alibaba (Qwen)Open weights

Vision-language Qwen2.5 model for image, document, video, and agentic visual grounding tasks.

Dense72B128K ctxJan 26, 2025

Doubao-1.5-pro

Available
ByteDance SeedProprietary

Doubao 1.5 Pro update positioned for stronger multimodal, reasoning, and agentic work in Volcano Engine.

Undisc. ctxJan 22, 2025

Kimi k1.5

Available
Moonshot AIProprietary

Moonshot's multimodal reinforcement-learning reasoning model, reported as matching OpenAI o1 on math, coding, and multimodal reasoning.

Undisc. ctxJan 20, 2025

MiniMax-01

Available
MiniMaxOpen weights

Open MiniMax generation with MiniMax-Text-01 and MiniMax-VL-01 long-context models.

Hybrid456B4M ctxJan 15, 2025

Step-2

Available
StepFunProprietary

Second-generation StepFun foundation model line with larger-scale multimodal and reasoning ambitions.

Undisc. ctxDec 23, 2024

Gemini 2.0 Flash

Deprecated
Google DeepMindProprietary

First Gemini 2.0 release, built for native multimodal input/output, tool use, and agentic product integrations.

Undisc.1M ctxDec 11, 2024

OpenAI o1

Deprecated
OpenAIProprietary

General release of OpenAI's o1 reasoning model with stronger deliberative reasoning and multimodal ChatGPT integration.

Undisc. ctxDec 5, 2024

Amazon Nova Pro

Available
AmazonProprietary

AWS-native multimodal model with a 300k context; size and architecture undisclosed.

Undisc.300K ctxDec 3, 2024

Amazon Nova Lite

Available
AmazonProprietary

Lower-cost multimodal Nova understanding model for text, image, and video inputs.

Undisc.300K ctxDec 3, 2024

Claude 3.5 Haiku

Deprecated
AnthropicProprietary

Fast, lower-cost Claude 3.5 model for latency-sensitive coding, tool-use, and customer-facing workloads.

Undisc.200K ctxOct 22, 2024

Llama 3.2 90B Vision

Available
Meta AIOpen weights

First Llama family release with native vision models, alongside smaller edge-oriented 1B and 3B text models.

Dense90B128K ctxSep 25, 2024

Molmo 72B

Available
Allen Institute for AI (Ai2)Open weights

Open multimodal model family trained for strong image understanding, pointing, and visual grounding.

Dense72B ctxSep 25, 2024

Pixtral 12B

Available
Mistral AIOpen source

Mistral's first open multimodal model, adding image understanding to a Mistral text backbone.

Dense12B128K ctxSep 17, 2024

Grok-2

Retired
xAIProprietary

Second-generation Grok release with Grok-2 and Grok-2 mini for chat, coding, reasoning, and image-enabled product experiences.

Undisc. ctxAug 13, 2024

MiniCPM-V 2.6

Available
OpenBMBOpen weights

8B vision-language model for local image, multi-image, OCR, and video understanding, with llama.cpp and Ollama support.

Dense8B ctxAug 2, 2024

Claude 3.5 Sonnet

Retired
AnthropicProprietary

Major Sonnet upgrade that became Anthropic's default high-intelligence workhorse for coding, writing, and visual reasoning.

Undisc.200K ctxJun 20, 2024

GPT-4o

Retired
OpenAIProprietary

The 2024 omni-modal model that defined a generation of assistants. Deprecated in Feb 2026 and fully retired across ChatGPT on April 3, 2026.

Undisc.128K ctxMay 13, 2024

Falcon 2 11B

Available
Technology Innovation InstituteOpen weights

Falcon 2 generation, including text and vision-language 11B models under a permissive TII license.

Dense11B8K ctxMay 13, 2024

Step-1V

Available
StepFunProprietary

StepFun's first major vision-language model, released after the Step-1 language model.

Undisc. ctxApr 12, 2024

Claude 3 Opus

Deprecated
AnthropicProprietary

Highest-capability Claude 3 model, launched with Sonnet and Haiku and Anthropic's first major vision-capable Claude family.

Undisc.200K ctxMar 4, 2024

Gemini 1.5 Pro

Deprecated
Google DeepMindProprietary

Gemini generation that introduced production-scale long context, eventually expanding to a two-million-token window.

MoEUndisc.2M ctxFeb 15, 2024

GLM-4

Available
Z.ai (Zhipu AI)Proprietary

Zhipu's GLM-4 flagship generation, launched as the successor to ChatGLM3 with stronger tool use and multimodal variants.

Undisc.128K ctxJan 16, 2024

Gemini 1.0 Ultra

Deprecated
Google DeepMindProprietary

Google's first natively multimodal Gemini flagship, since superseded by the 1.5/2/3 lines.

Undisc.32K ctxDec 6, 2023

GPT-4 Turbo

Deprecated
OpenAIProprietary

Lower-cost GPT-4 generation with a 128K context window, introduced at OpenAI DevDay.

Undisc.128K ctxNov 6, 2023

ERNIE 4.0

Available
BaiduProprietary

Baidu's fourth-generation ERNIE flagship, announced with stronger understanding, generation, reasoning, and memory.

Undisc. ctxOct 17, 2023

LLaVA 1.5 13B

Available
LLaVAOpen weights

Open vision-language assistant and one of the most widely run early local multimodal models.

Hybrid13B ctxSep 30, 2023

EXAONE 2.0

Retired
LG AI ResearchProprietary

Second EXAONE generation, improving bilingual Korean-English performance and enterprise deployment options.

Undisc. ctxJul 19, 2023

GPT-4

Deprecated
OpenAIProprietary

The model that brought reliable multi-step reasoning to the mainstream; size never disclosed.

Undisc.8K ctxMar 14, 2023

EXAONE 1.0

Retired
LG AI ResearchProprietary

LG AI Research's first EXAONE foundation model generation, introduced as a large multimodal expert AI.

Undisc. ctxDec 14, 2021