The model release tracker
Updated Sep 14, 2026 / latest tracked event Sep 4, 2026
Every large language model release, sourced and tracked over time.
Frontier and open-weight. Filter by access, license, size, modality, and country of origin. Each entry links to a primary source and keeps a full history of changes, deprecations, and retractions.
349 models
GPT-6 Astra Pro
AvailableThe higher-quality reasoning tier of OpenAI's GPT-6 Astra, released Sep 4, 2026 (API id gpt-6-astra-pro). It is the same underlying model as GPT-6 Astra served with reasoning.mode set to 'pro', trading latency for higher-quality responses on the hardest professional, coding, research, computer-use, and agentic tasks. Carries a 1.05M-token context window and up to 128K output tokens. Standard pricing is $10/$50 per Mtok input/output (cached input $1/Mtok, batch at half price), with a Fast mode at ~2x the standard rate. Access is limited to ChatGPT Pro, Business, and Enterprise users (and the pro reasoning mode on the API); off by default at launch and enabled per workspace. Shares GPT-6 Astra's positioning around state-of-the-art computer/browser use, finished professional artifacts, long-session Codex coding, and cybersecurity capability that crosses the 'Critical' threshold of OpenAI's Preparedness Framework — so exploit-creation capabilities ship gated behind OpenAI's Daybreak program.
Ling-3.0-flash-Sante
AvailableA health- and medicine-domain-tuned variant of Ant Group inclusionAI's Ling-3.0-flash, launched September 4 2026 (model id inclusionai/ling-3.0-flash-sante). Same efficient sparse Mixture-of-Experts base — 124B total parameters, ~5.1B active per token — post-trained for medical knowledge reasoning, clinical safety, evidence-based retrieval, and long-horizon medical tasks, while inclusionAI reports it retains strong general reasoning, coding, and agentic ability. Text-in / text-out only (no vision), with a 262,144-token (256K) context and up to 32,768 output tokens; supports reasoning and tool / function calling. Available first via hosted serverless API — Novita, OpenRouter (inclusionai/ling-3.0-flash-sante), and Vercel AI Gateway — with a time-limited free window at launch (free through Oct 4 on Vercel AI Gateway). Positioned as a developer API for research, retrieval, summarization, and workflow assistance, explicitly not a medical device or a substitute for clinical judgment; no public benchmark table at launch, so treat domain claims as unverified. Like the Fin variant, Sante-specific open weights were not confirmed posted at launch (the base Ling-3.0-flash family ships open-weight, MIT), so treat the open-weight status as announced-but-unconfirmed for this variant.
GPT-6 Astra
AvailableOpenAI's new frontier flagship, released Sep 3, 2026 (API id gpt-6-astra) as the successor to GPT-5.6 Sol and OpenAI's most capable and most aligned model to date. Positioned around three areas: state-of-the-art computer/browser use, a step change in producing finished professional artifacts (documents, slides, spreadsheets, and websites via Sites in ChatGPT that follow the user's templates), and a jump in cybersecurity capability that crosses the 'Critical' threshold of OpenAI's Preparedness Framework. It decides when to ask clarifying questions versus proceed on assumptions, holds constraints across mid-task steering, and — with an updated Codex harness — keeps searchable notes across context windows instead of lossy compaction. Carries a 1M-token context window; available as gpt-6-astra in the OpenAI API and on Amazon Bedrock (with Microsoft Azure), plus a GPT-6 Astra Pro tier for ChatGPT Pro, Business, and Enterprise (off by default at launch, enabled per workspace). Standard API pricing is $10/$50 per Mtok input/output, with a Fast mode at ~2x price (~$20/$100) for up to 2.5x speed, plus separate cache rates. Vendor-reported benchmarks include OSWorld 2.0 72.6% (at ~47% less time per task than Sol), FrontierMath Tier 4 v2 97.6%, GPQA Diamond 96.0%, Terminal-Bench 4.0 57.7%, and ExploitBench 100%; the marquee ARC-AGI-3 ~99.9% depends on a stateful provider-adapter harness (stateless runs score far lower), and it trails Claude Fable 5.1 on Humanity's Last Exam with tools (57.2% vs 65.0%). Because it crosses the Critical cyber threshold, exploit-creation capabilities ship gated behind OpenAI's Daybreak program; OpenAI also flags a regression in chain-of-thought monitorability as a research priority.
Gemini 3.8 Flash
AvailableGoogle DeepMind's fast, cost-efficient Flash model released Sep 2, 2026 — its fourth Flash model in under four months and the successor to Gemini 3.7 Flash. At high reasoning it scores 59 on the Artificial Analysis Intelligence Index (up 3 points from Gemini 3.7 Flash and level with sub-maximum efforts of GPT-5.6 Sol and Grok 4.6), 57 at medium (matching GPT-5.6 Terra and Muse Spark 1.2), and 52 at low (matching Gemini 3.6 Flash at ~30% lower cost per task). The improvement is driven mainly by agentic evaluations — t^3-Banking tool use (+12 points to 45%), Terminal-Bench v2.1 coding, and GDPval-AA v2 real-world tasks. It keeps a 1M-token context window and multimodal input (text, image, video, speech) with text output. Pricing matches Gemini 3.7 Flash's current discounted rate of $0.75/$3.75 per Mtok input/output through the end of 2026 ($1.50/$7.50 at standard pricing), with cached input keeping a 90% discount; a ~30% rise in average output tokens per task (to ~48k) lifts cost per task to ~$0.58 at high reasoning despite unchanged per-token pricing. Available in the Gemini app for AI Pro and Ultra subscribers, AI Mode, and Gemini in Google Sheets, and for developers via Google Antigravity, AI Studio, and the Gemini API. Benchmark figures are vendor/third-party-reported.
Gemini 3.8 Flash Cyber
PreviewThe cybersecurity-tuned sibling of Gemini 3.8 Flash, introduced by Google DeepMind on Sep 2, 2026 in the same launch. It shares the same foundational intelligence as Gemini 3.8 Flash but ships with a more permissive set of mitigations for defensive security work, and is available only to trusted defenders — government authorities, critical-infrastructure operators, and software maintainers — through the new Fairwind Program. Google reports frontier-level autonomous vulnerability discovery (surpassing Gemini 3.5 Flash Cyber and much larger frontier models on the CyberGym benchmark), a success rate exceeding 70% on an internal real-world vulnerability benchmark spanning 20 programming languages, and 47.2% pass@1 on the external CWE-Bench patching benchmark (on the Pareto frontier versus a leading frontier model at 47.8%, at far lower cost). Google is already using it internally: the Chrome Security team reports 2.6x more correct patches than the best larger commercial models, and its Cloud Vulnerability Research team found a critical foundational vulnerability in under two hours. Keeps a 1M-token context and multimodal input with text output; prioritizes vulnerability fixing over offensive capability. Not publicly token-billed (trusted-access only), so no list price is recorded. Benchmark figures are vendor-reported.
Muse Spark 1.3
AvailableMeta's successor to Muse Spark 1.2, released Sep 2, 2026. A multimodal reasoning model built for long-running agentic, multi-agent, and coding workflows: it is designed to keep track of information across extended tasks, work through conflicting inputs, and request clarification or confirmation when needed, with an emphasis on concise execution. Accepts text and image input over a 1M-token (1,048,576) context and returns text. Standard-tier API pricing is $1.25 / $4.25 per 1M input/output tokens ($0.15 cached input); as with prior Muse Spark releases a lower-cost muse-spark-1.3-contributor tier is offered in exchange for permission to train future Meta models on prompts and completions. Hosted by Meta and served via OpenRouter (meta/muse-spark-1.3).
Qwen3.8-Max-0902
AvailableAn updated snapshot of Alibaba's flagship Qwen3.8-Max, released Sep 2 2026 (model id qwen3.8-max-0902; alias qwen3.8-max-2026-09-02). A post-training upgrade on the unchanged 2.4-trillion-parameter Mixture-of-Experts base (~95B active per token), it sharpens coding for engineering-scale projects and long-horizon autonomous development, multi-tool agent orchestration, and native vision (chart reasoning, document parsing, multimodal perception). Accepts text, image, and video input and returns text, with a 1M-token context, up to ~131K output tokens, and an optional thinking mode carrying up to 256K chain-of-thought tokens. Pricing is unchanged from Qwen3.8-Max at $2.00 / $6.00 per 1M input/output tokens. Vendor-reported gains include a +22-point CodeArena jump to 1,691 (first on that leaderboard at launch); figures are unverified. Served API-first (Alibaba, OpenRouter); the -Max flagship tier is closed-weight, so treat this snapshot as proprietary despite the open-weight base family. Distinct from the base Qwen3.8-Max (2026-08-03).
Claude Mythos 5.1
AvailableThe restricted sibling of Claude Fable 5.1, announced Sep 1, 2026. It is the same underlying model as Fable 5.1 but ships with more permissive safeguards for vetted users, available only through Anthropic's trusted-access programs — the Cyber Verification Program (defensive security work) and the Life Sciences Verification Program (professional biology R&D, developed with the US government). Anthropic reports it has the strongest cyber capabilities of any model it has released (still within the lower risk tier of its Frontier Compliance Framework) and scores 60.9% on Terminal-Bench 4.0 (vs 55.8% for safeguarded Fable 5.1). Also powers Claude Security's codebase vulnerability scanning. Access is currently limited to a set of US organizations. Not publicly token-billed, so list pricing is not recorded.
Claude Fable 5.1
AvailableAnthropic's most capable widely released model, launched Sep 1, 2026 (API id claude-fable-5-1) as a direct upgrade to Claude Fable 5 for agentic coding, long-running problem-solving, and knowledge work. Anthropic reports gains across coding, computer use, multidisciplinary reasoning, and finance/analysis while producing more concise plans and summaries, and cites benchmark results including Terminal-Bench 4.0 55.8%, Terminal-Bench-Science 0.1 52.6%, Humanity's Last Exam 65.0% (with tools), and CursorBench 3.2.0 73.4%. Base pricing is unchanged from Fable 5 at $10/$50 per Mtok input/output, but cache reads were cut 75% to $0.25/Mtok, reducing typical workload cost ~25% (up to ~45% for highly agentic work). Carries a 1M-token context and up to 128K output tokens; available on the Claude API, AWS, Google Cloud, and Microsoft Azure. Cybersecurity safeguards were loosened to allow defensive vulnerability discovery (Claude Code sees ~60% fewer false-positive interventions), while dual-use cyber and life-sciences R&D requests still route to Opus models. First Anthropic release to carry the EU AI Act text watermark.
Mercury 2.5 Preview
PreviewThe latest diffusion large language model (dLLM) from Inception, released Aug 31, 2026. Instead of generating tokens sequentially, Mercury 2.5 produces and refines many tokens in parallel, reaching ~1,107 tokens/sec on standard GPUs. Inception positions it as the fastest reasoning LLM, reporting a 10+ point intelligence gain over Mercury 2 and quality comparable to cost-optimized frontier models (GPT-5.6 Luna Low, Gemini 3.5 Flash-Lite, Claude Haiku 4.5). It supports tunable reasoning levels, parallel tool calls, and schema-aligned JSON output, and targets latency-sensitive production workloads such as search agents, voice pipelines, and coding subagents. 260K-token context, up to 65,536 output tokens; API-only (proprietary) via Inception and OpenRouter (inception/mercury-2.5-preview). List pricing is $0.20/$0.75 per Mtok input/output (a launch promo ran at $0.04/$0.15). Architecture recorded as unknown: it is a diffusion LM rather than a standard dense/MoE autoregressive transformer.
Hy4 preview
AvailableTencent Hunyuan's fourth-generation flagship, released and open-sourced under Apache-2.0 on Aug 28 2026 (weights and an FP8 variant on Hugging Face as tencent/Hy4-preview). A Mixture-of-Experts model with 770B total parameters and ~49B activated per token: a 78-layer backbone with 256 routed experts plus a shared expert in most layers (8 routed experts selected per token) and a native multi-token-prediction layer for speculative decoding. Carries a 1M-token (1,048,576) context and is aimed at long-horizon software engineering (planning, debugging, verification across extended tasks with tool calling), office and financial analysis, cross-document work, game-development prototyping, and scientific workloads. Available via Tencent Cloud TokenHub and OpenRouter (model id tencent/hy4-preview, tool calling + structured outputs) or self-hosted on vLLM/SGLang; OpenRouter listed launch pricing of $0.834 input / $2.501 output / $0.042 cache-read per 1M tokens (Aug 28 snapshot). In Tencent's internal blind test (163 experts over 203 engineering tasks) Hy4 averaged 2.99 vs 2.92 for GLM-5.3 and 2.94 for Kimi K3 — narrow margins, and Tencent flags preview-stage tendencies to over-reason and over-verify. All figures vendor-reported and unverified by independent labs at launch.
Parse 5
AvailableCohere's enterprise document-intelligence vision-language model, released Aug 27 2026 (parse-v5.0). A compact 2.3B-parameter VLM that converts PDFs, slides, and images into structured Markdown at scale — detecting and interpreting tables, forms, diagrams, and embedded images, extracting semantic context rather than a flat character stream, and returning bounding boxes for visual elements to support retrieval, grounding, and citation. Trained on business documents across finance, insurance, and scientific research, with support for nine major world languages, an 8,192-token context window, and a ~4.6GB footprint. Priced at $1.50 per 1,000 pages via the Cohere API, with Model Vault for higher-volume/managed deployment and availability on Amazon SageMaker and Microsoft Azure; it can also be deployed privately. Cohere's own comparison places Parse 5 behind larger general-purpose frontier models (GPT-5.5, Opus 4.8, Gemini 3.5 Flash) on raw accuracy, positioning it on price-to-performance and cost per page. Part of Cohere's document-AI line alongside North Micro Vision. Vendor figures, unverified independently at launch.
Ling-3.0-flash-Fin
AvailableA finance-domain-tuned variant of Ant Group inclusionAI's Ling-3.0-flash, launched August 27 2026. Same efficient sparse Mixture-of-Experts base — 124B total parameters, ~5.1B active per token — post-trained on high-quality financial data (developed with financial institutions and domain experts) for real-world investment and banking workflows: annual reports, financial workbooks, multi-document research, information retrieval, investment analysis, and valuation modeling, with an emphasis on complex multi-step tasks and long-horizon planning. Retains a 262,144-token (256K) context and up to 32,768 output tokens, and supports tool / function calling (tools and tool_choice), though it does not enforce structured JSON output (no response_format). inclusionAI reports it preserves strong general reasoning, coding, and math ability alongside the finance gains, citing finance benchmarks such as FinFIRST, FinSearchComp, and SpreadsheetBench — vendor claims, unverified independently. Launched hosted-API-first with a one-month free window through OpenRouter (inclusionai/ling-3.0-flash-fin); the promised open weights followed as announced and are now on Hugging Face under MIT (inclusionAI/Ling-3.0-flash-Fin, confirmed posted by 2026-09-04, with third-party hosting on DeepInfra and community GGUF quantizations).
Qwen3.8-Flash
AvailableThe productionized, managed Qwen Cloud API model announced Aug 26 2026 alongside the open-weight Qwen3.8-Flash-Next, with official QwenCloud pricing confirmed Aug 27-28. It runs the same Qwen4-preview architecture as Flash-Next — a 125B-total sparse Mixture-of-Experts that activates ~6B parameters per token (~180B stored once a 51B n-gram embedding table and a multi-token-prediction module are counted; roughly 95% sparsity), built on Gated-DeltaNet + Qwen Sparse Attention with a gated residual stream and Muon-trained large linear layers — but ships as the hosted service rather than the self-host weights. On QwenCloud it defaults to a 1M-token context with built-in tools and accepts text/image/video in, returning text out. List pricing is $0.15 input / $0.47 output / $0.016 cache-hit per Mtok (domestic China Y0.8/Y2.7/Y0.1), roughly a third of DeepSeek-V4-Flash and about one-thirteenth of the Qwen3.8-Max flagship. Also reachable through OpenCode Go's flat-rate subscription. Distinct catalog row from the open-weight Qwen3.8-Flash-Next (self-host, Qwen Community License 1.0); this managed API is recorded proprietary/API-only. Self-reported benchmarks carry over from Flash-Next (SWE-bench Pro 62.5, DeepSWE 58.7) — all vendor numbers, unverified by independent labs at launch. Thinking on by default.
Qwen3.8-Flash-Next
PreviewAn open-weight, experimental preview of the architecture that will underpin Qwen4, released Aug 26 2026 (Qwen/Qwen3.8-Flash-Next). A sparse MoE with ~6B active parameters (headline 125B-with-6B-activated; ~180B stored once a 51B n-gram embedding table and 4B multi-token-prediction module are counted), 512 experts (10 routed + 1 shared), and a hybrid Gated-DeltaNet + Qwen Sparse Attention design. Native 262,144-token context, extensible to 1M via YaRN. Accepts text, image, and video in and returns text out. Distinct from the managed Qwen Cloud 'Qwen3.8-Flash' API (which defaults to 1M context and bundled tools); this Next build is catalog/self-host only with no hosted list price at launch, served via Transformers, vLLM, SGLang, and TokenSpeed. Weights under the Qwen Community License 1.0. Self-reported vs DeepSeek-V4-Flash-0731: DeepSWE 58.7 vs 54.4, SWE-bench Pro 62.5 vs 56.0, LiveCodeBench v6 91.9, GPQA Diamond 91.7, though NL2Repo 48.1 vs 54.2 is a regression; vision self-reports include AndroidWorld 84.5 and RealWorldQA 88.5 — all vendor numbers, unverified at launch. Thinking on by default.
GLM-5.3-Flash
AvailableZ.ai's first natively multimodal GLM-5 (text, image, and video understanding in one stack), released Aug 26 2026 and stealth-tested beforehand as 'ox-alpha'. A 320B-total / 18B-active MoE (45 layers, hybrid linear + sparse attention) with a 1M-token context window and MIT-licensed open weights (zai-org/GLM-5.3-Flash). It is a distinct model from the text-only flagship GLM-5.3 (753B, $1.40/$4.40) and is priced roughly 10x cheaper on input: list $0.15 / $0.03 cached / $0.50 per Mtok, with a 50% promo through 2026-09-09. Vision sits inside the coding/agent loop (self-visual judgment) rather than as a bolted-on VL head. Self-reported vs GLM-5.2: DeepSWE 63.4 vs 46.2, AutomationBench 48.8 vs 26.2, Terminal-Bench 2.1 84.3 vs 81.0; vision self-reports include CharXiv Reasoning 89.4 and Chartography 78.0, though BabyVision 53.4 trails Gemini 3.7 Flash — all vendor numbers, unverified at launch. Thinking is always on and cannot be disabled.
Granite 4.2 3B
AvailableThe smallest member of IBM's Granite 4.2 open reasoning family, released Aug 25 2026 under Apache-2.0 (ibm-granite/granite-4.2-3b; reports ~4B parameters on Hugging Face) and aimed at local / edge deployment. A dense, decoder-only transformer with the family's thinking / non-thinking switch and low-effort thinking mode, pre-trained from scratch on ~15T tokens with a five-phase curriculum (context extended to a claimed 512K; shipped configuration 131,072 tokens), then SFT on reasoning data and multi-stage RL. Supports native tool calling. Open weights on Hugging Face, Ollama, and GitHub; no hosted list price at launch.
Granite 4.2 8B
AvailableThe mid-size member of IBM's Granite 4.2 open reasoning family, released Aug 25 2026 under Apache-2.0 (ibm-granite/granite-4.2-8b; reports ~9B parameters on Hugging Face). A dense, decoder-only transformer sharing the family's thinking / non-thinking switch and low-effort thinking mode, pre-trained from scratch on ~15T tokens with a five-phase curriculum (context extended to a claimed 512K; shipped configuration 131,072 tokens), then SFT on reasoning/agentic-trajectory data and multi-stage RL. Like the 30B, it is trained to call tools and act inside real sandboxed environments for multi-step software-engineering, terminal, and search-driven tasks. Open weights on Hugging Face, Ollama, and GitHub; no hosted list price at launch.
Granite 4.2 30B
AvailableThe flagship of IBM's Granite 4.2 family, released Aug 25 2026 under Apache-2.0 (ibm-granite/granite-4.2-30b; reports ~29B parameters on Hugging Face). A dense, decoder-only transformer with a thinking / non-thinking switch so one checkpoint can either reason step by step or answer directly, plus a low-effort thinking mode that caps the reasoning budget on easy queries. Pre-trained from scratch on ~15T tokens with a five-phase curriculum that extends context to a claimed 512K (shipped configuration 131,072 tokens), then supervised fine-tuned on chain-of-thought / reasoning / agentic-trajectory data and post-trained with multi-stage RL. Aimed squarely at agent work — native tool calling, multi-step software engineering, terminal tasks, and search-driven workflows — with the 8B and 30B trained to act with tools inside real sandboxed environments. Open weights on Hugging Face, Ollama, and GitHub; no per-token hosted list price at launch.
Apodex 1.1 mini
AvailableThe open-weight, locally deployable member of the Apodex 1.1 release (Aug 24 2026): a disclosed 35B-parameter model that Apodex reports reaching the performance band of selected frontier systems across professional work, finance, and scientific research, with further gains from Agent Team coordination — self-reported, and because several compared systems do not publish parameter counts, the lab avoids explicit size-vs-capability claims. Designed to run under the open-source FrontierAgent harness as a local ReAct or multi-agent "Agent Team" system, keeping files, search, code execution, task state, and delivered artifacts in one workflow. Announced as open weights under Apache 2.0 (Hugging Face: apodex/Apodex-1.1-mini); the lab notes model weights and developer documentation were still rolling out at launch, so downloadable availability is announced-but-in-progress. No independently verified benchmark figures at release.
Apodex 1.1
AvailableThe flagship model of Apodex's Aug 24 2026 release, built for what the lab calls "working capability" — sustained, verifiable progress on long-horizon professional and scientific tasks rather than one-shot answers. It powers an online workbench where a single long-running task can read files (papers, CSVs, spreadsheets, images, code), search, execute analysis code, revise its own plan, and coordinate an asynchronous "Agent Team" of sub-agents, with a separate "Statement Review" step that independently checks key claims before delivery. Trained with SFT plus agentic RL (the lab's PIVOT-RL method concentrates learning on the consequential decision points of long trajectories) across two scaling axes: Environment Scaling and Agentic Coordination Scaling. Apodex reports the full model entering the leading band of current agentic systems on complex professional work, finance, and scientific research while retaining broad reasoning, math, coding, and deep-search coverage — self-reported figures (e.g. APEX-Agents 38.5, GDPVal 78.8, FrontierFinance 54.3, FrontierScience-Research 63.3) with no independent audit at launch. Parameter count is not disclosed; the lab explicitly declines size comparisons. Available now via the apodex.ai web workbench with bare-model API access rolling out through major platforms; open weights and full developer docs are in progress.
Thomson
PreviewThomson Reuters' first in-house proprietary large language model, launched August 24 2026. Built by starting from a strong (undisclosed) open-source foundation and applying state-of-the-art mid- and post-training on decades of proprietary Thomson Reuters content — Westlaw, Practical Law, Checkpoint, and Reuters — with hundreds of subject-matter experts involved from training-objective design through final evaluation, to the company's 'Fiduciary-Grade' standard. Total training investment is reported at ~$40M (talent + compute), positioned as reaching frontier-level quality at a fraction of typical frontier cost while remaining fully owned and controlled by Thomson Reuters. The company claims early evaluations put Thomson 'on par with the latest frontier models' across a range of tasks, with notable domain-specific uplift in instruction following and dense professional-content reasoning — vendor claims, corroborated only by limited external academic testing at launch and with no public benchmark table, so treat as unverified. Model size, architecture, and context window are undisclosed. First deployment is inside Tabular Analysis in CoCounsel Legal (which remains multi-model by design); broader rollout across the legal and tax portfolio plus sovereign-AI options is planned. A smaller version is being released as an open-weight model on Hugging Face for academic and non-commercial use; the flagship remains proprietary and product/API-only.
DeepSeek-V4-Flash-Vision-Exp
PreviewDeepSeek's first multimodal V4 model — an experimental vision-understanding checkpoint that went live on the DeepSeek API (model='deepseek-v4-flash-vision-exp') on Aug 21, 2026. It extends DeepSeek-V4-Flash with image understanding while keeping its full text capabilities (agents, reasoning, coding, and world knowledge), matching V4-Flash on text benchmarks. DeepSeek reports a major jump on multimodal agent benchmarks over V4-Flash, bringing multimodal-agent performance close to Opus-4.8 — a vendor-reported result, unverified by an independent harness at launch. Keeps V4-Flash's 284B-total / 13B-active sparse MoE architecture and 1M-token context, with up to ~393K output tokens; accepts text plus up to 600 images per request (8,192px per side, 64 MiB payload) and returns text only, with images billed at up to 384 tokens each. API pricing held at V4-Flash rates: $0.22 / $0.66 per 1M input/output tokens, with a $0.007 per 1M cached-input rate. API-only and experimental at launch — weights were not published, so treated as proprietary.
Hy-MT2-30B-A3B
AvailableThe flagship of Tencent Hunyuan's Hy-MT2 family of 'fast-thinking' multilingual machine-translation models, open-weighted on Hugging Face on Aug 20 2026. A Mixture-of-Experts model with 30B total and ~3B active parameters covering 33 language pairs plus five Chinese-dialect and minority-language pairs, with workflows for structured/delimiter-based, contextual, glossary-based, and style-guided translation. Runs a short 8,192-token context with up to 4,096 output tokens and is small enough to run locally. Tencent reports it outperforming open heavyweights such as DeepSeek-V4-Pro and Kimi K2.6 on translation quality, with even the smaller 1.8B sibling (Hy-MT2-1.8B, released the same day) beating commercial APIs from Microsoft and Doubao — vendor figures, unverified at launch. A specialized translation LLM (text in/out), included as in-scope; the smaller 1.8B and FP8 variants are not tracked separately.
Ornith-1.5-9B
AvailableThe smallest model in DeepReinforce's Ornith-1.5 family (released 2026-08-19, MIT, weights on Hugging Face): a 9B-parameter dense coding/agent model trained with the family's self-improving task-and-scaffold RL loop, and shipped with a quantized 'Ornith-1.5-9B-Mobile' build that runs on iPhone and Android. Vendor-reported, five-run-averaged figures put it at 47.0 on Terminal-Bench 2.1 and 70.6 on SWE-Bench Verified, which DeepReinforce places above larger models including Gemma 4-31B and Qwen3.6-35B-A3B. Figures are self-reported and unverified at launch.
Ornith-1.5-35B-A3B
AvailableThe mid-size model in DeepReinforce's Ornith-1.5 family (released 2026-08-19, MIT, weights on Hugging Face): a 35B-parameter Mixture-of-Experts that activates ~3B parameters per token, trained with the same self-improving task-and-scaffold generation loop as the 397B flagship. Vendor-reported, five-run-averaged figures put it at 68.5 on Terminal-Bench 2.1 and 79.0 on SWE-Bench Verified while activating only 3B parameters per token — which DeepReinforce reports as outperforming dense models of similar or larger size such as Meta's Muse-Glimmer-30B and Gemma 4-31B. Figures are self-reported and unverified at launch.
Ornith-1.5-397B
AvailableThe flagship of DeepReinforce's Ornith-1.5 family, released 2026-08-19 under the MIT license with weights on Hugging Face. A ~397B-parameter Mixture-of-Experts coding/agent model (per-token active count not disclosed) trained with a self-improving RL loop: rather than fixed human-curated tasks, the system proposes progressively harder tasks itself, generates a task-specific orchestration scaffold for each, and produces the solution rollouts used for reinforcement learning, with reward propagating across all three stages (all optimized with GRPO). Vendor-reported, five-run-averaged figures: Terminal-Bench 2.1 85.1 and DeepSWE 56.0 — which DeepReinforce puts on par with Claude Opus 4.8 (85.0 / 59.0) and ahead of GLM-5.2 and DeepSeek-V4-Flash-0731 at comparable scale — plus 92.8 GPQA Diamond and 86.6 BrowseComp. All numbers are the vendor's own and unverified by an independent harness at launch. Extends the self-scaffolding approach introduced in Ornith-1.0 (June 2026).
GLM-5.2 Turbo
AvailableA speed-optimized, hosted "Turbo" serving tier of Z.ai's GLM-5.2, surfaced Aug 17 2026 (API id glm-5.2-fast). It targets latency-sensitive coding and agent workloads at the premium fast tier and carries GLM-5.2's 1M-token context. Served through Z.ai and SCX.ai with list pricing around $1.99 input / $6.16 output per Mtok — well above standard GLM-5.2 ($0.55/$1.78), reflecting the dedicated fast-serving tier. Z.ai has not published a separate parameter count, architecture detail, or open-weight release for the Turbo variant, so size and weights are recorded undisclosed pending confirmation; the underlying GLM-5.2 base is an open-weight (MIT) ~753B MoE. Added this sweep to resolve a catalog gap prior runs flagged for lacking a solid spec source (now sourced from the LLM Gateway model page); fields beyond context, pricing, and providers are conservative and unverified.
GLM-5.3
AvailableZ.ai's 2026-08-14 coding model, pitched as the strongest open-weights coder on the market. At 743B parameters, it reuses the GLM-5.2 base model with all gains coming from expanded post-training ("more environments, more diverse tasks, more compute"), and keeps the 1M-token context and 128K output ceiling while consuming far fewer tokens per task. Live at launch through the GLM Coding Plan subscription and ZCode, with API access following. Open weights shipped 2026-08-28 on Hugging Face (zai-org/GLM-5.3) after a roughly two-week safety review that Z.ai attributed to unexpectedly strong multi-stage exploit-chaining behavior surfaced during evaluation (2,436 vulnerabilities found across 269 open-source projects). Unlike GLM-5.3-Flash's plain MIT, the flagship weights carry a bespoke "GLM-5.3 License": MIT-equivalent grants (use, modify, distribute, sell without restriction) for most users, with one divergence — companies whose aggregate revenue exceeds $10B over any consecutive 12 months must pass Z.AI's security review before using the weights or derivatives commercially as a Model-as-a-Service.
Qwen3.8-27B
AvailableThe open-weight, single-GPU sibling of Qwen3.8-Max, published by Alibaba on Hugging Face on Aug 14 2026 under Apache 2.0 — the smaller open release Alibaba had promised alongside the closed Qwen3.8-Max flagship. A 27B dense model (~28B counting the ~1B vision encoder) with 64 layers, hidden size 5,120, and a 248,320-token vocabulary. Uses a hybrid attention stack — 48 Gated DeltaNet linear-attention layers to 16 full Gated Attention layers (a 3:1 split) — for a native 262,144-token context, extendable to 1M via YaRN. Natively multimodal (text, image, and video input; text output) and ships with Multi-Token Prediction for speculative decoding. Quantized (Unsloth dynamic GGUFs) it runs in ~16-17GB of VRAM, fitting a single consumer GPU such as a 3090 or 4090 — positioned as one of the most capable local models of 2026.
Dots3-Note Preview
PreviewThe first open-weight release in the dots3 series from Dots Studio (rednote-hilab), the AI lab of Xiaohongshu (RedNote), open-weighted on Hugging Face on Aug 14 2026 under Apache-2.0. A Mixture-of-Experts model with ~280B total parameters and ~16B active per token, carrying a 512K-token context and multimodal understanding across text, vision, and audio (text output). Positioned less around benchmark-maxxing and more around combining reasoning, long-context processing, coding, and multi-step agent workflows in a compute-efficient architecture optimized for long-horizon real-world tasks. Introduces TEMPO, a reinforcement-learning method for long-horizon agents in which the model periodically checkpoints its own progress and updates its working memory mid-task. Comes from the same dots3 series whose internal build scored a perfect 42/42 at the 2026 International Mathematical Olympiad. Weights ship in BF16 and FP8; served free on OpenRouter (dots-studio/dots3-note-preview) at launch.
Gemini 3.7 Flash
AvailableGoogle DeepMind's 2026-08-13 workhorse Flash model, tuned this cycle for software engineering, agent workflows, and multi-step execution. Multimodal input (text, image, video, audio, PDF) over a 1,048,576-token context with a 65,536-token output ceiling; supports function calling, search-as-a-tool, and computer use. Introductory pricing is $0.75 / $3.75 per 1M input/output tokens through 2026-12-31, rising to $1.50 / $7.50 thereafter. Google-reported gains over Gemini 3.6 Flash: DeepSWE v1.1 65.3% (vs 49.0%), AutomationBench 30.4% (vs 17.0%), WebDev Arena Elo 1588 (vs 1538).
DeepSeek-V4-Pro-0813
AvailableThe dated GA build behind DeepSeek's 'deepseek-v4-pro' API id, pinned on the official pricing table with an OpenRouter listing dated Aug 12 2026 — the Pro-tier counterpart to the way V4-Flash graduated as DeepSeek-V4-Flash-0731. It is a quiet version pin (no separate launch post or benchmark card), keeping the V4-Pro architecture and the 1M-token context / ~384K max-output window. List pricing is cache-heavy: $0.435 input cache-miss / $0.003625 cache-hit / $0.87 output per Mtok, with concurrency 500 (vs Flash's 2500); DeepSeek warns a significant, undated API price increase is coming. Thinking is on by default at effort 'high' (requested medium/xhigh both collapse to high). Architecture figures (1.6T total / 49B active MoE, hybrid long-context attention) are carried over from the April 2026 V4-Pro preview and are not independently reconfirmed for the 0813 build; Hugging Face still hosts only the April preview weights (MIT), with no confirmed separate 0813 open-weight repo, so this row is recorded as proprietary/API-only.
Grok 4.6
AvailablexAI's frontier update to Grok 4.5, released 2026-08-12. Built on the same reported 1.5-trillion-parameter V9 foundation as Grok 4.5, with gains coming from upgraded supervised fine-tuning and reinforcement learning rather than a new base. Supports a 500K-token context and configurable reasoning effort (low, medium, high default, xhigh). List pricing is $2 / $6 per 1M input/output tokens below 200K prompt tokens ($0.50 cached input), rising to $4 / $12 above 200K.
North Micro Vision Instruct
AvailableCohere's compact document-focused vision-language model, published Aug 12 2026 under Apache 2.0. A 2.4B-parameter VLM combining a custom 400M native-resolution vision encoder, a 2B language model on the Command A+ architecture, and a projector; it preserves aspect ratio for images up to 1654x2339px (an A4 page at 200 dpi). Multilingual visual understanding across documents, charts, and natural images, text output. Vendor-reported: 0.921 DocVQA and 0.808 ChartQA on document tasks, 0.732 RefCOCO on visual grounding, and 0.687 MMBench on general VQA; text-only capability lags larger models. Part of Cohere's North product family alongside North Mini Code. Open weights on Hugging Face.
LFM2.5-VL-3B
AvailableLiquid AI's edge vision-language model, released Aug 12 2026 — a 3.1B-parameter VLM built on the LFM2.5-2.6B text base with an integrated SigLIP2 400M NaFlex vision encoder. Accepts text, images, and video frames and returns text, tuned for on-device screen understanding, visual grounding, and tool calling. Vendor-reported: 80.7 average on ScreenSpot-v2 screen understanding, 87.9 P@1 on RefCOCO grounding, 59.5 on ToolSandbox function calling, 81.0 on MMBench, and 69.4 averaged across 28 benchmarks. Runs on-device at ~228 tok/s on an Apple M5 Max and ~116 tok/s on an AMD Ryzen AI Max+ 395; supported in llama.cpp, MLX, vLLM, SGLang, and ONNX. Open weights on Hugging Face.
Nemotron 3.5 Lightning
AvailableNVIDIA's efficient open Mixture-of-Experts model, released Aug 11 2026 for long-running agents. A hybrid Mamba-2 + MoE + Attention design with ~31.6B total and ~3.6B active parameters and a 1M-token context, shipped alongside the NeMo Switchyard model router. NVIDIA reports performance comparable to gpt-oss-120b at roughly a quarter of the total parameters, up to 4x the output speed of similar-sized models, and 10,000 tasks completed ~30% faster than Qwen3.6-35B at similar accuracy. Vendor-reported BF16 figures: SWE-bench Verified 51.56, GPQA Diamond 75.44, MMLU Pro 81.94, PinchBench 85.37 — self-reported and unverified by an independent harness at launch. Ships under the permissive OpenMDW-1.1 license with weights, training data, and recipes released, and is available on Hugging Face, ModelScope, OpenRouter, and build.nvidia.com as an NVIDIA NIM microservice.
Namazu
AvailableSakana AI's Japanese-specialized reasoning and agent model, opened as an API on Aug 11 2026. Rather than pretraining from scratch, Sakana builds on Moonshot's open-weight Kimi K2.6 and adds training for Japanese language and business contexts — the lab's 'sovereign AI by tuning other labs' strategy. Handles Japanese instruction following, business writing, math, coding, research, and multi-step agentic workflows; accepts text, images, and files such as PDFs and returns text, over a 262,144-token context with up to 65,536 output tokens, and ships native web-search and code-execution tools through a unified OpenAI-compatible API. Priced on OpenRouter at $0.95 / $4.00 per 1M input/output tokens. Not available in the EU/EEA, the UK, or Switzerland. Served via API; weights not separately published (base Kimi K2.6 is open-weight).
MAI-Code-1.1-Flash
AvailableMicrosoft's upgraded inference-efficient agentic coding model, released Aug 11 2026 and rolling out in production across GitHub Copilot (CLI and VS Code). A post-training-focused successor to MAI-Code-1-Flash, tuned through hundreds of thousands of reinforcement-learning environments built from real GitHub Copilot usage rather than raw scaling. Microsoft reports a +22% gain on Terminal-Bench 2.1 for CLI tasks and +15% on .NET workloads over 1.0, alongside 25% greater token efficiency (25% fewer tokens per task, streaming 25% faster) — at roughly a quarter of the cost of MAI-Code-1.0. Text/code only; size and context window undisclosed, weights not released.
Muse Glimmer
AvailableMeta's first open-weight agentic model, released Aug 10, 2026 under an Apache 2.0 license — Meta's return to open weights after the closed Muse Spark line. A ~30B-parameter dense causal transformer (about 29.6B parameters across 52 layers) paired with a ~1.8B ViT-G/14 perception encoder, so it accepts interleaved text and images and returns text across more than 100 languages. Carries a 131,072-token context, a 202,048-token vocabulary, and a Jan 4, 2026 knowledge cutoff. Uses grouped-query attention (32 query heads, 2 KV heads) in a local/local/local/global pattern with a 2,048-token sliding window, plus speculative decoding for throughput. Quantized to roughly 4-bit it fits inside a ~24GB memory envelope, running on a single consumer GPU or an Apple-silicon Mac — the model is tuned for on-device agent workloads. Shipped the same week as the closed-weight Muse Spark 1.2 coding flagship; Meta's first agentic model to ship with both open weights and a permissive commercial-use license.
Solar Pro 4
AvailableThe flagship of South Korean lab Upstage's in-house Solar family, surfaced on developer platforms around Aug 10 2026 and formally unveiled Aug 14 2026. An agent-first reasoning model pitched as a reliable workhorse — 'save frontier models for frontier problems' — with strong office-productivity, document-intensive, and coding performance rather than raw frontier scores. Carries a 524K-token context and up to 131K output tokens. Upstage reports large gains over Solar Pro 3: ~71 on the AA-LCR long-document benchmark (about 2.3x its predecessor) and 23 on Tau3-Banking tool-use (a 2.6x jump); Artificial Analysis scores it 42 on its Intelligence Index. API pricing is $0.30 / $1.20 per 1M input/output tokens, with a launch promo cutting that 90% to $0.03 / $0.12. Upstage is the first Korean company registered as an official model provider alongside OpenAI, Anthropic, Google, and NVIDIA. Served via Upstage's API and OpenRouter; weights were not published at launch, so treated as proprietary.
GPT-5.6-Cyber
PreviewOpenAI's cybersecurity-specialized model, launched Aug 10 2026 as the successor to the limited-preview GPT-5.5-Cyber. Built on GPT-5.6-Sol and tuned for authorized security workflows — vulnerability discovery, exploit-chain development, privilege escalation, authentication-bypass analysis, patch validation, and red-teaming — with a deliberately lower refusal rate on dual-use security tasks than general GPT-5.6. OpenAI reports a 95% completion rate on advanced exploit-chain prompts (vs ~1.5% for GPT-5.6-Sol) and that the model surfaced a high-severity V8 JavaScript-engine vulnerability plus flaws in unnamed mobile-OS, database, and kernel targets. Not generally available: access is gated to vetted defenders through the expanded Daybreak program's new "Daybreak Red" tier (partners include Accenture, IBM, Palo Alto Networks, CrowdStrike, Fortinet, Akamai, and Cloudflare). Size undisclosed; weights not released.
LFM2.5-2.6B
AvailableLiquid AI's on-device agentic model, released Aug 6 2026 (surfaced on hosted platforms ~Aug 11) — a 2.69B-parameter dense model, distinct from the LFM2.5-8B-A1B MoE. Uses Liquid's hybrid stack across 30 layers: 22 double-gated short-convolution blocks plus 8 grouped-query-attention blocks, with a 128K-token vocabulary and a 131,072-token context, pre-trained on ~34T tokens across 16 languages. Text-only, built to plan, call tools, and complete multi-step tasks entirely on-device (phone, laptop, PC, or robot) so data never leaves the device and the marginal cost per run is near zero; Liquid reports tool-use and instruction-following competitive with models ~4x its size (e.g. leading Qwen3.5-9B on ToolSandbox, Multi-IF, and IFStruct), while explicitly not recommending it for agentic coding or knowledge-heavy work. Decodes at ~220 tok/s on an Apple M5 Max in under 2.5GB and ~30 tok/s on a phone. Open weights (base + post-trained) on Hugging Face under the LFM Open License (lfm1.0), shipped day-one in GGUF, MLX, and ONNX.
Ling-3.0-tiny
AvailableThe smallest member of Ant Group inclusionAI's Ling 3.0 family, open-weighted on Hugging Face on Aug 6 2026 under the MIT license — distinct from the (API-only at launch) Ling-3.0-flash. A sparse Mixture-of-Experts model with 7.9B total parameters and only ~1.3B active per token: 128 routed experts with 8 routed plus 1 shared expert active per token, using the same 3:1 alternating stack of Kimi Delta Attention (KDA, linear) and Multi-head Latent Attention (MLA) layers as the rest of the family, for a 262,144-token (256K) context. Pitched as a highly economical on-device agent/reasoning model; weights are provided in BF16, FP8, and INT4 for a wide range of hardware. Vendor benchmark figures are unverified at launch.
Muse Spark 1.2
AvailableMeta's flagship coding model, released 2026-08-05 and purpose-built for complex software engineering: debugging sprawling codebases, validating changes across thousands of files, and multi-step reasoning, with deep integration into persistent asynchronous background agents. Accepts text, image, video, audio, and PDF input over a 1M-token context and returns text. Standard-tier API pricing is $1.25 / $4.25 per 1M input/output tokens ($0.15 cached input); a new muse-spark-1.2-contributor tier drops to $0.10 / $0.20 in exchange for permission to train future Meta models on your prompts and completions. Shipped alongside Muse Code, a terminal-based coding agent powered by the model. Meta has signaled open weights are coming.
Qwen3.8-Max
AvailableAlibaba's largest model to date and the flagship of the Qwen3.8 line — a 2.4-trillion-parameter sparse Mixture-of-Experts (~95B active per query) that Alibaba positions just behind Anthropic's Fable 5 on overall performance. Previewed 2026-07-19 at the World AI Conference in Shanghai, it went generally available on 2026-08-03 with a published benchmark table, standard API access, and firm per-token pricing ($2 / $6 per 1M input/output tokens). Fully multimodal (text, image, video input) over a 1M-token context. Alibaba also committed to shipping open weights for both Qwen3.8-Max and a smaller Qwen3.8-27B checkpoint.
DeepSeek-V4-Flash-0731
AvailableThe production release of DeepSeek's V4-Flash tier — the April V4-Flash preview retrained on a substantially improved post-training pipeline targeting coding, agents, reasoning, and tool use, with no change to the base architecture. Retains 284B total / 13B active parameters (MoE) and the 1M-token context window. DeepSeek reports the 0731 build scoring higher than its own larger V4-Pro-Preview on all nine agent and coding benchmarks it published — a vendor-reported result, with independent replication still limited at launch. Weights released on Hugging Face under the MIT license; API pricing held at $0.14 / $0.28 per Mtok. The upgrade is silent for existing callers: same endpoint, same key, same deepseek-v4-flash model name, zero migration cost.
Qwen3.7-Flash
AvailableCost-optimized multimodal member of the Qwen3.7 line — a vision-language reasoning model with a 1M-token context, tuned for high-volume multimodal agent workloads (visual coding, screen perception, browser/computer use, search) where cost matters more than peak intelligence. Launched quietly on Jul 27, 2026 as an OpenRouter/API listing at $0.03/$0.13 per 1M tokens, making it the cheapest 1M-context multimodal model available at release. Closed-weights and API-only; Alibaba published no technical report, benchmark suite, or architecture details, though community speculation points to a small sparse-MoE design.
Claude Opus 5
AvailableAnthropic's flagship Opus model, released July 24, 2026 — positioned as the go-to model for most knowledge work and automation, approaching the capability of Claude Fable 5 in many categories at roughly half the price. Built for demanding reasoning, autonomous coding, software development, and long-horizon agentic work. Introduces a five-level 'effort' dial exposed to developers on the Claude API and Platform, letting them trade compute and tokens for capability — at lower effort it preserves much of its performance while using fewer tokens and costing less to run. 1M-token context window (available at standard token pricing, not a separate long-context surcharge) with up to 128K output tokens; text, vision, and code. Standard API pricing $5/$25 per Mtok (the same as its predecessor Opus 4.8), plus a Fast mode at $10/$50. Anthropic describes it as its most aligned Opus model and the least susceptible to being tricked into misuse. Becomes the default model for Claude Max subscribers and is available across Anthropic's paid plans; scores 61 on the Artificial Analysis Intelligence Index. Closed weights; architecture, parameter count, and training compute undisclosed.
Ling-3.0-flash
AvailableAnt Group's efficiency-focused Mixture-of-Experts model, released July 23 2026 by its inclusionAI lab: 124B total parameters activating only ~5.1B per token (1/64 expert activation). Ant claims it matches or beats the company's own ~1T-parameter Ling-2.6 flagship on most benchmarks it shows, at 1/8 the total and 1/12 the active parameters — a vendor claim with no public benchmark table or independent audit at launch, so treat it as unverified. Built for production-scale agents (MCP tool use, multi-agent coordination) rather than chat, with both thinking and non-thinking modes. Architecture is a native hybrid-linear attention stack interleaving Kimi Delta Attention (KDA) and Multi-head Latent Attention (MLA) at a reported 5:1 ratio, giving an economical 262,144-token (256K) context, with 1M cited as the scaling target. Ant docs claim peak inference up to 1,000 tokens/s and <100ms time-to-first-token on its own stack. Announced as open-weight under Apache 2.0, but as of July 24 no weights or model card were posted to the inclusionAI Hugging Face org — so the license and open-weight status are announced but unconfirmed (weights not yet downloadable; not self-hostable today). Usable now only via hosted API — free on OpenRouter (as inclusionai/ling-3.0-flash:free, hosted by Novita) and Vercel's AI Gateway through August 3 2026; no post-promo per-token price published at launch.
Gemini 3.5 Flash Cyber
PreviewA specialized, highly efficient cybersecurity model built on Gemini 3.5 Flash and fine-tuned to find and fix software vulnerabilities at a lower price per token than larger models. Deployed inside Google's CodeMender agent, where multiple 3.5 Flash Cyber agents collaborate to produce a single combined report, reaching competitive frontier performance on the CyberGym benchmark. Given its dual-use nature, it is not generally available: access is limited to governments and trusted partners via CodeMender as part of a limited-access pilot program.
Gemini 3.5 Flash-Lite
AvailableGoogle DeepMind's fastest and most cost-effective 3.5-class model, released July 21, 2026 for low-latency and high-throughput agentic workloads like agentic search and document processing. Runs at ~350 output tokens/s (Artificial Analysis) with configurable thinking levels and built-in computer use, priced at $0.30 / $2.50 per 1M input/output tokens. Multimodal over a 1M-token context and a large step up on 3.1 Flash-Lite: Terminal-Bench 2.1 54% (vs 31%), GDM-MRCR v2 72.2% (vs 60.1%), GDPval-AA v2 1140 (vs 642); on several agentic and coding evals it even surpasses 3 Flash (SWE-Bench Pro 54.2% vs 49.6%, OSWorld-Verified 74.0% vs 65.1%). Available in the Gemini API (AI Studio, Android Studio), Gemini Enterprise, the Gemini app, and rolling out in Google Search.
Gemini 3.6 Flash
AvailableGoogle DeepMind's July 2026 workhorse Flash model, built for scaling agentic workflows. Multimodal over a 1M-token context, it improves on Gemini 3.5 Flash in coding, knowledge work, and computer use while cutting output-token usage ~17% (up to 65% on some benchmarks like DeepSWE) and taking fewer reasoning steps and tool calls. Ships at a lower price than 3.5 Flash ($1.50 / $7.50 per 1M input/output tokens). Google-reported gains: DeepSWE 49% (vs 37%), MLE-Bench 63.9% (vs 49.7%), OSWorld-Verified 83.0% (vs 78.4%), GDPval-AA v2 1421 (vs 1349); knowledge cutoff advances to March 2026. Computer use is a built-in client-side tool. Available in the Gemini API (AI Studio, Android Studio, Antigravity), Gemini Enterprise, and the Gemini app.
Laguna S 2.1
AvailablePoolside's open-weight agentic-coding model and a scale-up of the Laguna XS family (same pre-training data as XS 2.1): a 118B-total / 8B-active Mixture-of-Experts that activates only ~6.8% of its parameters per token, giving larger-model behavior while staying cheap to serve, with a 1M-token context in both thinking and no-thinking modes. Pitched by Poolside as 'the West's most capable open-weight model' — the claim is about its weight class, not the outright frontier. Two modes (off / max, max default; the model sets its own test-time compute budget). Vendor-reported: Terminal-Bench 2.1 70.2% and SWE-bench Multilingual 78.5% (tops the published open disclosed-size table), plus SWE-bench Pro 59.4%, DeepSWE v1.1 40.4%, SWE Atlas 46.2%, Toolathlon Verified 49.7% — matching or beating models several times its size, though closed frontier models still lead outright. Trained in under nine weeks on 4,096 NVIDIA H200 GPUs (pre-training began 22 May 2026); first Poolside model with RL in FP8. Knowledge cutoff November 2025. Weights on Hugging Face under the permissive OpenMDW-1.1 license in BF16/FP8/INT4/NVFP4 with GGUF/MLX conversions and DFlash draft models; at 4-bit it runs on a single NVIDIA DGX Spark. Day-one support for vLLM, SGLang, and Ollama; hosted free at 256K context via OpenRouter and paid at the full 1M context ($0.10 / $0.20 / $0.01 per 1M input / output / cache-read tokens), also on Baseten, Kilo, Prime Intellect, and ZML.
Kimi K3
AvailableMoonshot's flagship open-weight agentic model and the largest open model released to date: a 2.8T-parameter MoE (896 experts, 16 active per token) using Kimi Delta Attention and Attention Residuals, with native multimodal input and a 1M-token context. Launched via API on Jul 16, 2026 at $3/$15 per Mtok (cached input $0.30); full open weights published to Hugging Face on Jul 26, 2026 — a day ahead of the announced Jul 27 target — under a Modified MIT license, making it freely downloadable and self-hostable.
Inkling
AvailableThinking Machines Lab's first model and the leading U.S. open-weights release: a natively multimodal Mixture-of-Experts with 975B total / 41B active parameters that reasons across text, image, and audio inputs and emits text. Pretrained on ~45T tokens; served with a 1M-token context from the Hugging Face weights (256K on the hosted Tinker API). Apache-2.0 licensed (BF16 + NVFP4 checkpoints on Hugging Face), built for developers fine-tuning on proprietary data — coding assistants, agents/tool use, chatbots, and RAG — with an explicit low-cost and censorship-resistance focus. Debuted at 41 on the Artificial Analysis Intelligence Index. Hosted pricing (256K) $3.74/$9.36 per Mtok reflects a limited-time 50% launch discount.
KAT-Coder-Air V2.5
AvailableThe efficient, ~32B variant of Kwaipilot's KAT-Coder V2.5, optimized through multi-stage training (supervised fine-tuning plus reinforcement learning). Shares KAT-Coder-Pro's 256K-token context and 80K max output, with function calling and tool use for agentic coding, at roughly a fifth of Pro's cost: $0.15 / $0.60 per million input/output tokens. Surfaced on release trackers on July 14 2026.
KAT-Coder-Pro V2.5
AvailableKwaipilot's flagship agentic coding model, from the KAT (Kwaipilot Agentic Tuning) series at Kuaishou. A Mixture-of-Experts model with ~72B active parameters, trained through large-scale agentic reinforcement learning in reconstructed, verifiable repository environments. Supports function calling, tool use, structured JSON output, and prompt caching, with a 256K-token context and up to 80K output tokens. Served via API (StreamLake / Atlas Cloud / OpenRouter) at $0.74 / $2.96 per million input/output tokens. Surfaced on release trackers on July 14 2026.
GPT-5.6
AvailableOpenAI's GPT-5.6 series umbrella row. Officially previewed June 26, 2026 as three durable capability tiers — Sol (flagship), Terra (balanced, for everyday work), and Luna (fast and affordable) — introduced with a new `max` reasoning effort for deeper reasoning and an `ultra` mode that leverages subagents to accelerate complex work. In the GPT-5.6 naming system the number marks the generation while Sol/Terra/Luna are tiers that can advance on their own cadence. Initially a limited preview via the API and Codex for a small group of vetted partners (after U.S. government review), with general availability across ChatGPT, Codex, and the API planned in the following weeks.
GPT-5.6 Terra
AvailableBalanced, everyday-work tier of OpenAI's GPT-5.6 series, officially previewed June 26, 2026. OpenAI positions Terra as competitive with GPT-5.5 while being roughly 2x cheaper. Shares the series' new `max` reasoning effort and `ultra` subagent mode and OpenAI's GPT-5.6 safety stack. Begins as a limited preview via the API and Codex for vetted partners after U.S. government review, with broad availability planned in the following weeks. Priced at $2.50 / $15 per million input/output tokens.
GPT-5.6 Sol
AvailableFlagship tier of OpenAI's GPT-5.6 series, officially previewed June 26, 2026 — OpenAI's strongest model to date. Adds a new `max` reasoning effort for the deepest reasoning and an `ultra` mode that uses subagents to accelerate complex work. Sets a new state of the art on Terminal-Bench 2.1 (command-line, agentic coding) and shows broad gains in long-horizon biology (GeneBench v1) and cybersecurity (ExploitBench, ExploitGym), paired with OpenAI's most robust safety stack and a phased release. Begins as a limited preview via the API and Codex for a small group of vetted partners after U.S. government review, with general availability planned in the following weeks; also launching on Cerebras at up to 750 tokens/sec in July. Priced at $5 / $30 per million input/output tokens.
GPT-5.6 Luna
AvailableFast, low-cost tier of OpenAI's GPT-5.6 series, officially previewed June 26, 2026 — the most affordable model in the family, bringing strong capability at OpenAI's lowest cost. Shares the series' capability-tier naming and GPT-5.6 safety stack. Begins as a limited preview via the API and Codex for vetted partners after U.S. government review, with broader availability planned in the following weeks. Priced at $1 / $6 per million input/output tokens.
Muse Spark 1.1
PreviewMeta Superintelligence Labs' first paid model, released July 9, 2026 in US public preview on the Meta Model API. A natively multimodal reasoning model (text, image, video, PDF, and audio input; text output) with explicit chain-of-thought reasoning and a 1M-token context window that the model actively compacts. Positioned for agentic work and coding — tool use, multi-step workflow coordination, and long-horizon autonomous tasks — and pitched by Meta as roughly a quarter of the price of comparable Anthropic and OpenAI models at $1.25 in / $4.25 out per Mtok (with $20 in free credits per new API account). Marks the first time Meta has charged businesses for one of its models, a departure from the open-weight Llama strategy. Closed weights, undisclosed size. Vendor and third-party benchmarks place it around the Opus 4.8 / GPT-5.5 tier — strongest as an agent/workflow model and in tool-augmented reasoning, competitive but not dominant on coding and multimodal tasks.
Grok 4.5
AvailableSpaceXAI's "Opus-class" agentic flagship, released July 8, 2026 — the first Grok model trained jointly with the coding startup Cursor (Anysphere), on trillions of tokens of real Cursor usage data plus STEM tasks, research papers, and other knowledge work. Targets software engineering, agentic tasks, and knowledge work, with explicit strength in legal and finance use cases (SpaceXAI claims the top spot on the Harvey Legal Agent Benchmark). A Mixture-of-Experts model reported to be built on a ~1.5-trillion-parameter "V9" foundation. Vendor benchmarks are mixed against Anthropic's Opus 4.8 — ahead on DeepSWE 1.0 and Terminal-Bench 2.1, behind on DeepSWE 1.1 and SWE-bench Pro — but markedly more token-efficient (~15,900 output tokens on SWE-bench Pro tasks, ~4.2x fewer than Opus 4.8) and served at ~80 tokens/sec. Base pricing $2/$6 per Mtok; Cursor lists a faster variant at $4/$18. Available in Grok Build (default model), Cursor (all plans), and the SpaceXAI console; initially unavailable in the EU (expected mid-July). xAI was absorbed by SpaceX in Feb 2026 and rebranded SpaceXAI.
Hunyuan Hy3
AvailableThe general-availability release of Tencent's third-generation Hunyuan (Hunyuan 3.0), officially launched and open-sourced on July 6, 2026 after April's "Hy3 preview". A 295B-total / 21B-active Transformer MoE with an additional 3.8B multi-token-prediction (MTP) layer and a 256K-token context, offering three selectable inference modes that blend fast and slow thinking. Positioned as a leading open model for its size and cost efficiency, with standout results in coding, search, and scientific reasoning: Tencent reports it rivals flagship open models such as GLM-5.2 and DeepSeek-V4 (at 2-5x the active parameters) and matches or surpasses GPT-5.5 on several science benchmarks. Vendor-reported scores include 78.0 on SWE-bench Verified and 57.9 on SWE-bench Pro. Now Apache-2.0 licensed (the preview used Tencent's community license), with weights on Hugging Face (tencent/Hy3) and ModelScope and a free two-week API route on OpenRouter (tencent/hy3:free) through July 21, 2026. Deeply integrated into WeChat and Tencent's core products.
Nemotron-Labs-3-Puzzle-75B-A9B
AvailableA deployment-optimized open-weight model from NVIDIA, released July 6, 2026 — a compressed variant of Nemotron-3-Super-120B-A12B produced with "Iterative Puzzle", a post-training compression framework that jointly prunes MoE experts, active-parameter budget, and Mamba state to boost inference efficiency while preserving accuracy. Reduces the parent from 120.7B total / 12.8B active to 75.3B total / 9.3B active, keeping the hybrid Mamba-Transformer LatentMoE architecture with Multi-Token Prediction. Delivers ~2x higher server throughput than Nemotron-3-Super on a single 8xB200 node at matched user throughput and raises sustainable 1M-token single-H100 concurrency from 1 to 8 requests. Targets collaborative agents, chatbots, RAG, complex instruction-following, and long-context reasoning across English, code, and six other languages. Shipped in BF16, FP8, and NVFP4 variants under the OpenMDW-1.1 license.
Mistral frontier open-weight MoE (unnamed)
RumoredRumored / early-access: Mistral AI has confirmed a new open-weight Mixture-of-Experts family — described by CEO Arthur Mensch as "fat but sparse" — aimed at closing the gap with frontier open-weight releases, with early access beginning in July 2026. Mensch confirmed the intent but disclosed almost nothing concrete: no parameter count, no benchmarks, no license terms, and no release date. Tracked as rumored until weights or an official product page land.
Laguna XS 2.1
AvailablePoolside's open-weight small coding model: a 33B-total / 3B-active Mixture-of-Experts built for agentic coding and long-horizon work on a local machine, served at 256K context. An upgraded XS.2 (same architecture) that lifts SWE-bench Multilingual by 5.4 points to 63.1% and improves terminal-style tasks. Ships with open-weighted DFlash speculator (draft) models for each checkpoint that roughly double local tokens/sec, plus BF16/FP8/INT4/NVFP4 quantized checkpoints; supported in vLLM, SGLang, TensorRT-LLM, HF transformers, and Ollama (llama.cpp coming). Newly relicensed under the fully permissive OpenMDW-1.1. Available free on Hugging Face and via a free OpenRouter tier, with paid API pricing of $0.10 / $0.20 / $0.05 per 1M input / output / cache-read tokens. Its predecessor Laguna XS.2 sunsets on Poolside's API one week after launch.
Claude Sonnet 5
AvailableAnthropic's most agentic Sonnet model yet, with performance approaching Claude Opus 4.8 at a lower price. Built for coding, tool use (browsers and terminals), and autonomous multi-step agentic work, with selectable effort levels up to xhigh. A substantial upgrade over its predecessor Sonnet 4.6 on reasoning, tool use, coding, and knowledge work, with gains shown on SWE-bench, OSWorld-Verified, BrowseComp, and Humanity's Last Exam. The default model on the Free and Pro plans and available to Max, Team, and Enterprise users; in Claude Code and via the Claude API as claude-sonnet-5. Uses an updated tokenizer (same approach as Opus 4.7). Ships with real-time cyber safeguards enabled by default. Introductory API pricing of $2/$10 per Mtok through Aug 31, 2026, then standard $3/$15.
LongCat-2.0
AvailableMeituan's open-weight flagship: a 1.6-trillion-parameter Mixture-of-Experts model (~48B active per token, dynamically routed between ~33B and ~56B) with a 1M-token context, built for agentic coding. Notable as the largest Chinese model trained — for both pre-training and inference — entirely on a ~50,000-card cluster of domestic Chinese AI chips (Meituan's use of the Huawei Collective Communication Library points to Huawei Ascend hardware), and the first trillion-parameter model Meituan claims completed full-process training on home-grown compute. Vendor-reported software-engineering results: 59.5 on SWE-bench Pro (ahead of GPT-5.5's 58.6), 70.8 on Terminal-Bench 2.1, and 77.3 on SWE-bench Multilingual, with overall quality positioned as comparable to Gemini 3.1 Pro (self-reported, not yet independently verified). Open-sourced under the MIT license with weights on Hugging Face and GitHub; follows LongCat-Flash (560B, Sep 2025) and the multimodal LongCat-Next (Mar 2026).
Base1
AvailableBase44's first in-house model, a general-purpose agent for 'vibe coding' — turning natural-language prompts into working web apps. Fine-tuned on top of an open-source foundation model (rather than trained from scratch) and specialized on a dataset generated from tens of millions of real interactions across the Base44 platform, it holds a conversation, writes code, and handles multi-turn requests, tool use, and backend operations while being faster and cheaper to run than the frontier models it sits beside. Now selectable in Base44's model picker alongside GPT-5.5 and Claude Opus 4.8. Rolled out from June 29 2026; platform-only, with no separately published API pricing or context figure. Base44 (founded by Maor Shlomo) was acquired by Wix in 2025.
Seed 2.1 Turbo
AvailableThe low-cost, low-latency tier of ByteDance's Seed 2.1 family (served as Doubao-Seed-2.1-turbo), built for large-scale production with full features and performance ByteDance positions as comparable to Seed 2.1 Pro. Shares the family's coding, long-chain agent, and multimodal-understanding focus and 256K-token context, priced for high-volume online calls. Proprietary; available via Doubao and Volcano Engine. List price ¥3 / ¥15 per million input/output tokens — roughly half the Pro tier.
Seed 2.1 Pro
AvailableByteDance's flagship next-generation agent model (served as Doubao-Seed-2.1-pro), built for the "coding and agent era." A deep-thinking model tuned for strong demand understanding, long-horizon planning, and continuous self-repair across complex coding, long-chain agents, and multi-step engineering delivery, with a 256K-token context. ByteDance reports its core coding, agent, and multimodal capabilities are comparable to GPT-5.5, with the highest score on GDPVal, top-tier results on the Agents' Last Exam, the highest score on MobileWorld, and SOTA results across several visual and video-understanding benchmarks (CharXiv-RQ, MeasureBench, TVBench, TOMATO). Proprietary; available via Doubao and Volcano Engine. List price ¥6 / ¥30 per million input/output tokens.
Fugu Ultra
AvailableSakana AI's frontier-class orchestration model: a single ~7B language model trained to coordinate a swappable pool of external frontier LLMs (model selection, delegation, verification, and synthesis happen internally), exposed behind one OpenAI-compatible API. The Ultra tier is tuned for maximum answer quality on hard, multi-step problems and coordinates a deeper, fixed pool of expert agents. Sakana reports it stands shoulder-to-shoulder with Anthropic's Fable 5 and Mythos Preview across coding, science, and reasoning benchmarks while routing around single-vendor/export-control risk. Current model ID fugu-ultra-20260615; vendor-reported scores include 95.5 GPQA-D, 73.7 SWE-Bench Pro, 93.2 LiveCodeBench, and 50.0 Humanity's Last Exam.
Fugu
AvailableSakana AI's orchestration model: a single language model that delivers a full multi-agent system behind one OpenAI-compatible API, dynamically routing tasks across a swappable pool of frontier LLMs (including recursive calls to itself). The base tier balances strong performance with low latency as an everyday default for coding, code review, and chat, and lets teams opt specific agents out of the pool for data/privacy/compliance needs. Released alongside Fugu Ultra on June 22, 2026; vendor-reported scores include 95.5 GPQA-D, 92.9 LiveCodeBench, and category-leading SciCode and long-context results.
Z.ai Fable-class model
RumoredA speculative Z.ai frontier model tracked after Z.ai founder Jie Tang responded to Elon Musk's prediction of a Chinese Fable 5-class model by saying it would not take that long. Name, architecture, weights, and launch timing remain unconfirmed.
Kimi K2.7 Code
AvailableMoonshot's open coding-focused agentic model built on K2.6, with native vision/video input, forced thinking mode, and stronger long-horizon software-engineering performance.
GLM-5.2
AvailableZ.ai's latest open flagship for long-horizon coding, agentic engineering, and million-token workflows, adding IndexShare sparse-attention reuse over GLM-5.1.
MiniMax-M3
AvailableNative multimodal MiniMax model with a one-million-token context, sparse attention, and agentic coding/cowork positioning.
DiffusionGemma 26B-A4B
AvailableAn open-weight text-diffusion model built on the Gemma 4 26B-A4B MoE backbone (25.2B total / 3.8B active). Denoises text in parallel 256-token blocks for up to ~4x faster generation (1,000+ tok/s on an H100), with a 256K context and text, image, and video input. Apache-2.0.
Claude Fable 5
AvailableThe public, guardrailed sibling of Claude Mythos 5 and Anthropic's most capable widely released model, built for long-horizon agentic work, coding, vision, and knowledge workflows. Launched June 9, 2026 across the Claude API, AWS, and Microsoft Foundry, then suspended three days later under a U.S. export-control directive. Anthropic says those controls were lifted June 30 and Fable 5 was restored globally on July 1 across Claude Platform, Claude.ai, Claude Code, and Claude Cowork, with cloud partner access being re-enabled. Its safeguards route flagged cybersecurity, biology/chemistry, and distillation requests to Claude Opus 4.8.
North Mini Code 1.0
AvailableCohere's first developer-focused model and the first in its North family of code agents. A 30B-total / 3B-active MoE for agentic coding with a 256K context and up to 64K output, sized to run locally for enterprise coding agents. Apache-2.0.
Unisound U2
AvailableUnisound's new-generation, general-purpose "native agentic" large model, built for task execution: it can autonomously decompose and advance complex real-world workflows of 100+ steps rather than single-turn Q&A. Unisound frames it around "high intelligence density x high token value" and reports ~25% lower thinking-token consumption. Available via the Unisound Token Hub.
Nemotron 3 Ultra 550B-A55B
AvailableNVIDIA's largest Nemotron 3 open-weight hybrid Mamba-Transformer MoE, tuned for agentic reasoning, coding, planning, and tool calling.
Qwen3.7-Plus
AvailableMultimodal sibling of Qwen3.7-Max that adds vision input and GUI grounding for screen perception, browser automation, and hybrid GUI+CLI agent workflows. 1M-token context; closed-weights and API-only. Previewed at the May 2026 Alibaba Cloud Summit and reached general availability in June 2026 at a low price point ($0.40/$1.60 per 1M tokens).
Gemma 4 12B
AvailableA dense 12B member of the Gemma 4 family with a unified, encoder-free multimodal architecture: vision and audio are projected straight into the LLM backbone. First medium-size Gemma to natively ingest audio; runs on a 16GB laptop. 256K context, Apache-2.0.
MAI-Thinking-1
AvailableMicrosoft's first in-house frontier reasoning model, unveiled at Build 2026. A sparse MoE (~35B active, 256K context) trained entirely on commercially licensed data without third-party distillation. Microsoft reports 97.0% on AIME 2025 and coding parity with Claude Opus 4.6 on SWE-bench Pro.
MAI-Code-1-Flash
AvailableAn inference-efficient agentic coding model from Microsoft (~5B active parameters), trained from the ground up on clean, traceable, enterprise-grade data without third-party distillation. Rolling out in GitHub Copilot; Microsoft reports a +16-point SWE-bench Pro lead over Claude Haiku 4.5.
Nex-N2-Pro
AvailableNex AGI's open-weight agentic flagship, post-trained on Qwen3.5-397B-A17B (397B total / ~17B active MoE) by the Shanghai Innovation Institute-led Nex alliance. Built around an "Agentic Thinking" framework for long-horizon coding, deep research, tool calling, and terminal execution; accepts text and image input and emits text with explicit reasoning traces and function calling. Apache-2.0, ~262K context. Nex reports parity with GPT-5.5 and Claude Opus 4.7 on several agentic and coding evals. A smaller Nex-N2-mini (35B/3B-active) was announced but is not yet open-sourced.
Step-3.7-Flash
AvailableStepFun's high-efficiency multimodal sparse-MoE successor to Step-3.5-Flash: a ~196B-total / ~11B-active vision-language model with native image and video understanding, a 256K context, and selectable reasoning tiers (high/medium/low). Tuned for coding agents and search workflows.
Claude Opus 4.8
AvailableAnthropic's most capable model, with strengthened agentic and long-running task performance.
LFM2.5-8B-A1B
AvailableLiquid AI's on-device Mixture-of-Experts model: 8.3B total parameters with only ~1.5B active per forward pass (32 experts, 4 active per token). Uses Liquid's hybrid architecture — 18 double-gated LIV convolution blocks plus 6 grouped-query-attention layers — for a 131K-token context that runs in under ~6GB of memory on consumer hardware. A reasoning-only model that emits an explicit chain of thought before its answer, with strong tool-calling and agentic performance for its size. Builds on the October 2025 LFM2-8B-A1B, expanding the context window to 128K and scaling pretraining from 12T to 38T tokens. Released May 28 2026 under the LFM Open License; caught in a July catalog-gap sweep.
MiniMax-M2.7
AvailableOpen-weight agentic model from MiniMax focused on real-world software engineering, office tasks, tool use, and self-improving training workflows.
Qwen3.7-Max
AvailableAlibaba's proprietary flagship in the Qwen3.7 "Agent Frontier" line — a text-only sparse-MoE model with a 1M-token context, tuned for long-horizon agentic, coding, and reasoning workloads. Parameter count is undisclosed; access is API-only via Alibaba Cloud Model Studio / DashScope (and aggregators such as OpenRouter).
Gemini 3.5 Pro
PreviewAnnounced at Google I/O 2026; emphasizes deep multimodal reasoning over a 2M-token context. Recent reporting says the broad launch slipped from June toward July while testers continue using it in Google Antigravity and LMArena.
Gemini 3.5 Flash
AvailableGoogle's fast, cost-efficient Gemini 3.5 tier, unveiled at I/O 2026. Multimodal over a 1M-token context and tuned for agentic and coding workflows; Google says it beats Gemini 3.1 Pro on coding and tool-use while running ~4x faster.
Qwen3.6-27B
AvailableDense 27B that punches far above its weight on agentic coding — easy to self-host on a single GPU node.
ERNIE 5.1
AvailableBaidu's flagship ERNIE 5.1, derived from ERNIE 5.0 by extracting an optimal sub-network from its elastic sub-model matrix — compressing total parameters to ~1/3 and active parameters to ~1/2 of ERNIE 5.0 while reaching leading performance at only ~6% of the pre-training compute of comparable models. A Mixture-of-Experts model trained with a disaggregated fully-asynchronous RL stack and a multi-teacher on-policy-distillation pipeline, tuned for agentic execution, reasoning, world knowledge, and creative writing. Vendor-reported results: 99.6 on AIME26 (with tools, second only to Gemini 3.1 Pro), GPQA and MMLU-Pro approaching leading closed models, surpassing DeepSeek-V4-Pro on τ³-bench and SpreadsheetBench-Verified, and ranking 1st among Chinese models / 4th globally (score 1223) on the LMArena Search Arena. Proprietary; served via ERNIE Bot, Baidu AI Studio, and the Qianfan platform.
GPT-5.5-Cyber
PreviewOpenAI's limited-preview cybersecurity model for vetted defenders in Trusted Access for Cyber. It is tuned for more permissive authorized security workflows such as vulnerability triage, patch validation, malware analysis, red teaming, and controlled exploit validation, but is not generally available.
GPT-5.5
AvailableOpenAI's May 2026 GPT-5.5 release: a stronger frontier workhorse positioned for deep reasoning, coding, multimodal analysis, and long-context agent workflows. OpenAI lists an 800K-token input context and 128K-token output limit, with API pricing at $3 / $20 per million input/output tokens.
Grok 4.3
AvailablexAI's agentic flagship with a 1M-token context and aggressive API pricing.
DeepSeek V4-Flash
AvailableEfficient V4 companion model with 284B total / 13B active parameters and the same one-million-token context window.
DeepSeek V4-Pro
AvailablePreview-series sparse MoE flagship with a one-million-token context window and 1.6T total / 49B active parameters.
Hunyuan Hy3-preview
AvailableTencent's third-generation Hunyuan, rebuilt from scratch in ~90 days and open-sourced as the "Hy3 preview". A 295B-total / 21B-active Transformer MoE (80 layers, 192 experts with top-8 routing, plus a 3.8B multi-token-prediction layer) with a 256K-token context, positioned as a leading open reasoning-and-agent model for its size with strong cost efficiency. Vendor-reported results: 74.4 on SWE-bench Verified, 54.4 on Terminal-Bench 2.0, and 70.2 on WideSearch, with strong STEM-olympiad performance. Open weights on GitHub and Hugging Face under Tencent's community license.
Hunyuan-A13B-Instruct
AvailableTencent Hunyuan open-weight fine-grained MoE model with 80B total parameters and 13B active parameters, optimized for agentic tool use.
MiMo-V2.5-Pro
AvailableXiaomi's open-weight flagship: a 1.02T-parameter Mixture-of-Experts model with ~42B active parameters, a hybrid-attention architecture, and a 1M-token context window. Tuned for frontier-class agentic coding and long-horizon tasks (sustaining 1000+ tool calls with a proper harness). Open-sourced under the MIT license with weights and tokenizer on Hugging Face.
MiMo-V2.5
AvailableXiaomi's open-weight sparse-MoE model: ~310B total parameters with ~15B active, trained on ~48T tokens, with a 1M-token context window. Shipped alongside the larger MiMo-V2.5-Pro under the MIT license.
GPT-Rosalind
PreviewOpenAI's frontier reasoning model for life sciences, named after Rosalind Franklin and built to accelerate drug discovery, genomics, protein reasoning, and scientific research workflows. Optimized for multi-step, tool-heavy tasks (literature review, experimental design, sequence-to-function interpretation) with access to 50+ scientific databases via a Codex Life Sciences plugin. A June 3, 2026 update folded in GPT-5.5's agentic coding and tool use while using ~31% fewer tokens. Available as a research preview in ChatGPT, Codex, and the API through OpenAI's trusted-access program; not openly available.
GLM-5.1
AvailableZ.ai agentic-engineering follow-up to GLM-5, with stronger coding performance and better long-horizon tool-use behavior.
Muse Spark
AvailableMeta's new frontier model behind Meta AI for U.S. users, identified in reporting as the public release of the former Avocado effort and positioned to compete with Gemini, GPT, and Claude on multimodal assistant tasks.
Claude Mythos 5
PreviewThe restricted sibling of Claude Fable 5, sharing the same underlying Mythos-class model with fewer safeguards for vetted defensive-security and research use. Anthropic disclosed Mythos Preview on April 7, 2026, upgraded approved Project Glasswing users to Mythos 5 on June 9, suspended access under a U.S. export-control directive on June 12, and restored limited access for approved U.S. organizations on July 1 while it continues expanding the trusted-access program.
Gemma 4 31B
AvailableGoogle DeepMind's Gemma 4 advanced-reasoning open model for personal computers, part of the April 2026 Gemma 4 family.
GLM-5V-Turbo
AvailableZ.ai's native-multimodal vision agent: the first GLM model designed from the start as a multimodal agent, taking image, video, and text input and producing agent-oriented output (tool calling, task decomposition, and GUI interaction). Served via API with a ~203K-token context.
Kimi K2.6
AvailableMoonshot's open native multimodal agentic model for long-horizon coding, visual interface generation, and autonomous tool orchestration.
Mistral Medium 3.5
AvailableDense 128B open-weight model with a 256k context and strong coding performance for its size.
Nemotron 3 Super 120B-A12B
AvailableOpen-weight hybrid Mamba-Transformer MoE designed for collaborative agents and high-volume enterprise workflows.
Mistral Small 4
AvailableMistral's March 2026 Small release: the first Mistral model to unify reasoning (Magistral), multimodal understanding (Pixtral), and agentic coding (Devstral) into one Apache 2.0 model. A 119B-total / ~6B-active Mixture-of-Experts (128 experts, 4 active per token) with native text+image input, a 256K context, and a configurable reasoning_effort toggle for fast or deep responses. API pricing is $0.15 / $0.60 per million input/output tokens.
Step-3.5-Flash
AvailableStepFun's Apache-licensed sparse MoE model for fast agentic execution, coding, math, browsing, and tool-use workflows.
Meta Avocado
RetiredHistorical rumor/codename row retained for provenance. Recent reporting identifies the public productized model as Muse Spark, now tracked separately in the catalog.
Sarvam-105B
AvailableApache-licensed Indian-context MoE from Sarvam AI, optimized for reasoning, coding, agentic tasks, and 22 Indian languages.
GPT-5.4
AvailableWorkhorse GPT-5 release with a dedicated Thinking mode; widely deployed across ChatGPT and the API.
Qwen3.5-9B
AvailableThe flagship of Alibaba's small dense Qwen3.5 models. Independent analysis (Artificial Analysis) rated it the most intelligent model under 10B parameters at launch — roughly double the score of the next-closest sub-10B models — and the most intelligent multimodal model under 15B, leading peers on MMMU-Pro (~69%). A dense 9B with native vision, a 262K-token context, and the Qwen3.5 family's unified hybrid thinking / non-thinking mode. Native weights are BF16; in 4-bit it needs ~6GB, within reach of consumer laptops. High intelligence comes with heavy reasoning token usage (~260M output tokens to run the Intelligence Index).
Qwen3.5-4B
AvailableA dense 4B in Alibaba's small Qwen3.5 family, rated by Artificial Analysis as the most intelligent model under 5B parameters at launch — outscoring several 7B–9B peers despite roughly half the parameters. Native vision, a 262K-token context, and the family's hybrid thinking / non-thinking mode; Apache-2.0 licensed. Scores ~65% on MMMU-Pro multimodal reasoning and runs in ~3GB at 4-bit, suitable for lightweight on-device agents.
Qwen3.5-2B
AvailableA dense 2B Qwen3.5 model built for high-throughput, low-latency edge and on-device use. Despite its size it matches a 7B-class peer on Artificial Analysis's Intelligence Index. Apache-2.0, with native vision, a 262K-token context, and the family's hybrid thinking / non-thinking mode; runs in under 2GB at 4-bit, fitting laptops and smartphones.
Qwen3.5-0.8B
AvailableThe smallest Qwen3.5 model — a dense 0.8B designed for the most constrained on-device deployments, operating in non-thinking (instruct) mode by default. Apache-2.0, with native vision, a 262K-token context, and the family's hybrid thinking / non-thinking mode; needs roughly 2GB of VRAM and runs under 2GB at 4-bit, targeting smartphones and embedded hardware. Notable for a sub-1B model, it still scores ~26% on MMMU-Pro multimodal reasoning.
Qwen3.5-397B
AvailableNative vision-language MoE supporting 201 languages with a 1M-token context.
Gemini 3.1 Pro
AvailableGenerally available multimodal flagship with native tool use and a 2M-token context.
GLM-5
AvailableZ.ai flagship for complex systems engineering and long-horizon agentic tasks, scaling the GLM line to 744B total / 40B active parameters.
Claude Opus 4.6
AvailableIntroduced genuinely autonomous multi-file coding and stronger computer use.
GPT-5.3-Codex
AvailableOpenAI's February 2026 Codex update, optimized for agentic software engineering in ChatGPT, Codex, and the API. GPT-5.3-Codex improves code quality, patch reliability, repository-scale reasoning, and long-running autonomous coding workflows while keeping the 400K-token input context and 128K-token output limit of the Codex line.
Qwen3-Coder-Next
AvailableApache-licensed Qwen3-Next coding-agent model with 80B total / 3B active parameters, 256K context, and long-horizon tool-use training.
Kimi K2.5
AvailableOpen multimodal Kimi model that adds native visual agentic intelligence, instant and thinking modes, and agent-swarm workflows on top of the K2 base.
GLM-4.7
AvailableCoding-focused GLM release with improved multilingual agentic coding, terminal tasks, tool use, and interface generation.
GPT-5.2-Codex
AvailableOpenAI's December 2025 Codex model for agentic coding, released after GPT-5.2 with stronger repository understanding, code generation, and tool-use behavior for software engineering agents. The API model is listed with a 400K-token input context, 128K-token output limit, and $1.25 / $10 per million input/output tokens.
OLMo 3 Think 32B
AvailableAi2's fully open thinking model with public weights, code, data, checkpoints, and training details across the OLMo 3 pipeline.
Nemotron 3 Nano 30B-A3B
AvailableEfficient Nemotron 3 MoE checkpoint for agentic reasoning and coding, activating about 3B parameters while supporting 1M-token contexts.
GPT-5.2
AvailableOpenAI's December 2025 GPT-5.2 general model release, positioned as a stronger default for reasoning, coding, vision, instruction following, and long-context analysis. OpenAI lists a 400K-token input context, 128K-token output limit, and $2 / $12 per million input/output tokens.
GLM-4.6V
AvailableOpen 106B-class vision-language model with native multimodal function calling for visual agents.
Mistral Large 3
AvailableMistral's largest open-weight MoE, aimed at frontier reasoning while remaining self-hostable.
DeepSeek-V3.2
AvailableReasoning-first agent model that adds DeepSeek Sparse Attention and thinking directly inside tool-use workflows.
DeepSeek-V3.2-Speciale
AvailableHigh-compute reasoning variant of V3.2, positioned for olympiad-level math, programming, and other deep reasoning tasks.
LFM2 1.2B
AvailableLiquid AI hybrid model for efficient CPU/GPU/NPU local deployment, using short convolutions plus attention blocks.
Kimi K2 Thinking
AvailableOpen K2 reasoning-agent variant that interleaves step-by-step thinking with tool calls and supports stable 200-300 step tool-use trajectories.
Kimi-Linear-48B-A3B-Instruct
AvailableMIT-licensed hybrid linear-attention model using Kimi Delta Attention, built for million-token contexts with much lower KV-cache usage.
Claude Haiku 4.5
AvailableAnthropic's fast, low-cost Claude 4.5 model, released in October 2025 for latency-sensitive coding, tool-use, and customer-facing agent workloads. Anthropic positions it as bringing near-Sonnet capability to the Haiku tier at substantially lower cost and higher speed.
GLM-4.6
AvailableAgentic reasoning and coding upgrade over GLM-4.5, expanding the text context window from 128K to 200K tokens.
DeepSeek-V3.2-Exp
PreviewExperimental checkpoint that introduced DeepSeek Sparse Attention as an efficiency bridge between V3.1-Terminus and V3.2.
Claude Sonnet 4.5
AvailableAnthropic's September 2025 Sonnet release, positioned as its strongest model for coding, agents, and computer-use workflows at launch. Proprietary API model with text, vision, and code capabilities, 200K context, and Sonnet-tier list pricing.
DeepSeek-V3.1-Terminus
AvailableStability update to V3.1 focused on language consistency, code-agent reliability, and search-agent behavior.
Kimi K2 Instruct 0905
AvailableSeptember 2025 K2 update with stronger agentic coding, better frontend generation, and a doubled 256K context window.
Gemma 3 27B
AvailableGoogle's open multimodal model: 128k context, 140+ languages, runs on a single GPU.
DeepSeek-V3.1
AvailableHybrid thinking/non-thinking release that upgraded tool calling, long-context training, and agent task performance.
Seed-OSS-36B-Instruct
AvailableByteDance Seed's Apache-licensed long-context reasoning and agent model, with controllable thinking budgets and a native 512K context.
DeepSeek R2
RumoredRumored successor to DeepSeek R1. Reports say development and launch timing were affected by hardware constraints around Huawei Ascend training and Nvidia availability; final specs, license, and release date remain unconfirmed.
GLM-4.5V
AvailableVision-language GLM based on GLM-4.5-Air, covering image, video, document, grounding, and GUI-agent tasks.
Grok 5
RumoredRumored next major Grok model. Elon Musk said after GPT-5's launch that Grok 5 would arrive before the end of 2025, but no broad public Grok 5 release is logged in this catalog yet.
gpt-oss-20b
AvailableSmaller gpt-oss reasoning model optimized for local inference on systems with about 16GB of memory.
gpt-oss-120b
AvailableOpenAI's larger open-weight reasoning model, a 117B-total / 5.1B-active MoE with 128K context for local and self-hosted deployment.
Claude Opus 4.1
AvailableAnthropic's August 2025 Opus point release, focused on stronger coding, reasoning, and agentic reliability over Claude Opus 4. Proprietary API model with text, vision, and code capabilities.
Gemini 2.5 Deep Think
AvailableGoogle's enhanced Gemini 2.5 reasoning mode for harder math, science, coding, and multimodal analysis. Previewed at Google I/O 2025 and later made available to Gemini app subscribers, Deep Think uses more deliberative reasoning for complex prompts.
Falcon-H1 34B
AvailableA hybrid attention + state-space-model (SSM) design that matches 70B-class models with fewer parameters.
GLM-4.5
AvailableOpen agentic, reasoning, and coding foundation model that marked Z.ai international rebrand and MIT-licensed GLM push.
GLM-4.5-Air
AvailableCompact GLM-4.5 companion with 106B total / 12B active parameters for efficient agentic reasoning and coding.
Gemini 2.5 Flash-Lite
AvailableGoogle's lowest-latency, lowest-cost Gemini 2.5 tier, designed for summarization, classification, extraction, routing, and other high-volume production tasks. Proprietary API model with a 1M-token context and multimodal support.
Qwen3-Coder-480B-A35B-Instruct
AvailableAlibaba Qwen's large open coding-agent model: a 480B-total / 35B-active MoE released under Apache-2.0, tuned for code generation, repository-level software engineering, tool calling, and long-horizon agent workflows with a 256K-token native context.
EXAONE 4.0 32B
AvailableLG AI Research's unified model with non-reasoning and reasoning modes, agentic tool use, and English, Korean, and Spanish support.
Kimi K2 Instruct
AvailableOriginal open K2 post-trained model: a 1T-parameter MoE optimized for coding, reasoning, and tool-using agentic workflows.
Grok 4
DeprecatedxAI's fourth-generation Grok line, preceding the later 4.x API updates already tracked in the catalog.
SmolLM3 3B
AvailableHugging Face's fully open 3B multilingual long-context model with optional reasoning mode and 128K context.
ERNIE-4.5-300B-A47B
AvailableBaidu's open ERNIE 4.5 language MoE, part of a 10-variant Apache-licensed model family built with heterogeneous multimodal MoE training.
ERNIE-4.5-VL-424B-A47B
AvailableBaidu's largest ERNIE 4.5 vision-language MoE, supporting text, image, and video inputs with thinking and non-thinking modes.
Kimi-VL-A3B-Thinking-2506
AvailableUpdated MIT-licensed Kimi-VL reasoning model with better multimodal reasoning, video understanding, high-resolution perception, and lower thinking-token use.
Kimi-Dev-72B
AvailableMIT-licensed coding LLM trained with repository-level reinforcement learning for software issue resolution.
Gemini 2.5 Flash
AvailableGoogle's faster, lower-cost Gemini 2.5 model for high-throughput multimodal and agentic workloads. It brought Gemini 2.5's reasoning improvements to a production Flash tier with a 1M-token context and broad text, image, audio, video, and coding support.
MiniMax-M1-80k
AvailableOpen Apache-licensed hybrid-attention reasoning model with 456B total / 45.9B active parameters and a native 1M-token context.
Magistral Medium
AvailableMistral's first dedicated reasoning model family, released in Small open-weight and Medium enterprise/API tiers.
Magistral Small
AvailableOpen-weight 24B reasoning model from Mistral's Magistral family, popular for local reasoning experiments.
DeepSeek-R1-0528
AvailableMajor R1 reasoning update with stronger math, programming, general logic, function calling, and reduced hallucinations.
Claude Opus 4
DeprecatedFirst Claude 4 Opus model, positioned for long-running agentic and coding work before the 4.x point releases.
Seed Thinking v1.5
AvailableByteDance Seed reasoning model focused on long-horizon thinking and problem solving.
Sarvam-M
AvailableSarvam's medium-scale open model for multilingual Indian-language chat, reasoning, and translation tasks.
Devstral Small 2505
PreviewMistral and All Hands AI's open coding-agent model, released as a 24B Apache-2.0 research preview for software engineering tasks. Devstral is optimized for repository navigation, issue resolution, and agentic coding and is available via Hugging Face and Mistral's API.
Gemma 3n E4B
AvailableGoogle's mobile-first Gemma 3n model variant, built with a MatFormer-style architecture for efficient on-device multimodal inference. The E4B variant has roughly 4B effective parameters, supports text, vision, audio, and video-oriented use cases, and is released under Gemma terms.
Mistral Medium 3
AvailableMistral's May 2025 enterprise workhorse model, positioned as a high-performance, lower-cost alternative to larger proprietary systems for coding, STEM, enterprise search, and multilingual workloads. Mistral lists API pricing at $0.40 / $2 per million input/output tokens and offers hosted and enterprise deployment paths.
Phi-4 Reasoning
AvailablePhi-4 reasoning-specialized model family for math, science, and chain-of-thought style tasks.
Granite 3.3 8B
AvailableGranite 3.3 text update for enterprise chat, RAG, and instruction-following workflows.
Qwen3-235B-A22B
AvailableLargest open Qwen3 MoE, introducing hybrid thinking/non-thinking modes and 119-language coverage.
Kimi-Audio-7B-Instruct
AvailableOpen audio foundation model for audio understanding, generation, speech recognition, audio QA, captioning, and speech conversation.
Kimi-VL-A3B-Instruct
AvailableEfficient MIT-licensed vision-language MoE for OCR, image/video understanding, long documents, and OS-style agent tasks.
OpenAI o3
AvailableReasoning model released alongside o4-mini with tool use, image reasoning, and stronger agentic problem solving.
GPT-4.1
DeprecatedAPI model family focused on coding, instruction following, and one-million-token long-context work.
Llama 4 Maverick
AvailableMeta's flagship open-weight MoE; highest MMLU among open models at release.
Llama 4 Scout
AvailableEfficient open-weight MoE designed for very long context on modest hardware.
Llama 4 Behemoth
AnnouncedMeta's announced but unreleased Llama 4 teacher model: a multimodal MoE with 288B active parameters and nearly 2T total parameters. Meta says it was still training when Scout and Maverick shipped and that those released models were distilled from Behemoth.
Llama-3.3-Nemotron-Super-49B
AvailableOpen Llama Nemotron reasoning model from NVIDIA's 2025 Nemotron family.
Qwen2.5-Omni-7B
AvailableLocal omni-modal Qwen model that supports text, image, audio, video, and speech generation in a 7B package.
DeepSeek-V3-0324
AvailablePost-R1 V3 update with improved reasoning, front-end coding, Chinese writing, search, and function calling.
Gemini 2.5 Pro
DeprecatedReasoning-focused Gemini 2.5 model that made thinking a core part of Google's flagship model line.
Mistral Small 3.1
AvailableApache-licensed Small update adding vision and a 128K context window to the efficient 24B line.
ERNIE X1
AvailableBaidu's reasoning model released alongside ERNIE 4.5 before the open ERNIE 4.5 weights.
OLMo 2 32B
AvailableA fully open model — weights, data, and training code all public — and the first such to beat GPT-3.5 / GPT-4o mini.
Command A
AvailableEnterprise-grade model tuned for RAG, tool use, and multilingual business workloads.
Granite 3.2 8B
AvailableGranite 3.2 update with reasoning controls and multimodal/document-oriented Granite variants.
Claude 3.7 Sonnet
RetiredAnthropic's first hybrid-reasoning Sonnet. Shut down May 11, 2026 as the 4.x line matured.
Moonlight-16B-A3B-Instruct
AvailableMIT-licensed 16B/3B-active MoE trained with Moonshot's scalable Muon optimizer experiments.
DeepHermes 3 Llama 3 8B
AvailableNous reasoning-oriented Hermes model trained to combine concise answers with optional deep reasoning traces.
Grok 3
DeprecatedxAI's third-generation model family, introduced with stronger reasoning, search, and coding modes.
Dolphin 3.0 Llama 3.1 8B
AvailablePopular local assistant model tuned for coding, math, function calling, and agentic workflows.
Mistral Small 3
AvailableA latency-optimized 24B dense model under Apache-2.0 — a popular local-deployment workhorse.
Qwen2.5-Max
AvailableProprietary MoE flagship for the Qwen2.5 generation, released through Qwen Chat and Alibaba Cloud APIs.
Qwen2.5-VL-72B
AvailableVision-language Qwen2.5 model for image, document, video, and agentic visual grounding tasks.
Doubao-1.5-pro
AvailableDoubao 1.5 Pro update positioned for stronger multimodal, reasoning, and agentic work in Volcano Engine.
DeepSeek-R1
AvailableBreakout open reasoning model trained with large-scale reinforcement learning and released with weights under MIT.
Kimi k1.5
AvailableMoonshot's multimodal reinforcement-learning reasoning model, reported as matching OpenAI o1 on math, coding, and multimodal reasoning.
MiniMax-01
AvailableOpen MiniMax generation with MiniMax-Text-01 and MiniMax-VL-01 long-context models.
DeepSeek-V3
AvailableThe 671B/37B-active MoE release that made DeepSeek a central open-model lab before the R1 breakthrough.
Step-2
AvailableSecond-generation StepFun foundation model line with larger-scale multimodal and reasoning ambitions.
Granite 3.1 8B
AvailableIBM's enterprise-focused open model with a 128k context, Apache-2.0 licensed.
Falcon 3 10B
AvailableUAE's TII open model designed to run on light infrastructure, including laptops.
Command R7B
AvailableCohere's smallest, fastest R-series model, tuned for RAG and tool use on modest hardware.
Phi-4
AvailableA 14B dense model that rivals far larger ones on math and reasoning, under a permissive MIT license.
Gemini 2.0 Flash
DeprecatedFirst Gemini 2.0 release, built for native multimodal input/output, tool use, and agentic product integrations.
EXAONE 3.5 32B
AvailableEXAONE 3.5 32B open-weight model for bilingual reasoning, coding, and long-context tasks.
Llama 3.3 70B
AvailableLate-2024 70B Llama update delivering much of the 405B instruction-following quality at lower serving cost.
OpenAI o1
DeprecatedGeneral release of OpenAI's o1 reasoning model with stronger deliberative reasoning and multimodal ChatGPT integration.
Amazon Nova Pro
AvailableAWS-native multimodal model with a 300k context; size and architecture undisclosed.
Amazon Nova Lite
AvailableLower-cost multimodal Nova understanding model for text, image, and video inputs.
QwQ-32B-Preview
AvailableQwen's first public reasoning-preview model, aimed at math, coding, and deliberate problem solving.
Tulu 3 405B
AvailableAi2's post-trained open instruction model line, scaling the Tulu recipe to Llama 3.1 405B.
DeepSeek-R1-Lite-Preview
RetiredReasoning-preview model exposed in DeepSeek Chat ahead of the open DeepSeek-R1 release.
Qwen2.5-Coder-32B
AvailableCode-specialized Qwen2.5 model family, with the 32B checkpoint as the flagship open coding model.
Hunyuan-Large
AvailableTencent's 389B total / 52B active open-weight Transformer MoE, released with a 256K pretraining context and 128K instruct context.
SmolLM2 1.7B
AvailableCompact on-device model family trained on 11T tokens, popular for lightweight local chat and experimentation.
Claude 3.5 Haiku
DeprecatedFast, lower-cost Claude 3.5 model for latency-sensitive coding, tool-use, and customer-facing workloads.
Sarvam-1
AvailableSarvam's 2B open model trained for ten major Indian languages.
Granite 3.0 8B
AvailableApache-licensed Granite 3.0 text model, part of IBM's push toward enterprise-friendly open models.
Yi-Lightning
Available01.AI's MoE API model that reached the global top-10 on Chatbot Arena, strong in Chinese, math, and coding.
Ministral 8B
AvailableSmall Mistral model line optimized for edge and low-latency workloads.
Llama-3.1-Nemotron-70B
AvailableNVIDIA-tuned Llama 3.1 70B instruction model optimized with Nemotron reward and alignment recipes.
Llama 3.2 90B Vision
AvailableFirst Llama family release with native vision models, alongside smaller edge-oriented 1B and 3B text models.
Molmo 72B
AvailableOpen multimodal model family trained for strong image understanding, pointing, and visual grounding.
Qwen2.5-72B
AvailableBroad Qwen2.5 foundation-model update spanning general, coding, math, and multimodal descendants.
Pixtral 12B
AvailableMistral's first open multimodal model, adding image understanding to a Mistral text backbone.
OpenAI o1-preview
RetiredOpenAI's first public reasoning-model preview, optimized to spend more inference time on hard math, coding, and science tasks.
Yi-Coder-9B
Available01.AI's compact code model trained for repository-scale programming and code completion tasks.
DeepSeek-V2.5
AvailableUnified DeepSeek V2 generation combining general-chat and coding strengths before the V3 series.
Hunyuan Turbo
AvailableTencent's faster, lower-cost Hunyuan update before the open Hunyuan-Large model card.
OLMoE 1B-7B
AvailableFully open sparse MoE model with 7B total and about 1B active parameters.
Jamba 1.5 Large
AvailableIsrael's AI21 hybrid Mamba-Transformer MoE, with a 256k context and strong long-document throughput.
Phi-3.5 MoE
AvailablePhi-3.5 mixture-of-experts model, scaling Microsoft's small-model line while preserving efficient active parameters.
Hermes 3 Llama 3.1 405B
AvailableLarge Hermes 3 instruction-tuned model built on Meta's Llama 3.1 405B.
Grok-2
RetiredSecond-generation Grok release with Grok-2 and Grok-2 mini for chat, coding, reasoning, and image-enabled product experiences.
EXAONE 3.0 7.8B
AvailableLG's first open-weight EXAONE model, a compact bilingual instruction model for Korean and English.
MiniCPM-V 2.6
Available8B vision-language model for local image, multi-image, OCR, and video understanding, with llama.cpp and Ollama support.
Llama 3.1 405B
AvailableMeta's first frontier-scale open Llama model, with 405B parameters, 128K context, multilingual support, and tool-use improvements.
Mistral NeMo
AvailableApache-licensed 12B model co-developed with NVIDIA, including a 128K context window and strong multilingual tokenization.
Gemma 2 27B
AvailableSecond-generation Gemma model, improving open-weight quality and efficiency at 9B and 27B sizes.
Claude 3.5 Sonnet
RetiredMajor Sonnet upgrade that became Anthropic's default high-intelligence workhorse for coding, writing, and visual reasoning.
DeepSeek-Coder-V2
AvailableOpen code-focused MoE built from DeepSeek-V2, expanding programming-language coverage and coding benchmark performance.
Nemotron-4 340B
AvailableNVIDIA's large open model family for synthetic data generation and reward modeling.
Qwen2-72B
AvailableQwen2's largest dense model, introducing stronger multilingual support, coding/math gains, and long-context variants.
GLM-4-9B
AvailableOpen GLM-4 9B model family, covering chat, long-context, and code-oriented variants.
Codestral 22B
AvailableMistral's first code-specialized model, trained for code generation, fill-in-the-middle, and multi-language programming tasks.
Aya 23 35B
AvailableOpen multilingual research model covering 23 languages, released by Cohere For AI.
Doubao-pro
AvailableByteDance's commercial Doubao foundation model line for text, code, and assistant workloads.
GPT-4o
RetiredThe 2024 omni-modal model that defined a generation of assistants. Deprecated in Feb 2026 and fully retired across ChatGPT on April 3, 2026.
Yi-1.5-34B
AvailableYi 1.5 update with stronger instruction following, coding, math, and multilingual performance.
Falcon 2 11B
AvailableFalcon 2 generation, including text and vision-language 11B models under a permissive TII license.
DeepSeek-V2
AvailableDeepSeek's first major MoE general model with Multi-head Latent Attention and low-cost API positioning.
Granite Code 34B
AvailableApache-2.0 code model from IBM's Granite Code family, used for local code generation and enterprise coding assistants.
Amazon Titan Text Premier
AvailableLarger Titan text model for enterprise RAG, summarization, and agent workflows in Amazon Bedrock.
Snowflake Arctic
AvailableApache-2.0 enterprise LLM with 480B total / 17B active parameters, optimized for SQL, code, and instruction following.
Phi-3 Mini
Available3.8B-parameter Phi-3 model released as a phone-capable small model with 4K and 128K variants.
Llama 3 70B
AvailableFirst Llama 3 release, with 8B and 70B open models and a stronger tokenizer, data mix, and post-training stack.
Mixtral 8x22B
AvailableLarger open Mixtral sparse MoE with 141B total and 39B active parameters, released under Apache-2.0.
abab6.5
AvailableMiniMax's commercial long-context abab model generation before the open MiniMax-01 and M series.
WizardLM-2 8x22B
AvailableMicrosoft's WizardLM-2 MoE chat model, widely mirrored and run locally after its model-card release.
Step-1V
AvailableStepFun's first major vision-language model, released after the Step-1 language model.
CodeGemma 7B
AvailableOpen code-specialized Gemma model for local code completion, generation, and instruction-following.
Command R+
DeprecatedHigher-capability RAG and tool-use model in Cohere's Command R family.
Grok-1.5
RetiredGrok update with stronger reasoning and a 128K context window.
Jamba
AvailableFirst Jamba hybrid Transformer-Mamba MoE model with open weights and a 256K context length.
DBRX Instruct
AvailableDatabricks' 132B-total / 36B-active open MoE model for code, math, RAG, and enterprise self-hosted workloads.
Step-1
AvailableStepFun's first public foundation model generation, introduced as a trillion-parameter Chinese model line.
Kimi 1M
AvailableLong-context Kimi upgrade advertised with support for million-character document and conversation contexts.
Command R
DeprecatedEnterprise RAG-focused model with tool use, citations, multilingual retrieval, and long-context support.
Claude 3 Opus
DeprecatedHighest-capability Claude 3 model, launched with Sonnet and Haiku and Anthropic's first major vision-capable Claude family.
StarCoder2 15B
AvailableNext-generation BigCode code model trained on 4T+ tokens and 600+ programming languages, with 16K context.
Mistral Large
DeprecatedMistral's first proprietary flagship API model, introduced alongside Le Chat and stronger multilingual/coding performance.
Gemma 7B
AvailableFirst Gemma open-weight text model family, derived from the same research lineage as Gemini.
Gemini 1.5 Pro
DeprecatedGemini generation that introduced production-scale long context, eventually expanding to a two-million-token window.
Qwen1.5-110B
AvailableLargest Qwen1.5 model, released as the bridge from the original Qwen line to Qwen2.
Qwen1.5-72B-Chat
AvailableLargest chat-tuned Qwen1.5 dense checkpoint, released with stronger human-preference alignment, multilingual support, and 32K context.
OLMo 7B
AvailableAi2's first fully open language model release, including weights, training data, code, logs, and intermediate checkpoints.
Stable LM 2 1.6B
AvailableSmall multilingual Stable LM release built for low hardware barriers and local experimentation.
GLM-4
AvailableZhipu's GLM-4 flagship generation, launched as the successor to ChatGLM3 with stronger tool use and multimodal variants.
DeepSeekMoE 16B
AvailableEarly DeepSeek sparse MoE research model that foreshadowed the later V2/V3 architecture direction.
Nous Hermes 2 Mixtral
AvailableNous instruction-tuned Mixtral model with strong open-chat and tool-use adoption.
OpenChat 3.5
AvailableCompact Mistral-based local chat model trained with C-RLFT, popular in early 2024 local leaderboards.
TinyLlama 1.1B Chat
AvailableCompact Llama-style 1.1B chat model trained for local experimentation and low-memory deployments.
Phi-2
Available2.7B-parameter Phi model showing strong reasoning and language understanding at small scale.
OpenHathi-7B
AvailableSarvam AI's first open Indic language model, adapted from Llama 2 for Hindi and Indian-language work.
Mixtral 8x7B
AvailableThe open sparse Mixture-of-Experts that brought MoE efficiency to the open ecosystem.
Gemini 1.0 Ultra
DeprecatedGoogle's first natively multimodal Gemini flagship, since superseded by the 1.5/2/3 lines.
Qwen-72B
AvailableAlibaba's first major open Qwen model and the start of a prolific open-weight line.
DeepSeek LLM 67B
AvailableFirst general DeepSeek language model family, with 7B and 67B base/chat checkpoints.
Yi-34B-Chat
AvailableChat-tuned Yi-34B checkpoint from 01.AI, released alongside quantized chat variants for bilingual open-weight assistants.
Claude 2.1
RetiredClaude update with a 200K context window, lower hallucination rates, and improved tool-use beta support.
Yi-34B
Available01.AI's strong bilingual open model, with a 200k-context variant.
GPT-4 Turbo
DeprecatedLower-cost GPT-4 generation with a 128K context window, introduced at OpenAI DevDay.
Grok-1
AvailablexAI's first Grok model, later released as open weights with a 314B-parameter MoE checkpoint.
DeepSeek Coder 33B
AvailableDeepSeek's first public code-model family, released before the general DeepSeek LLM line.
ERNIE 4.0
AvailableBaidu's fourth-generation ERNIE flagship, announced with stronger understanding, generation, reasoning, and memory.
Kimi Chat
AvailableMoonshot's first Kimi assistant release, establishing the long-context product line before the open Kimi model cards.
LLaVA 1.5 13B
AvailableOpen vision-language assistant and one of the most widely run early local multimodal models.
Amazon Titan Text Express
AvailableAmazon's first-party Titan text generation model exposed through Bedrock, initially alongside embeddings and image models.
Mistral 7B
AvailableThe 7B that punched far above its weight and put Mistral on the map.
Qwen-14B
AvailableSecond open Qwen size, expanding the first-generation Qwen language-model lineup.
Granite 13B
AvailableIBM's early Granite foundation model family for enterprise language and code tasks.
Hunyuan
AvailableTencent's first Hunyuan foundation model release, introduced as a general-purpose Chinese enterprise model.
Falcon 180B
AvailableAt launch the largest openly available model, from the UAE's TII.
Code Llama 34B
AvailableMeta's first code-specialized Llama model family, released in base, Python, and instruction-tuned variants.
Qwen-7B
AvailableAlibaba's first open Qwen checkpoint and the start of the Qwen open-model line.
Nous-Hermes-Llama2-13B
AvailableEarly Nous Hermes instruction model on Llama 2, widely used in the open-model fine-tuning ecosystem.
EXAONE 2.0
RetiredSecond EXAONE generation, improving bilingual Korean-English performance and enterprise deployment options.
Llama 2 70B
AvailableThe release that made capable open-weight models genuinely usable for production.
Claude 2
RetiredAnthropic's first widely-available Claude, notable for an early 100k-token context window.
ChatGLM2-6B
AvailableSecond open ChatGLM generation, improving long context, inference efficiency, and bilingual chat quality.
Phi-1
AvailableMicrosoft's first Phi small-language-model release, demonstrating strong code performance from textbook-quality synthetic data.
Falcon 40B
AvailableTII's breakout open Falcon model, released before Falcon 180B and trained on the RefinedWeb corpus.
PaLM 2
RetiredGoogle's improved multilingual, reasoning, and coding foundation model family introduced at I/O 2023.
MPT-7B
AvailableMosaicML's permissively licensed 7B model, an early favorite for commercial local fine-tuning and long-context variants.
Vicuna 13B
AvailableLMSYS instruction-tuned LLaMA model that became a landmark early local ChatGPT-style assistant.
ERNIE Bot
AvailableBaidu's public chat assistant launch, built on the ERNIE foundation-model line.
GPT-4
DeprecatedThe model that brought reliable multi-step reasoning to the mainstream; size never disclosed.
ChatGLM-6B
AvailableZhipu AI and Tsinghua KEG's first widely used open bilingual ChatGLM checkpoint.
Claude 1
RetiredAnthropic's first broadly announced Claude assistant model, launched through an API and select product partners.
Jurassic-2 Ultra
DeprecatedSecond-generation Jurassic model with better multilingual support, lower latency, and instruction following.
GPT-3.5 Turbo
RetiredOpenAI's first ChatGPT API model, bringing the GPT-3.5 chat-tuned line to developers at much lower cost than text-davinci-003.
LLaMA
AvailableMeta's first LLaMA, released to researchers; its leak catalyzed the open-weight movement.
Galactica
WithdrawnA science-focused model whose public demo was withdrawn after just three days over confidently wrong outputs — an early, instructive retraction.
BLOOM
AvailableAn open, multilingual 176B model (46 languages) from a global research collaboration.
PaLM
RetiredGoogle's 540B Pathways model; the API was later deprecated in favor of Gemini.
EXAONE 1.0
RetiredLG AI Research's first EXAONE foundation model generation, introduced as a large multimodal expert AI.
ERNIE 3.0 Titan
RetiredBaidu's 260B-parameter ERNIE 3.0 Titan model, an early Chinese frontier-scale language model.
Jurassic-1 Jumbo
RetiredAI21's first major API language model, launched through AI21 Studio.
GPT-3
RetiredThe 175B model that proved in-context learning at scale; its base API models were retired in 2024.
GPT-2
AvailableInitially withheld over misuse fears, then fully released in Nov 2019 — an early 'limited release' debate.
BERT
AvailableThe bidirectional encoder that reshaped NLP and seeded the transformer era.