Last 30 days
All reportsLLM release radar
A compact briefing of the most recent model releases and lifecycle changes, anchored to the newest tracked event in the catalog.
A fast read on what shipped recently: the newest releases and lifecycle changes across every lab we track, anchored to the latest event in the catalog and limited to the past 30 days. Every item is source-linked so you can verify it at the origin.
The window is based on tracked event dates, not publication time on this site. Sources are linked for every item where we have one.
Freshest events
- sourceMoonshot AI
Moonshot AI publishes the full Kimi K3 weights to Hugging Face under a Modified MIT license on July 26 — a day ahead of its announced July 27 target — making the 2.8T-parameter MoE freely downloadable, modifiable, and self-hostable, and cementing K3 as the largest open-weight model publicly available.
- sourceReleasedAnthropic releases Claude Opus 5Anthropic
Anthropic releases Claude Opus 5, its flagship model for demanding reasoning, autonomous coding, and long-horizon agentic work — pitched as the go-to model for most knowledge work, approaching Fable 5 capability in many categories at about half the price. Adds a five-level 'effort' dial on the Claude API/Platform to trade compute for capability, a 1M-token context at standard pricing, and up to 128K output tokens. Standard pricing $5/$25 per Mtok (matching Opus 4.8), plus a $10/$50 Fast mode; becomes the default for Claude Max subscribers. Anthropic calls it its most aligned Opus model.
- sourceAnt Group (inclusionAI)
Ant Group's inclusionAI lab releases Ling-3.0-flash, a hybrid-reasoning Mixture-of-Experts model with 124B total parameters and ~5.1B active per token (1/64 expert activation), built for production-scale agents. Ant claims it matches or beats its own ~1T-parameter flagship on most benchmarks shown at 1/8 the total and 1/12 the active parameters — a vendor claim with no public benchmark table at launch. Uses a KDA + MLA hybrid-linear attention stack at a reported 5:1 ratio for an economical 256K-token context. Announced as open-weight under Apache 2.0, but weights and a model card were not yet posted to Hugging Face as of July 24; usable only via hosted API, free on OpenRouter and Vercel AI Gateway through August 3 2026.
- sourceGoogle DeepMind
A cyber-specialized model built on 3.5 Flash for finding and fixing software vulnerabilities, deployed inside the CodeMender agent and reaching competitive frontier performance on CyberGym. Limited-access pilot restricted to governments and trusted partners.
- sourceReleasedPoolside releases Laguna S 2.1Poolside
Poolside releases Laguna S 2.1, a 118B-total / 8B-active open-weight MoE coding model with a 1M-token context, pitched as 'the West's most capable open-weight model' for its weight class. It scores 70.2% on Terminal-Bench 2.1 and tops the published open disclosed-size table on SWE-bench Multilingual at 78.5%. Trained in under nine weeks on 4,096 H200 GPUs, it ships weights on Hugging Face under OpenMDW-1.1 (BF16/FP8/INT4/NVFP4 + GGUF/MLX) and runs at 4-bit on a single NVIDIA DGX Spark. Hosted free at 256K context and paid at full 1M context via OpenRouter ($0.10/$0.20/$0.01 per 1M input/output/cache-read tokens).
- sourceGoogle DeepMind
Fastest, most cost-effective 3.5-class model (~350 output tokens/s) for high-throughput agentic workloads, priced at $0.30 / $2.50 per Mtok with a 1M-token context and built-in computer use. Large step up on 3.1 Flash-Lite and beats 3 Flash on several agentic/coding evals.
- sourceReleasedGoogle launches Gemini 3.6 FlashGoogle DeepMind
Google DeepMind ships Gemini 3.6 Flash, a multimodal 1M-context workhorse for agentic workflows that improves on 3.5 Flash in coding, knowledge work, and computer use while cutting output-token usage ~17% at a lower price ($1.50 / $7.50 per Mtok). Available in the Gemini API, Gemini Enterprise, and the Gemini app.
- sourceDeepSeek
V4-Flash, the efficient tier of the V4 family, exits preview alongside V4-Pro as DeepSeek retires its legacy deepseek-chat and deepseek-reasoner aliases (cutoff July 24, 2026).
- sourceDeepSeek
DeepSeek moves the V4 model family (V4-Pro and V4-Flash) out of preview and into general availability, closing a run of just under three months from the April 24 preview. The legacy deepseek-chat and deepseek-reasoner API aliases are retired on July 24, 2026 with no fallback; production traffic must reference deepseek-v4-pro / deepseek-v4-flash.
- sourceAlibaba (Qwen)
Alibaba launches Qwen3.8, a 2.4-trillion-parameter fully-multimodal model — more than double the size of its predecessor — which the Qwen team says ranks "second only to Fable 5" on overall performance (internal evals, no independent benchmarks yet). Qwen3.8-Max-Preview is available to developers via Alibaba's Token Plan subscription and the Qoder / QoderWork coding platforms, with open weights promised "soon" but no timeline, license, or architecture details disclosed.
- sourceGoogle DeepMind
Bloomberg reports Google delayed Gemini 3.5 Pro again after the rebuilt model fell short of internal quality goals on hallucinations and reliability — its third slipped target after June and early July. Google DeepMind's Logan Kilpatrick said on July 21 the company is testing it with partners and hopes to "land soon"; there is still no model card or public API entry.
- sourceReleasedMoonshot AI releases Kimi K3Moonshot AI
Largest open-weight model to date: 2.8T-parameter MoE (896 experts, 16 active) with Kimi Delta Attention, native multimodal input, and a 1M-token context. API launched Jul 16 at $3/$15 per Mtok; full weights slated for Jul 27. Reported 93.5% GPQA Diamond (strongest published open-weight result).
- sourceThinking Machines Lab
Mira Murati's lab ships its first model — a 975B-total / 41B-active multimodal MoE (text/image/audio in, text out) pretrained on ~45T tokens, released under Apache-2.0 with weights on Hugging Face (1M context) and hosted on the Tinker API (256K). Debuts at 41 on the Artificial Analysis Intelligence Index, the leading U.S. open-weights model.
- sourceKwaipilot (Kuaishou)
The efficient ~32B-active variant of KAT-Coder V2.5, sharing the 256K context and agentic tool-use focus at roughly a fifth of Pro's price ($0.15/$0.60 per Mtok).
- sourceKwaipilot (Kuaishou)
Kuaishou's Kwaipilot team releases KAT-Coder-Pro V2.5, an agentic coding MoE (~72B active) trained with large-scale agentic RL in verifiable repository environments, with a 256K context, 80K max output, and API pricing of $0.74/$2.96 per Mtok via StreamLake, Atlas Cloud, and OpenRouter.
- sourceOpenAI
OpenAI moves the GPT-5.6 family to general availability after the June 26 restricted preview. Sol, the flagship for difficult professional, coding, research, computer-use, and tool-heavy work, ships with a 1.05M-token context window, 128K max output, and standard pricing of $5/$30 per Mtok, alongside a new max reasoning effort and ultra subagent mode.
- sourceOpenAI
Luna, the fast, most cost-efficient tier of the GPT-5.6 family ($1/$6 per Mtok), moves from restricted preview to general availability alongside Sol and Terra.
- sourceOpenAI
Following U.S. government review and a two-week restricted preview, OpenAI rolls out the full GPT-5.6 family — Sol, Terra, and Luna — to general availability across ChatGPT (Plus, Pro, Business, Enterprise) and the API.
- sourceOpenAI
Terra, the balanced everyday-work tier of the GPT-5.6 family ($2.50/$15 per Mtok), moves from restricted preview to general availability alongside Sol and Luna, available to Plus, Pro, Business, and Enterprise users in ChatGPT and via the API.
- sourceMeta AI
Meta Superintelligence Labs releases Muse Spark 1.1 in US public preview on the Meta Model API, the first time Meta charges for one of its models. A natively multimodal reasoning model (text/image/video/PDF/audio in, text out) with a self-compacting 1M-token context, aimed at agentic and coding workflows and priced at $1.25/$4.25 per Mtok — about a quarter of comparable Anthropic/OpenAI models. Closed weights; benchmarks around the Opus 4.8 / GPT-5.5 tier.
- sourcexAI
SpaceXAI (the rebranded xAI) releases its most capable model to date — the first Grok trained jointly with Cursor (Anysphere), pitched as "Opus-class" but faster, cheaper ($2/$6 per Mtok), and ~4x more token-efficient than Opus 4.8 on SWE-bench Pro. Available in Grok Build (default), Cursor (all plans), and the SpaceXAI console; initially unavailable in the EU.
- sourceMistral AI
CEO Arthur Mensch confirms Mistral is preparing a new open-weight Mixture-of-Experts family — "fat but sparse" — entering early access in July 2026 and aimed at the frontier open-weight tier. No parameter count, benchmarks, license terms, or release date disclosed; tracked as rumored.
- sourceNVIDIA
A compressed variant of Nemotron-3-Super produced with "Iterative Puzzle", trimming the parent to 75.3B total / 9.3B active while keeping the hybrid Mamba-Transformer LatentMoE design and 1M context. NVIDIA reports ~2x higher server throughput at matched user throughput; released on Hugging Face in BF16/FP8/NVFP4 under OpenMDW-1.1.
- sourceTencent Hunyuan
Tencent officially launches Hunyuan 3.0 (Hy3), the GA of its rebuilt third-generation model, and open-sources it under Apache-2.0. A 295B-total / 21B-active MoE (plus a 3.8B multi-token-prediction layer) with a 256K context and three selectable fast/slow inference modes; Tencent reports it rivals GLM-5.2 and DeepSeek-V4 and matches or surpasses GPT-5.5 on several science benchmarks, with 78.0 on SWE-bench Verified. Weights on Hugging Face (tencent/Hy3) and ModelScope, with a free OpenRouter route (tencent/hy3:free) through July 21, 2026.
- sourceReleasedPoolside releases Laguna XS 2.1Poolside
Poolside releases Laguna XS 2.1, an upgraded 33B-A3B open-weight MoE coding model served at 256K context. It raises SWE-bench Multilingual by 5.4 points to 63.1% over XS.2, adds open-weighted DFlash speculator models that roughly double local tokens/sec, and moves to the fully permissive OpenMDW-1.1 license. Weights are on Hugging Face (BF16/FP8/INT4/NVFP4) and it's available free on OpenRouter, with paid API pricing of $0.10/$0.20/$0.05 per 1M input/output/cache-read tokens.
- sourceAnthropic
Anthropic restored Mythos 5 access for approved U.S. organizations and continues expanding the Glasswing trusted-access program, while Mythos remains restricted rather than generally available.
- sourceAnthropic
Anthropic says Claude Fable 5 is available globally again on Claude Platform, Claude.ai, Claude Code, and Claude Cowork after export controls were lifted; cloud partner access is being re-enabled and safeguards now route high-risk requests to Opus 4.8.
- sourceMeituan (LongCat)
Meituan releases LongCat-2.0, a 1.6T-parameter MoE (~48B active) with a 1M-token context for agentic coding, open-sourced under the MIT license on Hugging Face and GitHub. It was trained and served on a ~50,000-card cluster of domestic Chinese AI chips — the first trillion-parameter model Meituan says completed full-process training and inference on home-grown hardware. Vendor-reported results: 59.5 SWE-bench Pro (vs GPT-5.5's 58.6), 70.8 Terminal-Bench 2.1, and 77.3 SWE-bench Multilingual.
- sourceAnthropic
Anthropic releases Claude Sonnet 5, its most agentic Sonnet yet — performance approaching Opus 4.8 at a lower price, made the default model on the Free and Pro plans and available to Max, Team, and Enterprise users. Available in Claude Code and via the Claude API as claude-sonnet-5, with introductory pricing of $2 per Mtok input / $10 per Mtok output through Aug 31, 2026, then standard $3/$15.
- sourceAnthropic
Reporting says the White House lifted the ban on Anthropic's models after an agreement on additional safeguards, authorizing Anthropic to return Fable 5 to public release channels while Mythos remains limited to pre-vetted partners.
- sourceBase44
Base44 (a Wix company) rolls out Base1, a general-purpose 'vibe coding' agent fine-tuned on an open-source foundation model using data from tens of millions of platform interactions, selectable alongside GPT-5.5 and Claude Opus 4.8 in its model picker.
- sourceAnthropic
The June 26 government carveout restored only narrow Mythos access; Fable remains withdrawn from broad public availability.
- sourcePreviewOpenAI previews GPT-5.6 LunaOpenAI
Luna is the lower-cost GPT-5.6 preview variant, still gated by the government-review access process.
- sourcePreviewOpenAI previews GPT-5.6 SolOpenAI
Sol is the highest-capability GPT-5.6 preview tier, available only to a small vetted cohort.
- sourceOpenAI
Terra is the mid-tier GPT-5.6 preview variant in the Sol/Terra/Luna rollout.
- sourceAnthropic
U.S. Commerce officials reportedly allowed Anthropic to restore limited Mythos access under tighter controls, while broader Fable access remains restricted.
- sourceOpenAI
The Sol, Terra, and Luna variants entered a tightly restricted preview for vetted customers while U.S. officials review security risks.
Frequently asked questions
How recent is the release radar?
It shows lifecycle events from the last 30 days, measured against the newest event we have tracked — not the time you happen to load the page. That keeps the window stable even between crawls.
What kinds of events show up here?
New releases, previews, updates, benchmark and price changes, deprecations, retirements, and withdrawals. Each item links to the model and, where we have one, the original source.
How is this different from the release calendar?
The radar is a rolling 30-day briefing for “what just happened”. The release calendar groups the same source-backed events by month so you can scan the longer-term cadence.
Related