LLM Releases

Model family timeline

Last updated Jul 23, 2026

Ling model releases

A source-backed timeline for the Ling model family, collecting release dates, labs, access details, context windows, and major lifecycle changes.

1
Models
1
Labs
1
Open
1
Recent

1 model

Ling-3.0-flash

Available
Ant Group (inclusionAI)Open weights

Ant Group's efficiency-focused Mixture-of-Experts model, released July 23 2026 by its inclusionAI lab: 124B total parameters activating only ~5.1B per token (1/64 expert activation). Ant claims it matches or beats the company's own ~1T-parameter Ling-2.6 flagship on most benchmarks it shows, at 1/8 the total and 1/12 the active parameters — a vendor claim with no public benchmark table or independent audit at launch, so treat it as unverified. Built for production-scale agents (MCP tool use, multi-agent coordination) rather than chat, with both thinking and non-thinking modes. Architecture is a native hybrid-linear attention stack interleaving Kimi Delta Attention (KDA) and Multi-head Latent Attention (MLA) at a reported 5:1 ratio, giving an economical 262,144-token (256K) context, with 1M cited as the scaling target. Ant docs claim peak inference up to 1,000 tokens/s and <100ms time-to-first-token on its own stack. Announced as open-weight under Apache 2.0, but as of July 24 no weights or model card were posted to the inclusionAI Hugging Face org — so the license and open-weight status are announced but unconfirmed (weights not yet downloadable; not self-hostable today). Usable now only via hosted API — free on OpenRouter (as inclusionai/ling-3.0-flash:free, hosted by Novita) and Vercel's AI Gateway through August 3 2026; no post-promo per-token price published at launch.

MoE124B262K ctxJul 23, 2026