LLM Releases
← Catalog

Ling-3.0-flash-VL

Available
Ant Group (inclusionAI)Open source

The natively multimodal member of Ant Group inclusionAI's Ling-3.0 line, released Sep 10 2026 with MIT open weights (inclusionAI/Ling-3.0-flash-VL). A 124B-total / ~5.5B-active sparse MoE that inherits Ling-3.0-flash's language, reasoning, and long-context ability and extends it with native image and video understanding: a ViT visual encoder feeding a two-layer MLP projector, VideoRoPE positional encoding for video, and the family's 42-layer hybrid backbone alternating Kimi Delta Attention and Gated MLA layers at a 5:1 ratio. Carries a 256K-token context. inclusionAI frames the vision work around three axes — understanding complex visual information, reasoning over visual evidence, and interacting with interfaces (GUI agents) — and reports 42 on the Artificial Analysis Intelligence Index v4.1.1, four points above text-only Ling-3.0-flash at 38. Served free at launch via OpenRouter (inclusionai/ling-3.0-flash-vl) alongside self-hosting on vLLM. Vendor figures, unverified independently at launch.

Specifications

License
Open source · MIT
Weights
Downloadable
Architecture
Mixture-of-Experts
Parameters
124B · 5.5B active
Context window
262K tokens
Max output
33K tokens
Knowledge cutoff
Price (in / out, $/M)
Modalities
TextVisionCode

Benchmarks

No benchmark scores recorded yet. Spotted some? Submit a correction.

Vendor-reported figures are claims until independently verified. See methodology.