Agnes 3.0 Flash
AvailableA fast multimodal model from Singapore's Agnes AI, surfaced mid-September 2026 (model card and provider coverage around Sep 14). This row reflects the disclosed open-weights PREVIEW checkpoint (Agnes-AI/Agnes-3.0-Flash on Hugging Face, Apache 2.0): a 33B-parameter model with a 262,144-token context, text/image/video input and text output, using a hybrid-attention architecture that mixes recurrent and standard attention to hold memory down at long context β of 72 decoder layers, 54 run a gated delta rule (a recurrent mechanism whose per-layer state does not grow with sequence length) while 18 use standard global grouped-query attention (24 query / 4 KV heads) and are the only layers that accumulate a KV cache. At bf16 it needs roughly 66 GB of disk and a single H100/H200-class GPU, and ships custom modeling code (trust_remote_code=True). Note the production "Agnes 3.0 Flash" served through Agnes AI's API is a different checkpoint with a 1M-token context window; the specs here are the open-weights preview. Vendor-reported figures, unverified independently at launch.
Specifications
- License
- Open weights Β· Apache 2.0 (preview checkpoint)
- Weights
- Downloadable
- Architecture
- hybrid
- Parameters
- 33B
- Context window
- 262K tokens
- Max output
- β
- Knowledge cutoff
- β
- Price (in / out, $/M)
- β
- Modalities
- TextVisionCode
Benchmarks
No benchmark scores recorded yet. Spotted some? Submit a correction.
Vendor-reported figures are claims until independently verified. See methodology.