LFM2.5-VL-3B
AvailableLiquid AI's edge vision-language model, released Aug 12 2026 β a 3.1B-parameter VLM built on the LFM2.5-2.6B text base with an integrated SigLIP2 400M NaFlex vision encoder. Accepts text, images, and video frames and returns text, tuned for on-device screen understanding, visual grounding, and tool calling. Vendor-reported: 80.7 average on ScreenSpot-v2 screen understanding, 87.9 P@1 on RefCOCO grounding, 59.5 on ToolSandbox function calling, 81.0 on MMBench, and 69.4 averaged across 28 benchmarks. Runs on-device at ~228 tok/s on an Apple M5 Max and ~116 tok/s on an AMD Ryzen AI Max+ 395; supported in llama.cpp, MLX, vLLM, SGLang, and ONNX. Open weights on Hugging Face.
Specifications
- License
- Open weights Β· LFM Open License v1.0
- Weights
- Downloadable
- Architecture
- hybrid
- Parameters
- 3.1B
- Context window
- β tokens
- Max output
- β
- Knowledge cutoff
- β
- Price (in / out, $/M)
- β
- Modalities
- TextVisionVideo
Benchmarks
No benchmark scores recorded yet. Spotted some? Submit a correction.
Vendor-reported figures are claims until independently verified. See methodology.