Bonsai 2 27B
AvailablePrismML's flagship compressed model, released 2026-09-17 under Apache 2.0. Bonsai 2 27B is a ternary (1.58-bit) compression of Alibaba's open-weight Qwen3.8 27B: instead of storing each weight in 16 bits, PrismML's "ternary" scheme restricts weights to three values (+1, 0, -1), shrinking the 27.8B-parameter multimodal model to just 5.9 GB — a 9-10x memory reduction versus full precision — so it can run on a PC and, potentially, a high-end smartphone. Across a 20-benchmark suite spanning reasoning, math, coding, instruction following, vision and agentic tool use, it reports an aggregate score of 83.9, retaining 98.2% of Qwen3.8 27B's 85.4 (up from the first Bonsai 27B's 95% in July). It reaches up to ~143 tokens/second on an NVIDIA GeForce RTX 5090 and is optimized for low-latency local inference on consumer CPUs and edge GPUs, targeting local assistants, multimodal agents, private knowledge workflows, computer-use apps and long-running agentic tasks. Trained/prepared using Google v5 TPUs. Weights are free to download on Hugging Face (prism-ml/bonsai-2). Benchmark figures are vendor / self-reported.
Specifications
- License
- Open source · Apache 2.0
- Weights
- Downloadable
- Architecture
- Dense
- Parameters
- 27.8B
- Context window
- — tokens
- Max output
- —
- Knowledge cutoff
- —
- Price (in / out, $/M)
- —
- Modalities
- TextimageCode
Benchmarks
No benchmark scores recorded yet. Spotted some? Submit a correction.
Vendor-reported figures are claims until independently verified. See methodology.