LLM Releases
← Catalog

Bonsai 2 27B

Available
PrismMLOpen source

PrismML's flagship compressed model, released 2026-09-17 under Apache 2.0. Bonsai 2 27B is a ternary (1.58-bit) compression of Alibaba's open-weight Qwen3.8 27B: instead of storing each weight in 16 bits, PrismML's "ternary" scheme restricts weights to three values (+1, 0, -1), shrinking the 27.8B-parameter multimodal model to just 5.9 GB — a 9-10x memory reduction versus full precision — so it can run on a PC and, potentially, a high-end smartphone. Across a 20-benchmark suite spanning reasoning, math, coding, instruction following, vision and agentic tool use, it reports an aggregate score of 83.9, retaining 98.2% of Qwen3.8 27B's 85.4 (up from the first Bonsai 27B's 95% in July). It reaches up to ~143 tokens/second on an NVIDIA GeForce RTX 5090 and is optimized for low-latency local inference on consumer CPUs and edge GPUs, targeting local assistants, multimodal agents, private knowledge workflows, computer-use apps and long-running agentic tasks. Trained/prepared using Google v5 TPUs. Weights are free to download on Hugging Face (prism-ml/bonsai-2). Benchmark figures are vendor / self-reported.

Specifications

License
Open source · Apache 2.0
Weights
Downloadable
Architecture
Dense
Parameters
27.8B
Context window
— tokens
Max output
Knowledge cutoff
Price (in / out, $/M)
Modalities
TextimageCode

Benchmarks

No benchmark scores recorded yet. Spotted some? Submit a correction.

Vendor-reported figures are claims until independently verified. See methodology.