Open-weight language models keep getting larger, but a quieter countertrend asks a different question: how small can a capable model get before its intelligence falls apart? A recent release from Prism ML, called Bonsai 27B, pushes that question further than most. It compresses a full 27-billion-parameter reasoning model into ternary and 1-bit weights, turning a roughly 54 GB download into a model that runs on a standard laptop, with a compact variant sized for a flagship phone.
The approach builds on an idea Microsoft researchers formalized in 2024 with BitNet b1.58: large language models can represent most of their weights using only three values, -1, 0, and 1, without collapsing. What makes Bonsai 27B notable is that it applies this idea end to end across a modern, multimodal, thinking-capable model rather than a research prototype.

What 1-Bit and Ternary Quantization Means
Most large language models store their parameters as 16-bit floating-point numbers, abbreviated FP16. A 27-billion-parameter model therefore needs roughly 54 GB just for its weights. Quantization reduces that number by storing each weight with fewer bits.
Conventional quantization usually stops at 4-bit or 2-bit precision, where weights still take numeric values within a range. Extreme quantization goes further. A ternary model restricts each weight to one of three values: -1, 0, or 1. A binary model uses only -1 or 1. In a 1-bit LLM, most of the network’s floating-point multiplication becomes simple addition and subtraction.
The practical bit counts sit slightly above the raw label because a small 16-bit scaling factor is stored for each group of 128 weights. The ternary build works out to about 1.71 effective bits per weight, while the binary build is about 1.125 bits per weight.

From 54 GB to 7 GB: Inside Bonsai 27B
Bonsai 27B is based on Qwen3.6-27B, an open-weights multimodal model from the Qwen team. The developers quantized the language model end-to-end, across embeddings, attention projections, MLP projections, and the final output head, with no higher-precision “escape hatch” hiding behind a low-bit label. The separate vision tower ships in compact 4-bit precision.
The release offers two operating points:
- Ternary build (1.58-bit): about 7.2 GB deployed, retaining a reported 95 percent of the FP16 model’s intelligence, at a true 1.71 bits per weight.
- Binary build (1-bit): about 3.9 GB deployed, roughly 14 times smaller than FP16, at a true 1.125 bits per weight.
| Build | Weight values | Approx. deployed size | Reported avg. score (15 benchmarks) |
|---|---|---|---|
| FP16 (baseline) | 16-bit float | ~54 GB | Reference (100%) |
| Ternary Bonsai 27B | -1, 0, +1 | ~7.2 GB | 80.49 (95%) |
| Bonsai 27B (1-bit) | -1, +1 | ~3.9 GB | 76.11 (89.5%) |

The distinction matters: the binary 1-bit model is the smallest option, while the ternary model is the quality-focused one behind the headline 95 percent claim. Both are published in GGUF format for llama.cpp, with an MLX companion for Apple Silicon.
How Much Intelligence Survives the Shrink?
Extreme quantization has historically traded away capability, especially reasoning. According to the model card, the ternary build averages 80.49 across 15 thinking-mode benchmarks, which the developers report as 95 percent of the FP16 reference. For comparison, a conventional 2-bit IQ2_XXS build scores 72.73 while taking up more space.
Individual results are stronger in some areas than others. Mathematics lands at 93.40, within two points of full precision, while coding reaches 85.96 and agentic tool use hits 74.01. The binary 1-bit build averages 76.11, or 89.5 percent of FP16, with math at 91.66 and coding at 81.88.
The notable claim is not the raw numbers but where they hold. Thinking, reasoning, and tool-use behavior typically collapse below 4-bit precision in conventional low-bit formats. The developers report that these capabilities survive at sub-2-bit precision here. As with any vendor-reported benchmark, independent verification will matter before the figures are treated as settled. The later Bonsai 2 27B ternary release reports even higher retention, at 98.2 percent of the full-precision benchmark average while running a Qwen 3.8 27B model in about 5.9 GB.
Why Tiny Models Generate Text Faster
The speed gain is not mainly about doing less math. On-device inference of a large language model is usually limited by memory bandwidth, the rate at which data can be read from memory. Every generated token requires reading the full set of weights, so smaller weights mean far less data must move for each token.
Shrinking a model roughly 9 to 14 times cuts that memory traffic proportionally, which translates directly into faster output. The developers report about 26 tokens per second for the ternary build on an Apple M5 Pro laptop. On the CUDA serving path, a lightweight speculative-decoding layer adds a further 1.34x speedup with no loss of quality, part of a wider industry effort to lift token throughput per server.
The Techniques That Make Extreme Compression Work
Bonsai 27B gets its size and speed from several techniques working together:
- End-to-end ternary and binary weights across every language-model component, rather than only part of the network.
- Group-wise FP16 scaling that stores one small scaling factor per block of 128 weights.
- Hybrid attention from the Qwen3.6-27B backbone, where roughly 75 percent of layers use linear attention, keeping a 262,000-token context practical on device.
- 4-bit KV-cache quantization to shrink the memory needed for long conversations.
- Custom llama.cpp kernels that consume packed weights directly instead of expanding them back to FP16.
- Speculative decoding via a drafter layer to speed up generation.
The same playbook extends across a wider family. Smaller Bonsai models ship at several quantization levels, and a separate image-generation release, Bonsai Image 4B, applies binary and ternary weights to a diffusion model built from FLUX.2 Klein 4B. The 1-bit image model weighs about 0.93 GB, and the developers describe it as the first image model in its class to run on an iPhone.
Running 27B-Class AI on Laptops and Phones
The models run through llama.cpp on CUDA, Metal, and CPU, and through MLX for native Apple Silicon inference. The binary 1-bit MLX build lands at about 3.9 GB, small enough to fit an iPhone 17 Pro Max. Everything is released under the Apache 2.0 license, which permits broad reuse.
Apple is paying attention. The company is in early talks with PrismML about running larger AI models directly on the iPhone, according to a CNBC report, a move that could make Siri faster and keep more personal data on the device. The startup’s CEO described the discussions as very early and said it remains unclear where they will lead, while Apple did not respond to a request for comment.
This is part of a broader pattern of AI capability moving onto local hardware. OpenBMB’s MiniCPM5-2B, a 2.6B dense model that tops the Intelligence Index for sub-4B open-weights models, shows that even very small architectures can deliver agentic tool-use performance competitive with models four times their size. Smaller models from teams like Alibaba’s Qwen family have already shown that compact open-weight systems can outperform the assumption that bigger is always better. Alibaba’s XuanTie C950 RISC-V chip recently demonstrated native inference of a 27-billion-parameter model at 30 tokens per second with no GPU required. Audio models are following the same path: a compact speaker diarization model now runs entirely on a phone GPU or in a web browser, on a budget of under 100MB.
A Smaller Energy and Cost Footprint for AI
Extreme quantization has an environmental angle that matters as much as convenience. Fewer bits per weight means less memory, less data movement, and less energy per token. A model that runs on a laptop or phone avoids the energy and hardware costs of a data center inference server.
If 1-bit and ternary models continue to hold their performance, the practical effect is cheaper and more accessible AI and a smaller footprint for a technology whose energy appetite has been growing quickly. The same principle is driving how LLMs shrink without getting weaker across the industry.

Frequently Asked Questions
Is a 1-bit model as smart as the full-precision original?
Not quite. The developers report that the ternary build retains about 95 percent of the FP16 model’s benchmark performance, and the 1-bit binary build about 89.5 percent. That is a small loss in exchange for a roughly 7- to 14-fold reduction in size, but it is still a measurable difference in some tasks.
What is the difference between 1-bit and ternary models?
A 1-bit binary model stores each weight as either -1 or +1. A ternary model adds a third state, 0, which lets individual weights be switched off. The extra state costs a little more space but gives the model more representational flexibility, which is why the ternary build scores higher.
Can these models really run on a phone?
The 1-bit build is about 3.9 GB and is published in formats that fit an iPhone 17 Pro Max. The ternary build, at about 7.2 GB, is aimed more at laptops and single GPUs.
Why does reducing precision make the model faster?
Local inference is mostly limited by how fast the model’s weights can be read from memory, not by how fast they are computed. Fewer bits per weight means less data moves for every token, so output speeds up even though the model does the same kind of work.
The Road Ahead
Bonsai 27B is one of the clearest demonstrations yet that extreme quantization has moved from research papers into usable, multimodal, thinking-capable models. The direction is consistent with what Microsoft’s BitNet work predicted: that most of a language model’s precision is redundant and that the future of accessible AI may run on surprisingly little.
The open questions are durability and verification. Whether ternary and binary models hold up across a wider range of real-world tasks, and whether independent testing confirms the reported scores, will determine how quickly this approach spreads. For now, the idea that a 27-billion-parameter model can shrink to fit a laptop and nearly fit a phone is no longer theoretical.
