A generated title card reading 'Bonsai vs Bonsai 27B', subtitled 'One family name, sixteen times the parameters', above three cards reading 'Context: 1,024 tokens vs 262K', 'LLM weights: 0.43 GB vs 5.9 GB', and 'Target device: smart glasses vs phone, laptop, desktop', with a footer reading 'Vendor-reported figures; the two models are not benchmarked on a common suite.' The OrcaRouter logo sits in the bottom-right.
Guides & Insights

Bonsai vs Bonsai 27B: One Family Name, Sixteen Times the Parameters

Author

Elias Hawthorne

Date Published

Latest models · 20View all models
Benchmarks: Artificial Analysis · updated daily
Back to all posts

Anyone shopping this family by name will get the wrong model, and the reason is that the bare name "Bonsai" is not a model. It is the family, and as of this week the family contains a 2-billion-parameter vision-language model that runs on a pair of smart glasses and a 27-billion-parameter model that runs on a phone, a laptop or a desktop. The first is the 1-bit Bonsai 2B vision-language model announced on 23 September 2026 — a 1.7B 1-bit language model plus a 0.3B 4-bit vision encoder, 1,024-token context, built on Bonsai 1.7B. The second is Bonsai 27B, released on 14 July 2026 under Apache 2.0, built on Qwen3.6-27B with a 262,000-token context and a choice of a 5.9 GB ternary build or a 3.9 GB binary build. They share a name, a vendor, a compression philosophy and almost nothing else. Sixteen times the parameters, two hundred and fifty-six times the context, and two entirely different devices.

This page is for the person who searched for one of them and landed on the other.

The naming problem, stated plainly

PrismML's model menu currently lists Bonsai 2 27B, Bonsai 27B, Bonsai 8B, Bonsai 4B, Bonsai 1.7B and Bonsai Image. The glasses announcement adds a 2-billion-parameter vision-language model on top of the Bonsai 1.7B line. So "Bonsai" alone can mean any of at least four parameter classes and two generations, and the size suffix is the only thing that disambiguates. That is not a criticism of the naming — it is the normal shape of a model family — but it does mean the family name carries no information about what you are downloading.

The useful cut is not generation, it is device class. Everything in the 27B line assumes you have somewhere between 4 GB and 8 GB of memory to give a language model and are willing to spend it. Everything in the 1.7B line assumes you have a few hundred megabytes and no more. That single fact determines which half of the family is relevant to you before any benchmark does.

What each one actually is

• Parameters — 1.7B 1-bit LLM plus a 0.3B 4-bit vision encoder in the glasses model, against a 27B-class model derived from Qwen3.6-27B in Bonsai 27B.

• Context — 1,024 tokens on the glasses model, against a full 262,000-token context on Bonsai 27B.

• Language-model weights — 0.43 GB for the 1-bit 1.7B, against 5.9 GB for the ternary Bonsai 27B and 3.9 GB for its 1-bit build.

• Representation — binary {−1, +1} weights with FP16 group-wise scaling in both, which is the 1-bit recipe the family is built around; Bonsai 27B additionally offers a ternary {−1, 0, +1} build at 1.71 effective bits per weight.

• Target silicon — the Qualcomm Hexagon NPU on the Snapdragon AR1 Gen 1 glasses platform, compiled through a QNN SDK with 1-bit kernel support, against Apple silicon via MLX and NVIDIA via CUDA for Bonsai 27B.

• Decode throughput — 15.36 tokens per second on a 4 GB AR1 Gen 1 test platform for the glasses model, against roughly 11 tokens per second on an iPhone 17 Pro, 87 on an Apple M5 Max and 163 on an RTX 5090 for the 1-bit Bonsai 27B.

Every one of those figures is the vendor's, and the two sets were measured on different hardware against different baselines, so they describe two products rather than two points on a curve.

A screenshot of PrismML's announcement page titled 'PrismML Brings 1-Bit Bonsai Models to AI Smart Glasses Powered by Snapdragon', dated September 23 2026, stating a 2-billion-parameter vision-language model runs locally on AI smart glasses powered by the Snapdragon AR1 Gen 1 Platform with a 1,024-token context.

Where Bonsai 27B wins, and it is not close

Context is the first place, and the margin is not a nuance. A 262,000-token window holds a codebase, a long document, a transcript, or a working agent session with room to spare. A 1,024-token window holds one short instruction and one image. If your workload involves reading anything longer than a page, Bonsai 27B is the only one of the two that can do the job at all, and no amount of clever prompting on the glasses model closes that gap, because the limit is the memory reserved for key-value state rather than the size of the weights.

Capability is the second place. Bonsai 27B's ternary build is reported to retain about 95% of its full-precision baseline across a 15-benchmark thinking-mode suite, and its 1-bit build about 90%, with the largest losses in vision and tool use. Those are vendor figures on the vendor's own suite, and they are still a different order of thing from a 2-billion-parameter model whose only published quality statement is that it achieved "comparative benchmark results" against a 4-bit Qwen 3 1.7B. The glasses model is not being compared to a 27B model by its own maker, because the comparison would not be useful.

Ecosystem is the third. Bonsai 27B ships in GGUF, MLX and AWQ packings with documented runtimes, so there is a path to running it on hardware you already own. The glasses model exists inside a Qualcomm NPU toolchain.

A screenshot of PrismML's July 2026 launch post for Bonsai 27B, titled 'Announcing Bonsai 27B: The First 27B-Class Model to Run on a Phone', listing a ternary build at 1.71 effective bits per weight and 5.9 GB and a 1-bit build at 1.125 bits per weight and 3.9 GB, with a 262K-token context.

Where the 2B wins, and it is the entire point

A 0.43 GB language model fits in places a 5.9 GB one cannot go, and the glasses platform is one of them. That is the whole argument, and it is a stronger argument than the spec contrast makes it sound, because the devices in question have no alternative: a pair of glasses with 4 GB of shared platform memory cannot host Bonsai 27B in any build, and a phone cannot host the ternary build without giving up most of what the phone is for.

There is a second win that gets less attention: the 1-bit representation is not a compromise on this device class, it is the enabling trick. PrismML's own framing is that the 1-bit path delivers roughly the equivalent intelligence of the same model at 4-bit precision while using about a quarter of the memory and generating tokens at more than twice the speed. For an always-on device where every milliwatt and every megabyte is contended, that trade is the product.

So the two models are not competing. The 2B is what runs when there is nowhere else to run. The 27B is what runs when there is.

The tier boundary is the real decision

Almost nobody is choosing between these two. The realistic deployment has both: a small always-on model on the device for the requests that fit in 1,024 tokens, and something larger behind it for everything else. Once you accept that shape, the engineering question stops being "which Bonsai" and becomes "where is the line, and who enforces it".

Enforced in application code, that line becomes a branch that has to be rewritten every time the on-device model changes context length or capability — which, on the evidence of the last three months, is roughly monthly. Enforced as a routing policy, it becomes one rule about context budget and required capability, and the local tier stays a tier rather than becoming an assumption baked through the client. That is the layer OrcaRouter operates at: one endpoint in front of 200-plus hosted models, with failover and the escalation policy expressed in configuration. Neither Bonsai is routed here — both are downloads you run on your own hardware — so the honest framing is that the routing layer covers the tier above the local one, and with a 1,024-token local model that tier is where most of the traffic will land.

A generated two-column scoreboard titled 'Bonsai vs Bonsai 27B — the scoreboard'. The Bonsai 2B VLM column reads: parameters 1.7B 1-bit plus 0.3B vision; context 1,024 tokens; weights 0.43 GB; base Bonsai 1.7B; device smart glasses; runtime Qualcomm NPU toolchain. The Bonsai 27B column reads: parameters 27B on Qwen3.6-27B; context 262K tokens; weights 5.9 GB ternary, 3.9 GB 1-bit; base Qwen3.6-27B; device phone and laptop; runtime MLX, CUDA, GGUF. Footer reads 'Vendor-reported; different hardware and baselines.' The OrcaRouter logo sits in the bottom-right.

If you have to pick one

• You are building for glasses, wearables or any always-on device — the 1-bit Bonsai 2B vision-language model is the only option here, and the 1,024-token context is the design constraint you build around rather than against.

• You are building for a phone — 1-bit Bonsai 27B at 3.9 GB is the largest model in this family that fits an iPhone-class memory budget, and it buys you a 262K context the glasses model cannot approach.

• You are building for a laptop or desktop — the ternary Bonsai 27B build is the quality-oriented choice in the family, and the 1-bit build buys speed rather than capability.

• You are building a hybrid — decide the escalation rule first and the model second. The model you can swap; the boundary you cannot, once it is in the client.

What is worth watching is whether the NPU path that made the glasses release possible extends upward. If 1-bit kernels on mobile accelerators generalise, the 1.7B line gets bigger and the case for a 27B model on a phone gets stronger. Until then the family's two halves stay cleanly separated by device class, which makes the choice easier than the shared name suggests.