← All posts

Bonsai 2 27B on iPhone: a 27B model squeezed to 5.95 GB, and why it still doesn't fit

PrismML's ternary Bonsai 2 27B packs Qwen3.8-27B into a 5.95 GB file that runs on a laptop — but even the 12 GB iPhone 17 Pro tops out around 5.07 GB of weights, and the files need a custom llama.cpp build that stock apps don't have.

Bonsai 2 27B is one of the most-downloaded new models on Hugging Face this month. Its GGUF repository shows about 3.7 million downloads since PrismML put it up on September 16. The pitch is easy to repeat: a 27-billion-parameter reasoning model, shrunk from about 54 GB to 5.95 GB, keeping 98.2% of the original's benchmark score. "27B on your phone" is the obvious next thought.

Short answer: no iPhone can hold it. Not the 8 GB ones, and not the 12 GB iPhone 17 Pro either. It misses the biggest iPhone by less than a gigabyte, which is exactly why it is worth doing the arithmetic instead of guessing. The smaller Bonsai models from earlier this year are a different story, and we get to them at the end.

What actually shipped

Bonsai 2 27B is not a new model trained from scratch. It is Qwen3.8-27B — Alibaba's 27B hybrid-attention model, about 75% linear attention — with its weights re-encoded by PrismML. The model card lists 27.36B parameters in total: 24.35B in the language backbone, 2.54B in the embeddings and output head, and 0.46B in an optional vision tower. License is Apache 2.0.

The trick is the weight format. Every weight is one of three values, −1, 0 or +1 ("ternary"), with one 16-bit scale shared by each group of 128 weights. That works out to 1.72 bits per weight across the whole model. For comparison, the Q4_K_M files most phone apps use sit around 4 bits per weight or a little more — see our quantization explainer for what those letters mean. The weights are also stored in a rotated basis (a Hadamard transform), and the runtime has to apply the matching transform to activations. That detail matters later.

PrismML ships two GGUF packings of the same ternary weights:

FileBits per weightSize
PTQ1_0 (dense trits)1.755.95 GB
PQ2_0 (2-bit slots)2.137.21 GB
MLX 2-bit (Apple Silicon)—8.60 GB
Vision add-on (mmproj Q8_0)—0.63 GB

The GGUF sizes are from the model card; the MLX size is the model.safetensors file in PrismML's MLX repository.

Does it fit your phone

We use the same rule on this blog as the app uses on the phone: plan for no more than 55% of the device's physical RAM, keep a 20% safety margin inside that, and set aside 600 MB for everything a loaded model needs besides its weights (context cache, compute buffers, the app itself). The full reasoning, including the phone restart that taught us to be this careful, is in how much RAM an LLM needs on iPhone. That leaves this much room for weights:

DeviceWeight ceilingPTQ1_0 (5.95 GB)PQ2_0 (7.21 GB)
4 GB — iPhone 12/131.29 GB✗ too big✗ too big
6 GB — iPhone 12–14 Pro, 14, 152.23 GB✗ too big✗ too big
8 GB — iPhone 15 Pro, 16, 16e, 173.18 GB✗ too big✗ too big
12 GB — iPhone 17 Pro, 17 Pro Max, Air5.07 GB✗ too big✗ too big

The smallest file, PTQ1_0, is 0.88 GB over the ceiling of a 12 GB iPhone. By the same rule, it would need a device with roughly 14 GB of RAM (5.95 GB + 0.6 GB overhead, divided by 0.44). No iPhone has that.

Could an app just take more memory than that and hope? iOS lets an app with the right entitlement ask for a lot. But the kernel, SpringBoard and the GPU need those same gigabytes, and iOS starts killing processes before an app reaches its documented limit. Our rule is cautious on purpose. Getting it wrong costs a restarted phone, not a slower answer.

The second problem: stock llama.cpp can't read these files

Even with the RAM, a normal llama.cpp-based app could not run this model. PrismML's model card says it plainly: stock llama.cpp rejects PTQ1_0 and PQ2_0 as unknown types. It also says that stock llama.cpp will load the Q2_0 variant without any warning and produce garbage, because it doesn't know to apply the Hadamard transform. You need PrismML's llama.cpp fork (which does publish an iOS XCFramework) or their MLX fork.

That matters for iPhone apps. An app has to ship the fork's kernels, not only download the file. Pocket AI uses standard llama.cpp and does not ship PrismML's kernels, so it does not offer any Bonsai model today. Any app that lists one should be using the fork. If it isn't, the model card tells you what you'll get.

The benchmark claims

PrismML's own table, run in thinking mode on 14 benchmarks: Qwen3.8-27B at full FP16 precision averages 86.32; Bonsai 2 27B averages 84.78 (98.2% of it) at 5.95 GB. A conventional 2-bit build of the same base model (IQ2_XXS, 7.27 GB) averages 72.59, and a 4-bit build (UD-Q4_K_XL, 17.6 GB) averages 85.18. The biggest losses are in knowledge and reasoning (85.55 → 79.86) and vision (71.36 → 66.19). Math and coding stay almost level.

These are the vendor's numbers on the vendor's setup (EvalScope and vLLM on an H100). We have not reproduced them. What is independently checkable is the file size, and that part holds up.

Two practical notes from the model card. First, it thinks by default at its highest reasoning effort, and PrismML's example command reserves 16,384 tokens for the answer because a small cap "ends generation mid-thought." Second, their fastest Apple number is about 47 tokens per second on an M5 Max laptop, measured on an earlier build. Neither says much about a phone, where memory bandwidth is far lower and Pocket AI caps context at 4,096 tokens anyway.

What does fit: the smaller Bonsai models

PrismML released smaller ternary models in April: Ternary Bonsai 8B, 4B and 1.7B. These are not the new release, and they also need the PrismML fork. But their sizes do fit phones. From their Hugging Face repositories, the PQ2_0 files are:

ModelPQ2_0 sizeSmallest iPhone that fits
Ternary Bonsai 1.7B0.46 GB4 GB
Ternary Bonsai 4B1.07 GB4 GB
Ternary Bonsai 8B2.18 GB6 GB (just — the ceiling is 2.23 GB)

An 8B model in 2.18 GB is the interesting one for phones. It is a slower-moving story than Bonsai 2, though, and it depends on the fork's Metal kernels arriving in apps.

Should you care

If you have an Apple Silicon Mac with memory to spare, Bonsai 2 27B is worth trying. Macs are what PrismML measured on, and it is a 27B-class model in a 6–7 GB file.

If you only have an iPhone, skip it. It doesn't fit any model, and it would need a runtime most apps don't ship. For the best model your phone can actually run today, see which AI models run on your iPhone. On a 12 GB iPhone 17 Pro, that means 8B-class models like Ministral 8B, a 4.5 GB download that fits under the 5.07 GB ceiling.

The real signal here is the format, not this model. 1.72 bits per weight with a small score loss is what could bring 8B–14B models to 8 GB phones. We will write it up when a phone app can actually load one.