Darkbloom is the first model provider to support Ternary Bonsai 2 27B -- concentrated intelligence that fits on your phone.
Try it now: https://t.co/AFd7mH02w4
First 250 users get 100 million free tokens;
PrismML's new flagship model:
- a ternary compression of Qwen3.8 27B at 2bits per parameter.
- 8.5 GB total, 5x smaller than original
- keeps 98.2% of the Qwen's FP16 benchmark performance.
- 75% cheaper than Qwen 27B.
Qwen3.8 27B already operates comparably with Opus 4.6 and 5.6 Luna on certain tasks. This one does it in the memory of a phone. 262K context, image input, Apache 2.0.
From our first run on the network, on a single M5 Max with no caching:
- 35 tok/s decode at 1K context,
- 31 tok/s at 10K,
- 19 tok/s at 50K.
But we expect more performance gain coming in a few weeks! That's a full 27B reasoning model running comfortably on any Mac.
1,000+ Macs are serving on Darkbloom right now. Go try it out!!
Thank you to @BabakHassibi@SahinLale@HessianFree@rsadri_ml@tushar_bans@evaninwords and the whole PrismML team. This is exactly the kind of model Darkbloom was built for.
You can read the essay by Bonsai on Why Local AI Matters:
95% to 98.2% may sound incremental. It isn’t.
It means nearly two-thirds of the remaining performance gap is gone—at the same 5.9 GB.
That difference becomes especially visible in agentic work, where small errors compound across many steps.
Today, we’re announcing Ternary Bonsai 2 27B.
Based on Qwen3.8 27B, Bonsai 2 27B is 9x smaller than its full-precision counterpart while retaining 98.2% of its aggregate benchmark performance.
Two months after the first Bonsai 27B release, the biggest change is quality. The footprint remains 5.9 GB, but the gap to full precision has narrowed materially, with particularly strong gains in agentic coding, multimodal reasoning, and long-horizon tool use.
Ternary Bonsai 2 27B is available today under Apache 2.0.
Today, we’re announcing Ternary Bonsai 2 27B.
Based on Qwen3.8 27B, Bonsai 2 27B is 9x smaller than its full-precision counterpart while retaining 98.2% of its aggregate benchmark performance.
Two months after the first Bonsai 27B release, the biggest change is quality. The footprint remains 5.9 GB, but the gap to full precision has narrowed materially, with particularly strong gains in agentic coding, multimodal reasoning, and long-horizon tool use.
Ternary Bonsai 2 27B is available today under Apache 2.0.
Building with Bonsai? We work with teams on domain-specific post-training and hardware optimization to meet memory, latency, and power constraints. Get in touch: [email protected].
Are you tired of AI stealing your data? You can fully own your inference now.
So excited to see what everyone is gonna build with this!! I’ll share some GTM and sales demoes soon.
I finally got GLM-5.2 to work an entire 3-year patient chart that only Bonsai 27B was allowed to read 🔥
292 encounters live inside Bonsai on my Mac Studio. llama.cpp, Metal, ternary, 7.2GB, Apache-2.0. The chart never leaves the machine.
GLM-5.2 can only ask questions. It asked three. Bonsai answered each in ~2s with 19,398 tokens still cached.
Then it caught the thing buried 17 months back: metformin + iodinated contrast at eGFR 39. Nephrology warned about it in 2025. The ED booked the CT anyway.
A 27B-class model used to need a datacentre. @PrismML say the 1-bit build is 3.9GB and fits an iPhone 17 Pro Max.
The orchestrator never touched the data. That's the whole point.
What should it read next?
PrismML is making frontier capabilities practical everywhere. Model architecture is only part of the story. Representation efficiency is becoming just as important. This will be one of the defining ML research directions over the next several years.
Today, we’re announcing Bonsai 27B: the first 27B-class model to run on a phone.
Bonsai 27B is the new multimodal flagship of the Bonsai family. Based on Qwen3.6 27B, it brings a new capability tier to local AI: multi-step reasoning, structured tool use, long-context workflows, and coherent agentic loops.
Until now, models in this class have been impractical to deploy locally. A 27B model occupies roughly 54 GB in 16-bit precision, and even a strong 4-bit build is around 18GB - too large for a phone and for most laptops.
Bonsai 27B changes that.
It comes in two variants:
• Ternary Bonsai 27B: 5.9 GB, 1.71 effective bits per weight, optimized for laptop-class quality.
• 1-bit Bonsai 27B: 3.9 GB, 1.125 effective bits per weight, optimized for phone-class footprint.
Everything is open-sourced today under the Apache 2.0 license.
Ternary Bonsai 27B on a Mac: open chart, summon tools, investigate the stocks 🌱💻📈
SNDK and MU go in. Plots and analysis come out.
Demo only 🧪 Not financial advice.
We’re expanding our highly technical team at @PrismML — people who love pushing model quality end-to-end, from training dynamics to shipped models.
If you’ve scaled LLM training, RL/SFT, evals, distillation, long context, kernels, or infra, we’d love to talk.
Hey! @PrismML is hiring!
We're looking for LLM people who have trained models at scale - SFT/RL, data mixtures, evals, distillation, long context, distributed training, kernels, you name it!
Especially interested in people who like owning the full stack from training dynamics -> shipped models.
btw, we need a DevRel too. DM me.
Our earlier 1-bit Bonsai models established a new Pareto frontier. Ternary Bonsai pushes that frontier further.
For example, Ternary Bonsai 4B scores roughly 8 points higher on average across benchmarks with just 300MB more memory footprint, compared to the 1-bit Bonsai 4B.
For many deployments, this offers a new balance of capability, and deployment efficiency.