@thsottiaux Who thought disabling Cmd+C in the Codex CLI was a good idea? 😅
It’s such a small thing, but it makes the workflow surprisingly inconvenient. Would love to see the usual copy shortcut restored.
Based on published runs of the same quantization on a 1080 Ti and a 3080 Ti, I expect the 2080 Ti to generate around 35–45 tokens per second. My estimate is roughly four seconds per classification: about 850 checks an hour with one request running at a time, or 2,400–3,300 hour
The time I’ve been waiting for has finally come! The new Bonsai model, based on Qwen3.8 27B, is good enough to replace DeepSeek and runs on an affordable graphics card.
I’m planning to use Bonsai for the first pass of my safety checks.
Today, we’re announcing Ternary Bonsai 2 27B.
Based on Qwen3.8 27B, Bonsai 2 27B is 9x smaller than its full-precision counterpart while retaining 98.2% of its aggregate benchmark performance.
Two months after the first Bonsai 27B release, the biggest change is quality. The footprint remains 5.9 GB, but the gap to full precision has narrowed materially, with particularly strong gains in agentic coding, multimodal reasoning, and long-horizon tool use.
Ternary Bonsai 2 27B is available today under Apache 2.0.
A conventional 4-bit version would need more than 15 GiB for the weights alone. I haven’t benchmarked the 2080 Ti yet, so the performance numbers below are estimates. On the Mac, I measured 20 tokens per second when generating output and 144 when processing the input prompt.
It’s a fairly ordinary machine: a six-core i5-12600, 32 GB of DDR4, and a single RTX 2080 Ti with 11 GB of VRAM. That GPU came out in 2018. PQ2_0 brings the model’s weights down to 6.7 GiB, making it possible to fit a 27B model on the card.
With a hard timeout, that smaller spread matters more to me than a faster typical response. Next, I’m moving inference from the Mac to my own GPU box. The aim is to lower the cost per call, keep first-pass screening data on my hardware, and have more control over response times.
Both incorrectly flagged one of the 13 negative cases. Bonsai followed the JSON schema on all 49 calls; DeepSeek broke it three times. Bonsai was slower on the median request: 12.1 seconds versus 9.6. But its slowest response took 19 seconds, compared with 83 for DeepSeek.
Same system prompt, same temperature, same 49 test cases drawn from real production transcripts and curated negative examples. The models agreed on 78% of cases. Of the 36 cases that should have been flagged, Bonsai caught 23 and DeepSeek caught 18.
Keeping the first pass cheap is what lets me check every message. Before switching, I tested Bonsai 2 27B locally on my M1 Max, using PQ2_0 quantization and llama.cpp, against deepseek-v4-flash. Both had to return JSON matching the same strict schema.
That’s where most of the work happens: a cheap classifier screens every message, and a more expensive model reviews anything it flags. The second model costs several times more per output token, but it handles well under 0.1% of traffic.
@steipete@mah_moniem Agreed. I got blocked on Claude for no clear reason, switched to Codex — and now I’m grateful.
With Claude, I needed 2–3 iterations per feature to make it stable. Codex gets it right first try, and it’s $100 instead of $200.
I didn’t believe Peter’s posts before. Now I do.
@bcherny@bcherny
I’ve been using Claude Code for ~a year. Never violated any rules.
Today I got blocked.
Appeal flow:
- Enter email
-Enter organization ID — which you can’t access because you’re blocked
How does that make sense?
AI agent generated 13 content blocks.
With AiTML + PostgreSQL:
→ Query all maps across all docs
→ Full-text search on block.text
→ Update one block, not the whole doc
→ GIN index on https://t.co/cxuZlUwflF
Each block = a row. Each doc = a query.
https://t.co/MImJUMo3ZA
If your AI content format uses a nested tree, you're optimizing for humans, not agents.
AiTML = flat arrays:
→ LLMs stream one block at a time
→ PostgreSQL: one block = one JSONB row
→ Frontends: .map(), no tree traversal
Flat is a design decision.
An AI-generated map in AiTML = one JSON block: type, center coords, zoom, markers with labels + a text fallback for any LLM.
Not a prompt. Not HTML. A typed block any frontend can render.
14 block types. Zero ambiguity.
https://t.co/MImJUMo3ZA
Every AI framework solves how to call tools.
None solve what the agent returns.
7-day itinerary with maps and bookings — now what? A markdown blob?
AiTML: typed blocks that DBs store, frontends render, agents read.
https://t.co/MImJUMo3ZA