Congrats on the launch. The "decisions instead of output tokens" part is measurable on frozen open models, so I ran it: Qwen3-4B, CLINC150, each schema field as a lettered question, answer read from the logits of the option letters.
Against grammar-constrained JSON on the same model: same accuracy on closed enums, 4× faster on short inputs, ~2× on long ones with a shared prefix. Not your 20×, but that is the free part, no training involved =) Calibration, abstention and dependent fields are where a trained decision model should earn its keep.
Write-up and a small bench: https://t.co/KLbCtWKLfG
Jev-style "typed decisions" without training anything.
On a frozen Qwen3-4B: turn each schema field into a lettered question and read the option letters' logits. No output tokens.
Same accuracy as grammar-constrained JSON on closed enums. 4× faster on short inputs, up to 2.4× on long ones with a shared-prefix cache. Strings and numbers still need generation.
Write-up, figures, teaching bench: https://t.co/KLbCtWKdq8
What stops your AI coding assistant shipping garbage, other than a weak flesh proxy skimming the chat history? Almost nothing, until the harness and the engineering culture around the agent grow to match what it produces.
Today I'm sharing a placebo technique to improve AI coding: https://t.co/EEjabTSVhB
Last week I noticed a cool feature in Claude Code: sessions can now send direct messages to each other using their names (/rename ...).
Locally, it works like this: any session can send a SendMessage targeting another session by name. The target session listens on /tmp/cc-socks/<pid>.sock and injects the message into its context right after its current turn ends.
Previously, you could only inject context during events triggered by the session itself (SessionStart, UserPromptSubmit, or a rejected PreToolUse hook). You couldn't push a message into another session externally on your own initiative. Now you can.
What I tried out:
Knowledge injection. Knowledge bases stored as a collection of files are often ignored by agents. I split mine into atomic chunks that trigger via regex on the PreToolUse hook. The relevant knowledge is then injected directly into the session accessing that specific file or running that targeted operation. Sockets are crucial here because a successful PreToolUse stdout doesn't make it into the context, and we want to inject context without interrupting the command.
Inter-agent interaction. Since parallel sessions often mess up each other's code, I built a harness that blocks writes to the main tree. If an agent violates this rule, it gets a prompt explaining how to spawn its own tree using EnterWorktree. But here’s the cool part: if the agent disagrees with the restriction, the prompt suggests it "file a complaint" with a dedicated harness-manager session, which then resolves the issue. Since rolling this out, parallel agents have already proven 5 times that the rule was too strict!
The big picture: This mechanic points toward keeping a pool of "on-call" sessions. Each one's job is to monitor other agents and handle issues within its domain (like the harness-manager, which gets pinged whenever repetitive machinery needs to be added, modified, or fixed).
P.S. Any program running under the same user as Claude Code can write to the socket (sending a request to the LLM)—socket permissions are 0600.
Your CLAUDE.md rules mostly don't work. A rule holds only if something fails when it's broken.
Seven rungs, strongest first:
1. Regeneration - the edit is erased on next build
2. Hook refuses the tool call
3. Won't compile
4. AST guard
5. Cross-artifact guard
6. Runtime detector
7. Written in a file <- where most of yours are
Give the link to your Claude, point it at your own rules file, and ask which ones need a real guard.
Pulled together what Anthropic, OpenAI, Google and Qwen actually document about prompting, plus 20 papers, and checked it against my own measurements. They contradict each other more than you would expect.
Skill: https://t.co/B82MkJanXg
Analysis: https://t.co/jTO8TOyAIq
Ever tried to remember what that one Claude session is about when you have 10 open?
Use /rename <session-name> manually or ask Claude to add text on the screen to CLAUDE.md.
Now Claude will suggest a ready-to-use session name as soon as the purpose of the session becomes clear.
Much easier to find the right session later.
@ClaudeDevs Analyzed logs of our team using Claude: nobody runs /loop or /goal by hand. But the agents that main sessions spawn use them constantly. Seems like the orchestration primitives aren't for humans - they're perfectly consumed machine-to-machine.
Each fix became a rule, the rules became a standard.
MUST fixes a capability, SHOULD recommends a vendor — so it survives tooling debates. Release-gate checklist included. CC BY 4.0.
https://t.co/VtlBBs8FJE
Converged on different answers? Open an issue.
We run MCP servers for the whole company - an analytics gateway over the warehouse and AI-Lens. Production broke them in ways no demo shows.
Here I want to share issues I've spent most time debugging them.
4. One fat JSON cell blows past the byte limit and the user sees "File content exceeds maximum allowed size" instead of data.
Truncate by rows AND bytes.
Same discipline for instructions: Claude Code truncates them at 2 KB.
7/ Build an ironclad harness, and a local LLM stops being a black box. When something breaks, you find exactly which node failed and patch there.
Full breakdown of my LLM-as-a-judge implementation (architecture, code, and exact invariant logic):
https://t.co/cqDZKNkJGm
1/ Keeping sensitive data local — medical transcripts, raw terminal logs — often means running your own small LLM. Nothing leaves the box.
But small models bluff. I put one in charge of a 36-item grading checklist and it gamed the prompt.
Here's how I fixed the eval harness.
6/ And I keep one strict invariant: the score after verification must equal the final score.
The architecture says everything downstream is cosmetic. Instead of trusting my own code comment, I made it a test. If a summary layer secretly moves a grade, CI fails.