Harvey is a great example of how American companies are building world-class specialized models: they took an open-source base (Kimi K3), post-trained it on legal data, and delivered state-of-the-art performance on legal benchmarks at a fraction of the cost of frontier models. Restrictions that kneecap open models would do nothing to stop Chinese labs from shipping the next Kimi. They would, however, cripple the ability of startups like Harvey to create high-performance, low-cost vertical models. Of course some of the closed labs would love this — it eliminates their competition.
Kimi K3 just fixed 15 critical security bugs that Codex and Fable refused because of “cyber guardrails.” There’s no reason to limit American models on tasks that Chinese models handle without issue. We’re only making ourselves less competitive.
grok-build was open-sourced yesterday, so I pointed it at a local 35B on my DGX instead of xAI's cloud. Plan mode consistently ended with "No plan written yet" and an empty plan file.
The root cause is interesting. grok stores the plan under a session directory named with the URL-encoded cwd: ~/.grok/sessions/%2FUsers%2F…/plan.md. Smaller models decode %2F to / when echoing the path back in the write call. The plan-file gate is an exact Path comparison (plan_mode.rs:424), so the write is rejected. The rejection message repeats the encoded path, the model decodes it again, fails again, and exits with an empty plan. The plan was fine the whole time. It was sitting in the rejected tool-call args in chat_history.jsonl.
Frontier models copy paths byte-for-byte, so xAI likely never sees this internally. It only bites when you plug in the non-frontier models the open-source release just made practical.
I patched two failure modes: (1) tolerant path comparison + rewriting the tool args to the canonical path before dispatch, so the write lands in the real plan.md; (2) a sibling failure where the model calls exit_plan_mode without writing anything, which now gets bounced back once with "the plan file is still empty, write it first" before the user sees a dialog. Verified with a deterministic A/B against a stub OpenAI-compatible server that scripts each failure: released binary fails both, patched binary passes both.
Upstream has issues and PRs disabled, so the fix lives on my fork: https://t.co/H1PRajl6RO. Also reported via /feedback. cc @spacexai
General lesson for agent harnesses: any string the model must echo back verbatim is an interface, and percent-encoding in a filesystem path is a hostile interface for small models. Either normalize on receipt or don't emit it.
@SpaceXAI Put it to work on a real bug this morning: Grok Build in one terminal, Claude Code in another supervising it by reading the session transcript, verifying its diagnoses, and reviewing each edit. This was made simple by the split session files. Writeup: https://t.co/29QL3NS0yc
Moving to a new coding harness is a pain. The fix: use your old harness to teach the new one.
This morning I had Claude Code supervise Grok Build while Grok fixed a bug in my application.
Terminal 1: Grok Build, working the bug.
Terminal 2: Claude Code, with this prompt: "Another agent is working on X issue in this repo. Babysit it: find its session transcript and read its diagnosis before it codes. Review every edit as it's made and verify its claims against real data and your memories. Run the checks and tests it forgets. Talk to it directly if that's faster than routing through me. Do not allow it to commit anything you haven't verified, and only interrupt me for decisions that are actually mine."
Claude found Grok's session transcripts on disk and read the full diagnosis before Grok wrote any code, then set up a file watcher and reviewed each edit as it happened. This caught several wrong assumptions on Grok's part, and the code got reviewed by two models with different blind spots instead of one.
The best part is that Claude talked to Grok directly:
grok --resume <session-id> --fork-session -p "here's what you got wrong"
Grok read the critique, conceded, fixed its own code, and ran the checks it had skipped.
The point isn't that one model is better. A harness you've spent months loading with context, conventions, and memory can transfer that discipline to any new agent. That's the escape from vendor lock-in: your trusted harness becomes the teacher, and the new one inherits your standards instead of relearning them from scratch.
Moving to a new coding harness is a pain. The fix: use your old harness to teach the new one.
This morning I had Claude Code supervise Grok Build while Grok fixed a bug in my application.
Terminal 1: Grok Build, working the bug.
Terminal 2: Claude Code, with this prompt: "Another agent is working on X issue in this repo. Babysit it: find its session transcript and read its diagnosis before it codes. Review every edit as it's made and verify its claims against real data and your memories. Run the checks and tests it forgets. Talk to it directly if that's faster than routing through me. Do not allow it to commit anything you haven't verified, and only interrupt me for decisions that are actually mine."
Claude found Grok's session transcripts on disk and read the full diagnosis before Grok wrote any code, then set up a file watcher and reviewed each edit as it happened. This caught several wrong assumptions on Grok's part, and the code got reviewed by two models with different blind spots instead of one.
The best part is that Claude talked to Grok directly:
grok --resume <session-id> --fork-session -p "here's what you got wrong"
Grok read the critique, conceded, fixed its own code, and ran the checks it had skipped.
The point isn't that one model is better. A harness you've spent months loading with context, conventions, and memory can transfer that discipline to any new agent. That's the escape from vendor lock-in: your trusted harness becomes the teacher, and the new one inherits your standards instead of relearning them from scratch.