@K_L_M Local only is the right call. In MCP mode scrub_text takes the raw text as an argument, so the model has already read the dump by the time it calls the tool. Passing a path to scrub_file is what keeps secrets out of the context. Would you hide scrub_text from agents?
Every wrong value was a real number from the same text in the wrong unit. ARR instead of MRR 8 times out of 8. So the unit conversion belongs in the engine too. It is one probe per cell on synthetic tasks, so only the big gaps hold up.
https://t.co/yMTfNdEDws
Moving arithmetic into code is a 2022 idea (PAL, PoT). I measured what it costs in my stack. 330 of 330 calls are clean while every parameter is named in the engine's units. When the input has to be derived first, extraction scores 71.5% and the model computing scores 75.6%.
On MRR the code removes 18 to 24 multiplications in a row. Extraction gains 43.6 points there. On payback it removes one division but adds turning a quarterly payment into a monthly one. That costs 36.4 points.
Where the tool returned nothing, 23 of those 35 answers came back empty too. So I can't separate a model that hid the failure from one that had nowhere to put it. The judge is deterministic code against what the tool actually returned.
https://t.co/4rWjg9OpNw
ToolMaze, SilentProbe and Guardrails as Scapegoats all showed this year that agents pass broken tool output along. I ran a small replication and one prompt line mattered most: asking for exactly one line cut error reporting from 91.4% to 37.1%. Same tool, same 7 models.
@jonringer117 Read the PR, nice that the log tool says how many lines it cut. One catch: every tool returns String, so "Error: ..." goes out through rmcp as a success with isError unset. Returning Result<String, String> lets rmcp flag it. Then a missing log stops looking like a log.
@PiGCodingAgent The divergence ledger is the best part of this release. 31 active, each with a marker at the call site and a test. How do you decide when an upstream bug is part of the contract?
@fframes_rust Nice speedup. I read the prebuilt script in your FFmpeg crate: the archive comes from the 9.0.0 release via curl, is unpacked with no checksum and cached for every project on the machine. Is a pinned sha256 per feature key planned? A swapped asset would then fail the build.
38.5% of npm packages declare 0 dependencies. Another 28.2% declare 1 or 2. Only 2.2% declare more than 25.
Dependency hell is that 2.2%, counting runtime dependencies of the newest version. The rest of the registry has almost nothing underneath it.
X ships its For You ranking weights with a warning above them: they scale the predicted probability of your own actions, so "1 report cancels 468 likes" does not follow. That claim is the headline of an explainer about the release. I cannot tell which came first.
Cold and warm point reads on the same Postgres index differ by 47x: 236us against 5us. A read benchmark that does not say which it ran has said nothing.
I re-ran uuid v4 against v7 at 20M rows. Write gap 2.8x, confirmed. Read gap: none, cold or warm.
https://t.co/wC8BhPWkVc
@omarsar0 Scoring a harness on downstream feedback assumes it notices its own failures. I took the tools away and asked for something only a tool could know. 59 calls, 3 models, not one reply said the tools were gone. Haiku 4.5 invented a value 14 times of 20.
https://t.co/NEs7mIhbMh
@delali The number you will see for npm install hooks is 6.36%. It is 3.66x too high. I crawled all 4,296,340 packages: 74,664 run code on install, 1.74%. The gap is prepare, which no tarball install runs. Your own tree is a different count.
https://t.co/I8tPE1huqv
From the same run: CPI-U level for Dec 2015 came back right from memory in 13 of 20 cells with the tool down. That value has never been revised. I have one revised series against one unrevised one, so this shows the mechanism and not how often it fires.
https://t.co/wmBWSjJqDo
All 4 models answered 3.7% for US unemployment in June 2019. BLS returns 3.6% for that month today. 3.7% is what got printed at the time and the series was revised afterwards. The model is quoting an old print. Nothing here was hallucinated.
Correction to the card above. The 30-50% cell for Opus 5 was empty because I had not run it. Now I have: 10 of 10 caught, 0 false alarms on the matched clean set. That also takes the clean total from 110 to 120.
Sonnet 5 and Haiku 4.5 summarize a P&L that has 1 planted bad subtotal in it. Neither catches it in 10 tries. Both quote the wrong figure. The same documents with the error 5x bigger get Sonnet to 6 of 10 and leave Haiku on 0.
6 words move them. "and check whether it adds up" takes both to 10 of 10, with 0 false alarms on 110 clean reports. Opus 5 and Fable 5 need no hint, but it is 1 defect type on synthetic books. FinVerBench does this at bigger scale on real filings.
https://t.co/3KTZ1camfJ
I pressed a button on my Galaxy Watch and asked a question out loud on a walk. A model on my laptop read the answer back. $0 in API.
The watch recognises speech on device. A dependency-free Node server runs headless claude -p on my subscription.
https://t.co/yy8ySqm6BE