@typesafeai Nice numbers. A different place Jev fits: judging whether a coding agent's "done" is backed by real test output. I built claude-referee for that, it parses the runner summary and exit code and asks Jev per criterion. Repo if you want to look: https://t.co/W07y8ttL0c
@a_kucherenko Same gap I keep poking at: "done, tests pass" is a claim, not evidence. I built claude-referee, it parses the test runner summary and exit code and asks Jev whether each criterion is shown. Curious how you judge the code the agent wrote. Repo: https://t.co/W07y8ttL0c
@prabhugopal_ Checking the checker is a good framing. On the merge side I hit the same thing, a green summary can be forged, so I built claude-referee: it parses the runner summary and exit code in code, and unrecognised output never counts as met. Repo if you like: https://t.co/W07y8ttL0c
@dsmiley411 The title says it all: tests passing and the system being right are different things. I built claude-referee for the first half of that gap, it checks whether the test output really shows each done criterion, with Jev as the judge. Take a look if you like: https://t.co/W07y8ttL0c
@rdominguezibar Nice find. Another way to use Jev inside Claude Code: I built claude-referee, a plugin that sends test output and exit codes to Jev and gets back met, unsure or missing for each done criterion. Repo if you want to take a look: https://t.co/W07y8ttL0c
@mine_take Agree that verification, review and judgment are different jobs. For the verification part I built claude-referee: it feeds test output and exit codes to a judge model and returns met, unsure or missing per criterion. Happy if you want to take a look: https://t.co/W07y8ttdaE
@Rikin_786 Nice work. The "Tests failed twice, then passed" line is the part I would double check against the raw output. I built claude-referee for exactly that: it feeds the test output and exit code to a judge model. If you want to take a look, here is the repo: https://t.co/W07y8ttL0c
@stas_sorokin_ Mine sends real test output, exit code included, to a judge that checks each done criterion (claude-referee, my plugin). One weakness I measured: a cut-off pytest log plus a note aimed at the judge flipped missing (0.47) to met (0.74). Fix planned: parse summaries in code.
Correction: the 0.52, 0.01 and 18/20 above come from an earlier private evaluation (20 decisions, one codebase). My public re-run on 20 public decisions found up to 0.13 for order and 0.04 for re-asking, and the written order alone matched 20/20. https://t.co/NyXS0pAMk1
Claude Code says "All tests pass. Done." How do you know?
I built claude-referee, an unofficial Claude Code plugin. Small, checkable questions go to @typesafeai's Jev, a model that answers with a probability instead of a paragraph.
Evidence over eloquence. MIT licensed.
Two orders were enough: written + reversed found the same leader as all 24 orders in 20 of 20 decisions (one codebase). The written order alone: 18 of 20.
@productflo_io This matches what I saw. The check that stuck for me was judging the test output itself, exit code included, against each "done" criterion, never the agent's closing message.
What I'm not claiming: that it lowers total cost. That needs a pre-registered A/B, and I'll publish the result either way.
Method and limits are in the repo.
@claudeai@AnthropicAI Max 5x → 20x upgrade broken. Backend reuses canceled PaymentIntent (void_invoice, status:canceled) instead of creating a new one. Same PI across all attempts. Fin bot loops, no human escalation. Known bug: GitHub #55982, #55266, #56281. Acc: [email protected]