ONE JEV SUGGESTION CUT WRONG TOOL LOADS FROM 16.8% TO 7.3% THEN BROKE 7 REQUESTS THE AGENT ALREADY GOT RIGHT
That is exactly the kind of benchmark AI launches usually hide
TypeSafe tested 488 requests against a Hermes roster containing 182 skills
Without Jev, Claude Haiku loaded the wrong skill on 16.8% of covered requests and loaded a skill unnecessarily on 9.8% of uncovered requests
With a Jev suggestion, those rates fell to 7.3% and 4.0%
But the router was not free intelligence
It fixed 37 covered requests and damaged seven requests the agent had previously handled correctly
That is the real agent-routing problem. A cheap model can shrink the tool space before an expensive model thinks. It can also confidently remove the tool that would have solved the task
So routing needs a real NONE option, measured thresholds and logs showing when the suggestion changed the final choice
Cheaper decisions help
Unmeasured decisions just move the failure upstream
screening 200,000 items with claude costs ~$120. anthropic's enzyme run burned 210M tokens. both numbers are right.
the gap: their 950 agents also read literature and reproduced results in long loops. that's research. screening is the cheap part, if you keep the workers stupid.
stupid means: read a slice of 200, return JSON, die. no memory, no chatter, no idea other workers exist.
4 rules. the 3rd one is where the $120 comes from:
- SLICE. 200 items per call, one verdict per item: match, confidence, evidence, why_not. the prompt is the whole system, not the orchestration.
- STRIP. no tool loop, no shared state, no agent-to-agent talk. every call runs the same prompt, so all 200,000 get the same judgment.
- CAP. effort: low. ~150 tokens per item × 200 = ~30k input per call. ~1,000 calls = ~30M tokens ≈ $120 at $4/M. screening is pattern work, you don't pay for reasoning.
- GATE. keep confidence ≥ 0.6, then tune it on a labeled sample. starting line, not a law.
cost: a stupid worker misses what a smart one would catch. that's what ranking and the human check are for.
full funnel + filter prompt in the article below.
one smart agent reading all 200,000, or 1,000 dumb ones that never talk?