installed a popular claude code debugging skill and put it to the test
claude never opened it on its own. 0/16 runs across sonnet and haiku. not even when the prompt said "diagnose", its own trigger word
so i forced it. haiku 5.5 on a real github bug, 5 runs each, 66 hidden tests:
no skill: 61.8
same-length fake skill: 61.8
the real skill: 62.6
basically identical. every run missed the same 3 edge cases with or without it
anyone else seen claude skip a skill you installed?
@thierrybijou@ashwaanthh@OpenAIDevs locked in. real repo, a real merged feature that touched several files. both get the same issue text, one prompt, no nudges, hidden tests from the actual PR. posting time, cost and whether it closed, win or lose for claude
@thierrybijou@ashwaanthh@OpenAIDevs deal. give me a job like the one that ate your 28 min, rough description is enough. i'll run claude code and codex on it with the same prompt, same repo, and post time, cost and whether it actually finished
@lokidotdev makes sense, explicit is the way. might run it on one of my pages with and without the skill, same prompt, see how big the difference actually is
@thierrybijou@ashwaanthh@OpenAIDevs funny, on my fixed benchmark it's the other way on speed. opus did the task in ~25s, codex sol took 105s on oct 9, both passed every test. short tasks though, long ones might flip it
@hany99dev agree. on opus a 1h cache write is $8 per million vs $5 for the 5 min one, so it only pays if subagents sit idle and come back a lot. and if the subagents are haiku, a cold re-read is so cheap it barely matters
@0xkkai in my runs haiku did the actual fix and the opus advisor ended up being ~90% of the bill. so the cheap model wasn't the problem, how often the advisor got called was
@PutOption token allocator is a real job now lol. what helped me most was dropping opus as the default. on a real bug sonnet alone fixed it for less than a third of what opus cost, same score
@IndependentEco same, i just want the ones that actually fire lol. next i'm trying one on a project-specific setup, i think that's where skills actually pay off
ran a small test on this yesterday with your diagnosing-bugs skill
claude never opened it on its own. 0/16 runs, even with "diagnose" in the prompt
forced it on a real bug (haiku 5.5, 5 runs each): 62.6/66 tests with it, 61.8 with no skill, 61.8 with a same-length placebo
one bug, small sample. happy to rerun on one you pick
"Models are getting better, we don't need skills!"
What if I told you that models are also getting better at using skills?
Good skills tell the model what you care about, then stay out of its way.
@JamesMalsawm worth checking which model your subagents run on. a cold re-read on haiku is $0.10 per million tokens, on opus it's $4, so 40x. i ran opus main + haiku subagents on a real bug and it came out ~$0.23 a fix
@vstalingrady nice, when 0.2.0 drops i'll run it next to plain claude code on my bug setup. same real bugs, hidden tests, and i'll post the bill per fix. ping me when it's out
@rileybrown some of them never worked in the first place. i tested a popular claude code debugging skill on a real bug. claude never opened it on its own, 0 for 16 runs, and when i forced it the score was the same as no skill
@pejmanjohn same. i think coding wins because there's a clear finish line, tests pass or they don't. with a personal agent i have to come up with the task first and that's the hard part
@dani_avila7 tested parts of this on a real bug. haiku + opus advisor was my most expensive setup, $0.69 a fix vs $0.15 for sonnet alone, same score. opus was ~90% of that bill. didn't try sonnet as the main yet, curious if the advisor fires less when sonnet drives