Independent AI engineer, Paris. Coding agents in production, Claude Code training and audits on your own repo. Eight years of ML in prod, ex-Amazon, ex-Akeneo.
@Vtrivedy10 the two worst instabilities i measured on a narrow agent came from a contradiction between the prefix and a tool reply. i found them by reading traces, the scores did not show it. at scale, how do you surface those: clustering trajectories, or an agent reading them?
@Voxyz_ai how do you know which subagents will idle past 5 minutes before paying the 2x write? on a narrow agent i measured, explicit caching came out a bit slower than inline, so i check first. do you read the idle gaps from the transcript? might not transfer
@dani_avila7 on a narrow agent i measured, a refusal ended the call while a plain fact in the tool reply let the model fix itself. when your read-only guard blocks a write, does the subagent get a reason it can act on, or just a deny?
@pauliusztin_ with 19 tasks and a 0/1 verifier, one task flipping moves the score by about 5 points. do you rerun each task a few times before comparing two prompts, or accept that a one-task move is noise? for me a single run said little, might not transfer to coding agents
@danielrupawalla how many samples per task do you use to place it in the 0.2 to 0.6 band? on a narrow agent i measured, 8 runs separated two arms but gave only a rough level, so i wonder how tight a band can really be.
@Vladic_ETH is the $7.67 vs $3.46 averaged over several runs per task? on a narrow agent i measured, 8 runs could tell two arms apart but gave a poor level, so i would read the spread before the gap. might not transfer.
@sermakarevich how many runs per task sit behind the every-time number? on a narrow agent i measured, 8 runs could tell two arms apart but said little about the level itself, so how do you read a task that passes 4 runs out of 5? might not transfer
@samongaro_ on a narrow agent i measured, two rules went from 0/20 stated abstractly to 20/20 with one worked example written next to them. have you tried putting an actual short reply you'd accept right under the rule in claude.md? might not transfer to claude code
@alexgetmancom how did you see that the 5 minute expiry was costing you on subagents? one that finishes in a minute or two never hits it, so is it only the long idle gaps between calls, or did you spot it in the cost per session?
@fleyta88 how do you tell from the traces that a turn was a dead one and not one that needed the bigger model? in my experience a trace shows where the spend went, and i only saw if a cheaper model reached the same result by replaying the same task on both
@0xShoopy do you run every model x effort cell on all 20 tasks? i measured a narrow agent where 3 runs on everything, then 8 only where the two arms diverged, cost $1.56 instead of $3.26 for the same verdict. might not transfer to support tickets
@aditlal 24,000 prompts in, how do you tell that a hook or the second model reviewer paid off? on a narrow agent i measured, the same commit gave 24 fails out of 200 one morning and 39 that evening, so i only trusted a change against a control run in the same minutes. might not transfer
@cyrildelattre with 4 runs on 46 scenarios, how do you tell 83 from 87 on policy compliance? on a narrow agent i measured, only a gap like 0/8 vs 8/8 on the same replay felt safe to call (fisher p = 0.000155), smaller ones drifted. might not transfer to support agents
@RajeswarSai with 57 criteria that all must pass, how many times do you rerun the referee on the same transcript before you trust a verdict? i measured a narrow agent where the same commit gave 24 fails out of 200 one morning and 39 that evening. might not transfer to graders
@Damir_Akaza does quoting the line in the reason matter, or would a bare "not done" send it back just as well? i measured a tool error with one more sentence on a narrow agent, 0/8 to 8/8 on replay, so i suspect the quoted line does the work. might not transfer to stop hooks
@vesko_st on the six tasks, how many held-out examples per task before you called opus equal or better than roberta? in my experience a small gap only means something against a control run in the same minutes, since one run on a stochastic model says little. might not transfer
the fix was not a 28th. type the argument, let the model pick from fields. when the model gets a form wrong, give it a dumber argument, don't parse harder