@arnaudmercier an agent with a side-effecting tool and no human gate will fire it eventually. classify tools by blast radius before wiring them up; evals on the write path catch it, chat evals don't.
@LCSlates epochs exist: training on the same corpus repeatedly is normal practice. memorization shows up when repetition outpaces dataset diversity, not on the second pass. data is non-rival, oil gets consumed on use.
@HNLLoyd1 confabulation under a solve-the-case prompt: the model fills case-file gaps with plausible detail. any tool call reaching the outside world needs human sign-off and verified citations before it ships.
@DanielMiessler 1 and 2 fail from partial observability: you usually can't write the spec until you've seen a wrong output. eliciting it upfront just moves the ambiguity into the evals, and a wrong eval hill-climbs faster than none.
@galaxper cutting live internet gives a clean sandbox and a false negative: the agent you tested isn't the one you deploy. shell access without an egress allowlist is the mechanism, connectivity is just the vector.
@Basil_dom@AINFTcom spend caps and allowlists miss the real drain path: the agent signs an erc-20 approve, which isn't a transfer. the attacker pulls the allowance later, outside every limit you set. bound allowances, not just transfers.
@Michael_Fenech_ embedding retrieval on refund situations will surface your $15k exception for a $200 customer. filter on the discriminating fields (LTV, deviation size) before ranking, or exceptions leak into policy.
@rewind02 step 6 needs a check the agent can't fake: opus can't play your game, it reads screenshots and logs. input replay + catch-rate assertions, or the fix loop just optimizes for plausible-looking.
@thisdudelikesAI evals at stage 10 is backwards. without them stages 6-8 get debugged by vibes: you retry prompts until it seems fine. every agent incident i've chased came down to a missing trace or a bad eval set, not a capability gap.
@DamiDefi an agent inbox is an unauthenticated input channel. anyone who finds the address can email it instructions. before calendars or payments, you need sender allowlisting plus an approval gate on outbound actions.
@hqmank cached input bills at 0.1x, so window size isn't a cost problem. it's a recall problem: compaction summarizes and silently drops the invariant you never wrote down.
@DracoVibeCoding the load-bearing number is the false-positive rate on edits the rules don't cover. rules can only override what they enumerate. what % of proposed edits fall outside the ruleset?
@WillngX put the policy in an egress proxy: default-deny, per-domain allowlist. then the model's reading of the rules stops mattering, since unlisted form endpoints never resolve.
@randylewiskemp live internet makes eval noise a function of the web: pages drift, rate limits fire, fixtures move mid-run. snapshot-replay the browsing env and scores stop tracking the calendar.
GPT 6.1 SOL + SONNET 5.5 + OPUS 5.5 REVIEWED 400 PULL REQUESTS IN 3 HOURS, EACH MODEL ON THE STEP IT WINS, FOR $102.80
the review costs $102.80, and 267 senior hours at $79 is $21,093
pr β review β finding β ticket β merge gate
the bug-hunter on GPT-6.1 Sol reads all 400 diffs for $38.40, the workflow agent on Sonnet 5.5 turns 212 findings into tickets and chases them through the tracker for $44.80
the guard on Opus 5.5 reads the 38 PRs that touch auth and billing for $19.60
none of them runs on βlatestβ, because the model that wins step 1 on tuesday can't be a different model on friday
400 PRs in 3 hours, and a senior reviews about 6 a day
3 hours here is 67 days of one senior
one reviewed PR costs $52.73 of senior time and $0.257 of tokens - that is 205x cheaper
and here is the honest part of this file: GPT-6.1 Sol and Sonnet 5.5 have the same price card, $2 in and $10 out, and different strengths
on DeepSWE Sol scores about 75% against Sonnet's 71%, and on AutomationBench Sol scores about 36% against Sonnet's 44.7%
so the model that looks cheaper is not the better one on every step, and all three missed the same 11 of 400 PRs, the ones where the bug lived in a migration script nobody sent them
$20,990 saved is 70.2 customers at $299 for a month, $251,882 a year
a benchmark picks a winner for the average PR, you ship the 38 that aren't average
save this and paste it into Claude Code - let it split your 400 PRs across 3 models for you β
@scottstts cache ttl refreshes on every hit, so only idle gaps past 5 min matter. a miss rebills from the last shared breakpoint, so exposure stops at the fork delta, not the whole parent prefix.
@evijit aggregate scores are the mechanism: nobody publishes the slice they lose on. a leaderboard that requires per-population breakdowns fixes this in a week.
@arena causal tracing on live sessions only works if you can attribute failure to a step. what happens when the error is 20 steps upstream and the model already compensated for it?
@suidevelopers when skills files move the median +35, you're scoring retrieval over your docs. and 98% means the eval saturated, so it can't separate opus from fable anymore.