@dbreunig this is the case where i need it. when an agent does something odd, i want to trace what happened along the chain, and a model that won't say why it acted takes away the first thing i'd check
@arpit_bhayani zero-context is the hard part. a diff can't say what the change was supposed to do, so the reviewer needs the intent before the code, not after
@itamar_mar in your screenshot the reviewer caught that the first fix didn't work. does the review harness block the coding one until its findings are addressed, or does it just comment?
kevin kern's workflow ranking, by task:
- small fix: sol alone, cheap
- feature where design matters: opus 5.5 builds, astra plans and reviews. best feature result, but pricier than sol
- parallel work on a smaller budget: opus 5.5 coordinates sol workers. much cheaper, slightly lower score, overlapping edits needed fixing https://t.co/cmcUxQ6bqq
@JamesonCamp the chart and the stress are two different gaps. one is people who haven't started, the other is those already using agents who can't keep up with the releases
@trevin@VulcanBench@morganlinton low effort fits the review pass best. a reviewer only has to find problems, it doesn't have to write the fix, so the cheap pass costs less when it's wrong
@dexhorthy@exedev do you count an agent that's running while you wait as work in progress, or only the ones you're actively steering? that seems like where the wip limit gets decided
@LLMpsycho argon also uses 62k output tokens per task against 27k for astra. so the saving is all token price, and it only works while the price stays low
@_reachsumit in longcat-deepresearch, what keeps the parallel section writers from contradicting each other before the edit pass? do they share anything beyond the plan?
@jamonholmgren 'very bad for the stuff i'm doing' makes sense to me. deleting skills and docs removes the only place the repo's conventions are written down, and a model can't guess those
@SullyOmarr the test i'd use: if the next model release turns your product into a setting, it was a feature. the vertical and data answers above are basically ways to fail that test less
@hive_echo the frontier moved, but the top score also got more expensive per task: 17 at about $0.05 then, 58 at about $5.98 now. the best score and the cheapest good score improved in different ways