@samuelcolvin It’s also more auditable in addition to the papers shown by others and easier for reviewing agents to understand exactly what and how it was done more than just a diff. Fewer typos / hallucinations in syntax …. (Fewer, not none.)
Opus is not an orchestrator. It is a fantastic builder / sub agent. It requires a robust harness and orchestrator, but if those are in place it sings. Without it, it's like trying to hammer with a screwdriver.
and Opus 5 is up. Dark factory operating. 60% reduction in cost achieved on initial benches, but I need to dig in and re-bench 4.8 to make sure that's not just because I basically rebuilt the factory from ground up.
4.7 is where it hit for me as well. I thought it was just harness issues and eventually got it to work, but every model since has been a massive pain to evaluate because I feel like they think they "know better" and the alignment breaks. I wouldn't put Opus in production anywhere at this point... or maybe ever.
@0xSero that's sick. Wonder if you could even constrain the context and do some cache fan outs for concurrency gains. TTFT is rough if that's warm, but I assume this is cold yeah?
@MiaAI_lab@Kimi_Moonshot Check your cache especially if you are using another harness. There are specific cache requirements for Kimi and if you aren't careful you'll blow through your quota in minutes. Check cache invalidation as well. Made a huge difference (90%)
@tenobrus 100% - my dark factory had no problem using I as a skeptical builder meant to catch itself and it's orchestrator. As soon as it got put in the orchestrator seat, it went in circles and was not able to operate hardly at all. (sol, grok and Kimi could)
I think I figured out Opus 5. Had to retrofit my entire dark factory to a red light green light system with squid game level mechanics... but it's working now. Will drop a lessons learned once I get this smooth.
@trevin Ironically, when plugging my harness into Claude Code and struggling to align... I didn't try it without the Claude Code piece. Maybe they didn't get it right with their harness and I'm fighting 2 layers.
@adamlyttleapps about 25% as far as I can tell so far, but sessions are longer meaning more turns. not enough data to quantify, but yeah, seems more tokens, more chatty than fable.
The Opus 5 experience: spend dozens of turns exhausting every excuse and counterargument until it finally says it completely understands... then watch it immediately do the opposite.
I have had to switch to fable to get back on track too many times now.
@dani_avila7 Careful, I've been having to correct and argue with Opus like crazy... just after you call it out. Put Sol as an adversarial review and on plan. It even argues with my harnesses. Seems like it's saying ok ok ok human whatever. And then just does what it wants.
If you have a tuned harness that you refined through countless evals... Opus 5 sucks. It literally ignores my harness if not forced.
Complete harness rewrite required...