@hegargarcia Haha, exactly. Sol is similar in this regard. They just need a strict harness with restrictions, which on the other side is a good thing in its own. You have to saddle the horse well if you want it to reach the finish line.
@ChiragAsarpota I saw more people having this opinion. But please take into consideration that openai gets $200 month after month and many people will not spend every dollar. So that $17k is not guaranteed and the actual difference isnt that big. Looks to me like a pricing strategy.
When you build your harness, what is the practical difference between sometimes needs checking vs always needs checking? Just picking one example, Sol sometimes do ABC when asked to do A, so it always needs checking - unless your harness is so strict it never allows him to do anything outside the scope of what was asked.
@Kappaemme1926@jonnno_ Install Stats and look at Ram usage, that can be a culprit. I even replaced VS Code with Zed and iTerm with Ghostty and that helped a bit.
But if you want to scale the harness, the agents have to run elsewhere, on Mac mini or VM or some cloud.
@GohilHardy One of the main reasons for me is that there are more CLIs at the same time. :-) Usually 3 terminals & 3 agents, and I do have diffs and files; when I need to look at some file, I just ask agent to open it in Zed.
I experienced majority of those issues but some might be mitigated and are prominent also with other models. I found the harness should be more strict, for example its quite common that agent (not only Opus) declares unfinished task as finished. Thats because agent should only produce evidence of work being done and other agent checks if work is done, not him.
Not following claude.md explicit instruction - first thing is to look at the file size and also memory.md size; it tends to grow and after certain size agent is not capable to follow everything. I suggest to ask agent to list all autoloaded files and inspect them. When I looked at memory.md, there was a lot of garbage, some of it not even true. That can also contribute to problems described.
@LyraInTheFlesh We would need to see the whole session log to make judgment why the agent drifted like this, but I'd say your harness isn't strict enough. During work, agent shouldn't be even doing such elaborations.
Yes, it is perfectly fine to migrate, I for one use both 50:50 in CLI. The permissions settings are a bit different, you might spend some time there. What I don't like about GPT models is that they very often do textwalls, like they often don't know how to dialogue, it feels like they just throw unedited text at you. Fable is subjectively better in this regard.
But if you only run a strict harness, software dev etc., it is more or less the same. I often let Claude do the coding and codex to review, or vice versa. I'd say I am satisfied with using both.
I was building something similar, with agents communicating via tmux. But when agents just discussed between themselves and then one did it, I noticed it drifted towards my experiencing cognitive surrender, where I didnt already understand fully what is going on under the hood. Then I concluded this isnt scalable or can be only used for simpler things.
It now feels like an anti-pattern, but of course I might just miss something.
Minority opinion: I actually like Opus 5. A workhorse that's also a bit wild. Needs a heavy harness, likes strict rules and clear direction. Not at his best when I let him run around aimlessly or just explore stuff. Under a firm hand, he finishes his homework.
@AstraiaAI Yes because Fable is the only model who at least tries to speak like a human, unlike all other models who just throw textwalls at me. I sometimes chat with Grok too, since it knows what happens now. Other models are good for work, but I try to not read what they say.
@arturovilla Good for you, but from my experience, I cannot agree with point 4 - I regularly use Fable, Sol, Grok, Opus and Terra and this is how I would order them by overall capability. Grok is quickly catching up, which is welcomed, but Claude isn't falling anywhere.