@AlexanderKalian@p_droaraujo Do give that estimate. Do not forget to mention that we are heading into the singularity.
Ultimately that’s what it will come down to. If ya believe that’s where we are heading. Shit will be fast.
If you do not believe we’re heading there. You believe people are insane.
@gro_tsen But few actually knew you could get these sort of results a couple of days ago. So there has literally not been enough time for all those questions to be asked.
Claude is expensive. But Codex is so inefficient. Extreme difference in the harnesses.
Running tests, waiting and whatnot? Codex defaults to checking in every damn minute. Spending hundreds of thousands of tokens.
Claude? Sets a notification and sleeps until it needs to act.
@thsottiaux I'm sure you'll get it right eventually. Do not envy you in trying to sort that bit out. But it really has to be crystal clear for non-dev to get it to work.
@thsottiaux I'm trying. Capabilities are a bit unclear though. They've seen me steer my desktop from the iOS app, and try similar things. They end up on cloud agents rather than remote ones.
All of the frustration basically boils down to "it can't see my mail/calendar/browser"
GPT-5.6 Sol and Luna are ahead of Terra at every point on the Intelligence vs Cost per Task chart. GPT-5.6 Luna stands out as a particularly cost efficient model
Charting the Artificial Analysis Intelligence Index shows the trade-off between intelligence and Cost per Intelligence Index Task. Across reasoning efforts, each GPT-5.6 model pushes past GPT-5.5 on the Pareto frontier (excluding non-reasoning).
However, Luna and Sol are always ahead of Terra. This means for any Terra effort level, there is a Luna or Sol effort level that is more intelligent at no extra cost, or as intelligent at lower cost.
@lyc_aon You do get the point, don’t you? I used to run 3-10 agents for a good amount of time during 5.2 and earlier.
Now I’m sitting here with 10% of my 5h usage left. I’ve used 1 or 2 agents. Luna xHigh is 80% of that usage (yes I checked) Sol high the rest.
@boon_jin@simonw Never use 1m context. Use 200k. Tell it to spawn haiku sub agents when applicable.
Works decent for me.
With 1m context it wasn’t useable.
@normaltyp Hahah. Ska verkligen inte gnälla, imponerande jobb att kasta ihop så snabbt! Har du tur finns det DB dumpar från Football Manager, har du ha jobbet gratis!
@daniel_mac8@extliqprovider@GaelBreton But there's zero reason to believe the ROI will be there for every dev use. It'll be for certain work. The other percentage? Devs will still use AI for that, just not Mythos.
If a cheaper 5.6 beats whatever Anthropic provides at the same price? It'll be used.
Not really. The iterative process where you can get and receive feedback from business, design and QA is vital as a developer.
Even more so with AI. Especially if you're trying some sort of "spec driven development". There is no initial "complete" spec. It needs revision. Sprints give time for that.
@ryanprasad_ai@CWood_sdf Hmm, this sounds interesting. Sounds like we need a term for this. It sounds agile. Maybe we should call it that?
Agile software development. I like the sound of it.
Brilliant!
@zuess05@EscoTrajan Seniors are reviewing code on their level. Juniors are reviewing code they do not understand.
It’s like having a teacher grading stuff they do not understand. Versus one who does.