“GPT-6 Sol is just a cheaper GPT-5.6 Sol”? Two requirements in my test say otherwise.
Same starting point. GPT-6 Sol vs GPT-5.6 Sol: 22 vs 81 min, 12M vs 50M reported tokens, and $3 vs $23.50 API-equivalent.
GPT-6 Sol found conflicts in two requirements and held off. GPT-5.6 Sol implemented both incorrectly without warning.
@HarryKarl_@ForwardEditor@OpenAI Skill issue! Same $200 Pro plan. I’m using Sol xhigh to implement a complex engineering standard: 6,000+ requirements and 18,000+ tests that must pass. I’m at ~0.6% of my weekly limit/hour. Stop dismissing others’ work and audit your setup before blaming the model.
Early GPT-6 Sol (xhigh) impression: ~0.75% of my weekly limit per hour, versus ~2.3% with Astra (xhigh). It’s only been 2h40m, but the results look just as good to me. 6 Sol also feels sharper than 5.6 Sol. If this rate holds, I might actually make it through the week.
@pvncher Have them implement a complex new engineering standard from scratch: using a reference implementation that isn’t fully compliant and a specification with contradictions.
That’s pretty much what I’m doing right now with SysML v2.
Of course. 😂
I spend 7+ hours testing the experimental context management, conclude that it works surprisingly well for my long-running workflow…
…and OpenAI pulls the experiment.
Bummer. Bad timing on my part.
GPT-6 Astra’s experimental context management is changing how I use the model, and my first quota numbers are interesting.
I’m on the $200 Pro plan, using Astra to turn a complex technical standard into working code. Lots of reasoning, implementation, testing, validation, and correcting code produced by older, less capable models.
My previous setup started a fresh automation task every ~30 minutes at High effort. That was intentional. With traditional compaction, very long tasks meant compaction after compaction, with increasing risk of losing important context.
That setup used around 2.5 to 3% of my weekly quota per hour.
Now I’ve enabled Astra’s experimental context management. This makes long-running tasks practical for my workload, so instead of constantly starting fresh, I gave one Goal task the job and let it keep working with its context.
And I raised reasoning from High to XHigh.
It ran for 7+ hours.
Usage: roughly 2% of my weekly quota per hour.
So I increased reasoning effort while observed hourly quota usage actually went down.
This isn’t a controlled benchmark. The work packages differ, and Astra spent some of that time finding and fixing code produced by older models that had previously been considered done. But the code is demonstrably more correct and the pace of meaningful progress looks comparable.
I’m not claiming XHigh is cheaper than High. But it’s another data point against the idea that simply lowering reasoning effort is the obvious way to save quota.
For my workload, task lifetime and context continuity seem to matter too.
GPT-6 Astra’s experimental context management is changing how I use the model, and my first quota numbers are interesting.
I’m on the $200 Pro plan, using Astra to turn a complex technical standard into working code. Lots of reasoning, implementation, testing, validation, and correcting code produced by older, less capable models.
My previous setup started a fresh automation task every ~30 minutes at High effort. That was intentional. With traditional compaction, very long tasks meant compaction after compaction, with increasing risk of losing important context.
That setup used around 2.5 to 3% of my weekly quota per hour.
Now I’ve enabled Astra’s experimental context management. This makes long-running tasks practical for my workload, so instead of constantly starting fresh, I gave one Goal task the job and let it keep working with its context.
And I raised reasoning from High to XHigh.
It ran for 7+ hours.
Usage: roughly 2% of my weekly quota per hour.
So I increased reasoning effort while observed hourly quota usage actually went down.
This isn’t a controlled benchmark. The work packages differ, and Astra spent some of that time finding and fixing code produced by older models that had previously been considered done. But the code is demonstrably more correct and the pace of meaningful progress looks comparable.
I’m not claiming XHigh is cheaper than High. But it’s another data point against the idea that simply lowering reasoning effort is the obvious way to save quota.
For my workload, task lifetime and context continuity seem to matter too.
As AI models get better, our role needs to evolve too. That’s human–AI co-evolution: stop optimizing how you work around yesterday’s AI limitations. Some of your “best practices” are just scars from weaker models.
Did OpenAI change something? My weekly quota suddenly jumped from around 30% to 70%. I thought it might be an early reset, but my next reset is still September 14, five days away. Not September 16, as I’d expect if it had reset today.
I am getting more done in less time with Astra, and spending much less time steering it. I get the complaints about budget burn, but I think the work you get out of it matters too, along with how much overhead you have carried over from older models.
After about four hours using Astra with reasoning effort set to High, I have used 7% of my weekly limit implementing a complex systems engineering standard. Compared with Sol, I am getting through more problems faster, with a lot less back-and-forth.
Before switching, I removed Superpowers and a bunch of process rules I had added to work around the limitations of Sol. That already helped. Astra now has more room to figure things out, and so far, that trust seems well placed.
Running out of x20 Pro limits after two days still sucks. But I would also take a look at whether all those rules and workflows are still helping. Some of mine were just eating budget.
I’m currently stress-testing Codex: my MacBook is 3,000 km away and I only have the ChatGPT iOS app.
The machine stays online and existing threads remain accessible. But the problems already started on day two. One long-running Codex task could no longer create a new task through the app UI, so Codex itself decided to launch it via the Codex CLI instead. That worked, but the new task never appeared in the ChatGPT/iOS UI.
Since then, existing turns have degraded too. I can send a message, but often only get the first paragraph back: no tool activity, no progress updates, no visible completion and none of the normal end-of-turn controls. Automations still run, but show the same behavior.
So the remote connection is fine and work may even still be happening. But without a reliable view of task state or a way to recover the UI remotely, I’m effectively locked out of the system.
Apparently Codex means I can leave my laptop 3,000 km away and keep shipping from my phone.
Cute story.
After a few days of real remote use: tasks won’t spawn, answers stop after paragraph one, restarts do nothing, and I’m effectively locked out of my work.
Great demo material. Much harder to call it a serious productivity tool when a few days of actual remote use leave you completely stranded.
No tipping for FAFO tokens.
Don’t upsell me on cleanup.
The $80 Codex reset experiment hits a nerve because a lot of my quota is already spent keeping the model useful. Let it run autonomously for a day or two and there’s a good chance it starts drifting, overengineering, adding process around the process, or solving a different problem than the one I actually gave it.
So I keep tuning the setup, changing agent patterns, adjusting reasoning effort, simplifying the harness, reviewing the reviews, and cleaning things up just to get back to productive work. That’s the FAFO tax: not failed experiments, but the constant cost of keeping an experimental system pointed in the right direction.
I’m happy to pay more when more capacity gives me more useful output. But charging another $80 after a good chunk of the quota went into keeping the product from wandering off is a pretty spicy upsell.
Hey @thsottiaux,
Feature request for long-running Codex tasks:
I’d love a conversation-first view where my interventions and Codex’s direct replies stay front and center, instead of getting buried in progress updates.
Keep execution details available, but treat them more like a collapsible build log.
The main UI should be the actual human-agent conversation and the final outcomes.
This would make long-running, steerable tasks feel so much better.
Same here. Burned ~80% of my quota in 24 hours, Ultra + Fast mode. 😅
One thread. One project. One specific problem.
And I’d already optimized my setup before the resets to avoid Auto Review as much as possible.
Still had to hit the brakes to save enough quota for the work I actually need to get done today.
Tibo announcing the next reset before people have even burned through the current one is pretty clever.
It tells every Codex power user: stop budgeting. Go brrrrrr.
If OpenAI is validating recent fixes, auto-review included, that’s one hell of a distributed production test.
The reset playbook:
🚀 launch / celebrate
📈 boost usage / adoption
🔥 stress-test fixes / gain real-world insights at scale
Monday feels a lot like #3.
If Codex Auto Review is consuming a surprisingly large share of your usage, don’t disable it blindly. Treat it as an optimization problem:
- Analyze which recurring actions trigger the most reviews.
- Group them by command and underlying workflow.
- Pre-approve only narrow, repetitive, low-risk operations.
- Fix workflows that generate unnecessary retries or polling.
- Keep Auto Review for genuinely sensitive or unusual actions.
- Re-run the analysis regularly via automations and refine the rules based on actual usage.
The goal is to make sure Auto Review is invoked only when its judgment adds real safety value.