found an interesting chatgpt voice mode usecase
i had it sit in the background during a mock interview with this prompt:
“stay in the background. listen to the conversation and only answer when i explicitly ask you something. keep responses under one sentence.”
i played both the interviewer and interviewee.
i never told it who was speaking or when to answer.
it still knew when i was asking it a question and jumped in with the answer.
the only thing slowing this down is voice latency.
@thsottiaux can we get an option to turn off agent voice in voice mode and stream responses as text instead?
It would unlock a whole new class of real-time ai assistants!!’
@leodev Really? I personally find it significantly better over the last few days and it feels like the June usage levels now. Which plan are you on again?
@jay_bizzz Really? What plan are you on?
I'm on the 20x and my usage has returned to normal and feels as good as it used to be, even ran some sol high fast agents and it felt great.
Tibo has responded to everyone's cry on usage limits and it's now time to address which concerns I outlined were dealt with indirectly, what's left to be dealt with and what is next that they should do.
Firstly what is now addressed and resolved or somewhat resolved.
I mentioned that
"I am aware that 5.6 loves subagents, and the parallelism of more agents at once could lead to significantly higher token utilization counts, but even then why are my limits still being blasted through so much? I really don't understand how this is happening."
And Tibo basically addressed that by saying that
" - GPT-5.6 Sol is much more willing to work for longer, make additional tool calls, and coordinate complex workflows across tools and subagents. That makes it better at solving hard problems, but some tasks were using far more than we intended.
- Sol also works harder at the same reasoning effort than previous models. High on Sol can use more tokens than High did on GPT-5.5.
- Programmatic tool calling, also referred to as code mode, gives Sol much more flexibility to run tool calls in parallel or continue working while waiting. But it also led to more responses per turn, more cached input tokens, and higher usage than expected.
- This was particularly noticeable when Sol was waiting for tool calls to finish or running many web searches. We’ve improved how we handle both cases and are continuing to make code mode more efficient."
These things that he mentioned were probably the underlying causes and reasons why on arguably lower setups like GPT-5.6-sol-high-normal > GPT-5.5-xhigh-fast the usage utilized was much higher even if it was doing the task inline (that agent by itself with no subagents). The programmatic tool calling is amazing for the model to do things that it needs, as a much more customizable and per task basis tool it is amazing. But the resulted cached input token from more responses is albeit strange that it even happened, as arguably a custom tool would be leaner and more specialized, thus smaller, but not hard to see how something like that would happen.
Also, "- The impact was also very uneven. The median user actually found Sol quite token efficient, while some power users working on harder tasks saw their usage drain much faster. We were very focused on average and median usage before launch and missed some cases where the long tail could use significantly more usage." which is honestly a huge mistake from their part, you cannot look and resolve the median user's concern without even addressing or looking into your outliers, be it on the low or on the high end. Huge mistake on their part, glad they mentioned it and glad they fixed it.
I also mentioned "4. Fix 5.6's token usage problem on a model and inference side." which he indirectly addressed by saying "We’ve been digging into what was happening and have landed several improvements." which I am assuming is just model and inference and architecture side improvements. Though I think it would be much better if he highlighted what they did, remembering the thinking juice change controversy a week+ ago...
Next, what is left to be dealt with and problems I think that they have left behind and not looked at it at all or at least not enough.
One of the main critiques is that usages were not displayed clearly and easily for the user to see.
People had to click on settings, then under a little tab called "Usage Remaining" to see a percentage and a little bar in a small pop out window in the middle of their screen. Codex used to tell you how much usage you had left next to the model selector itself within the chat box, then it moved to the bottom left next to settings, then into settings. This is definitely a mistake and a wrong move. By making it harder to users to see, and thus track their usage, people might end up just unknowingly burning and wasting their usages until a 10% left pop up pops out and then it is already too late. They should make it extremely easy to see at a glance how much usage a user has left, no matter if its at 100% or at 0%. I used to use Claude and my most visited site was https://t.co/2zhGOVFILF, and when I initially switched to Codex around April, I never had to go to any site to just see the usage because it was there in my chat box. But now that they moved it and hid it away for some reason, my most visited site is https://t.co/3cx2vlJ48Q just to see some green bars. Ridiculous, and I honestly have just been coping and not thinking about how dumb and stupid this is compared to what they where doing before.
More suggestions I have are points 1-2 of what solutions I thought of in my post linked below.
Lastly, what they should really do next. (Before the miscellaneous and little things I want to talk about.)
Tibo just said that today would be the last day of no 5h limits, and it would be back tomorrow. Many people are upset, saying things like "Removing the 5h limit was the best thing that Codex did" which is understandable but honestly a bit ridiculous. My thoughts on the 5h limit again in points 1-2 of my initial post, which again is what I think Tibo and team should really implement along side the more upfront and easier to see usage remaining bars.
And of course just talking with users and gathering their honest feedback, from people like @theo to random users on X (me included pls), or adding a little popup within Codex asking if people want to share their opinions for those without X. This would allow them to get the outlier's and less or more than median user's thoughts, so that they do not repeat the same mistake as above again.
Random or little things, I honestly love using Codex as a platform so much better than when I was using Claude, due to how much nicer their PR team and communications have been (though it's most likely just to spite Anthropic, that if Anthropic didn't exist, they probably wouldn't be doing) and the constant resets as well. Also I really hope that Tibo and team implement my 5h limit suggestions, PLEASE, no one would complain and everyone would literally be so much happier and their experience using Codex would be so much better, ALSO MAKE THE USAGE REMAINING VISIBLE FOR GODS SAKE.
Overall, huge kudos to Tibo and team for investigating and being so upfront and communicative with the people of Sol, massive win for us, besides the 5h limit coming back but nonetheless beautiful and another day of using Codex!
I think it is time to talk about the rate limits on Codex now.
Back on gpt-5.5, on the $200/month plan, I could freely run 2-3 gpt-5.5 agents on xhigh on fast mode nearly 24/7. I would never even have to personally think of the usage limits that I had and would even struggle to think of things to do to burn the limits when I was close to a 7 day reset. The 5h limit was not even a consideration at all, it is as if it never existed to be honest.
Now on gpt-5.6 Sol, on the $200/month plan, I run 1-2 High agents on normal speed maybe 12-16h a day. With the same credit count per million tokens as gpt-5.5, that should mean that my usage limits last more because
1. 1-2 agents instead of 2-3
2. high thinking instead of xhigh
3. normal speed instead of fast mode (2.5x more limit usage btw)
I also heavily optimize my local Codex settings and systems, thanks again to @theo for his article and YouTube video on those. So theoretically on paper my limits should last significantly longer, also because OpenAI claims the model is more token efficient than 5.5. However, I still absolutely burn through my limits insanely fast. Back when the 5h limit was still in place (which they said they temporarily removed but haven't added it back) I could burn it in 2-3h with just 2 Sol agents. I am aware that 5.6 loves subagents, and the parallelism of more agents at once could lead to significantly higher token utilization counts, but even then why are my limits still being blasted through so much? I really don't understand how this is happening.
All of the resets @thsottiaux has been giving us is just to tide us over. The removal of the 5h limit alongside the banked resets every week, the global resets every other day for users, is still barely enough for users, people are constantly begging and asking for resets. Users are also highlighting the problems once the 5h limit and mass resets are gone, people will no longer be able to sustain any usage on the plans. I am personally also very scared of what will happen when both of the above get reverted back to normal.
I think that there are a few solutions to this.
1. More detailed and customizable usage tracking. The concept of a 5h limit on paper is fine, a lower and smaller limit to make sure that the user does not burn their entire weekly in a day. But people found it too restrictive.
The solution: Allow users to simply choose the amount of limits they want, provide users with a number of the weekly limit. Users can add on another limit bar and then set how much percentage of their weekly it should be limited to. For example, I can set a 1 day limit to 14.2% of my weekly (around 1/7 of 100%), this would make it so that I manage and ration my limits out well.|
2. Model transparency and displaying of how much tokens and cost they actually do per thread and task in Codex. Claude has /usage which shows the user the API equivalent cost, Codex does not have this feature, I think that this is actually really useful and helpful for users to see costs. Each chat and thread in Codex should also display more detailed usage. I would love to see my cost per per message, the cached message and on my limits, 3 different rate limit estimates which are also customizable. A past day estimate of tokens in the last 24h vs the tokens a user has left to see at the past day's rate how fast the user would run out. A past hour estimate and a 5min estimate. This would allow users to easily just get an estimate of real time use for the past day, hour and 5mins to match their current and existing usage patterns. Some people use it very consistently, some people us it at peak times, some just use it when they use it. This would allow for users of all types to see their estimated time their limits will run out.
3. More transparency with limits. Codex currently operates on a system of "Credits" where 1 credit is $0.04 in API equivalent. I personally think that this is a layer built just for obfuscation and confusion, no one knows the Credits that they have, and let alone the amount of money that is worth. And if possible they should make the limits in terms of Credits, or even better just dollar amounts.
This would allow for users to truly understand how much actual money they are burning instead of just a 100% bar that they don't know how much it communicates and converts to.
4. Fix 5.6's token usage problem on a model and inference side.
This comes with some drawbacks, mainly to OpenAI as a company and not to users, people might end up being able to min-max their subscriptions even more and end up draining OpenAI's money. But us users have already been able to reverse engineer the usage limits. Where the $200 plans give like $14k in equivalent API credits so honestly not much effect on that front.
As much as OpenAI has been moving forward in a lot of fronts, it's time I address the Codex Voice feature, what it's doing good, and what it's terrible at (light crashout later).
First, some praise. This is genuinely a pretty cool feature. It's pretty amazing to be able to just talk to my computer, have my computer do stuff, and have it control my other computers. And other people, including people working at OpenAI have mentioned that explicitly as well. Also, apparently I can do this for my phone as well, but I don't use Mac, so I'm unable to actually get codex to directly control my other machine besides using SSH. I'm unable to use codex voice on my phone to control my codex on my computers. OpenAI, please add this to Android and Windows soon, I've been really wanting this!
Next, it's time we addressed a few issues. We'll start off with the fundamentals.
1. Why am I not able to set a hockey to mute my mic? I typically don't want Codex to be listening the entire time. I would love it to be able to configure hockey in my settings to just mute the codex microphone in the codex app. I checked in settings, and it does not exist. Why.
2. Why is there not an option to change the codex speaker volume? The only option I have is just mute or unmute this codex speaker. Why do I not have the option to change how loud it is? Make it softer, make it louder. Like all my other applications have these features that I would consider pretty basic.
3. Why does the codex's voice just disconnect itself with no warning? Two minutes ago, I was talking to codex. Codex was talking back to me, it was running some commands, whatever, and then I just heard the sound that shows that the codex disappeared. The voice just disconnected, but the background codex was still running its commands. We need an explanation of why the codex voice just disconnects. We can't just have it disconnect whenever it feels like it like that. It needs to be more reliable.
4. It's the light crash out time
The underlying model of the Codex voice is GPT Real-Time 2. Compared to the previous models, it's really, really good, especially at holding a conversation. The issue is that, in an environment like this within Codex, most people aren't really going to be using it to hold a conversation. People are going to be using it to do actual things, which it is kind of terrible at.
I was doing my math revision for tests, and I pulled up Codex voice. I sent it my paper that I was doing, and I was like, "Hey, help me out with this." I decided to get the Codex to actually open an Excalidraw whiteboard so that it can show me visually how it wanted me to do things so that on paper I had the same steps as what it was thinking on the whiteboard.
The problem was that the Codex GPT Real-Time model was just honestly kind of dumb. It was unable to format things properly or use the whiteboard properly. It might just be an issue with GPT models. Computer use is amazing, but still not perfect. Obviously, they still have work cut out for them.
When I was discussing what to do for a part of the question, GPT Realtime 2 messed up the formatting of Excalidraw, and it was not able to fix it by itself. I was mostly fine with it. I was just kind of annoyed. I told it to fix it, but then the problem started coming out: GPT Realtime 2, is kind of dumb. I asked it, "x is more than 3," and then "x square is more than 3. x square is positive," which obviously it is. It didn't just say that; it just went ahead and said, "Let me check that." I've done similar things with Gemini 2.5 Flash-Lite in the Google AI Studio, and honestly, I've had a better experience communicating and doing these kinds of things with the model with Gemini 2.5 Flash-Lite. It is insane for me to say that Gemini 2.5 Flash-Lite is better than GPT Realtime 2 but that that task, it just fucking is better. I'm sorry. Like, when I said, "x is more than 3, and so that means x square is positive," which is VERY obvious, Gemini would just say, "Yes, it is." It doesn't need to say, "Let me check that. Let me think." It doesn't need to think because it knows that that is correct, but GPT decided to go ahead and start thinking.
And then it messed up the formatting of the whiteboard. It wasn't explaining things to me clearly because it just decided, "I'm not going to try to explain how to do the question to me. It was just going to go ahead and do the entire problem by itself on the whiteboard." I was really pissed because I explicitly told it before not to do that, and I said, "Explain to me step by step. Don't jump ahead." It did that. I told it, "What are you doing? Undo what you did." Instead of undoing it, it just closed the tab and said, "Oh, I can't do anything. I don't know what to do now." I was really pissed at this point, and I just told the GPT Realtime model to just spin up a separate Sol-light tread to go ahead and do it instead of it doing it for me itself. Sol-light is honestly so amazing and fast and reliable at computer use tasks like this, but then the GPT Realtime 2 just fucked up the instructions to Sol, saying wrong formatting things and just made Sol mess up the formatting as well.
In short, GPT Realtime 2 is amazing at holding a conversation, just don't ask it to actually do things for you, but it itself doing the task or it delegating the task to someone else, because if it isn't something that is concrete and relies a lot on perception and understanding and formatting, then you might have a better time with a Google model, that's a first for sure.
Most people just say all of the good things about the features they are using, but some need to point out the bad things, to make the feature get better for them and for everyone.
Anyways, I need to get back to studying for my math test, thanks GPT Realtime 2 for wasting an hour of my time and making me crashout on X about this.
Tibo has responded to everyone's cry on usage limits and it's now time to address which concerns I outlined were dealt with indirectly, what's left to be dealt with and what is next that they should do.
Firstly what is now addressed and resolved or somewhat resolved.
I mentioned that
"I am aware that 5.6 loves subagents, and the parallelism of more agents at once could lead to significantly higher token utilization counts, but even then why are my limits still being blasted through so much? I really don't understand how this is happening."
And Tibo basically addressed that by saying that
" - GPT-5.6 Sol is much more willing to work for longer, make additional tool calls, and coordinate complex workflows across tools and subagents. That makes it better at solving hard problems, but some tasks were using far more than we intended.
- Sol also works harder at the same reasoning effort than previous models. High on Sol can use more tokens than High did on GPT-5.5.
- Programmatic tool calling, also referred to as code mode, gives Sol much more flexibility to run tool calls in parallel or continue working while waiting. But it also led to more responses per turn, more cached input tokens, and higher usage than expected.
- This was particularly noticeable when Sol was waiting for tool calls to finish or running many web searches. We’ve improved how we handle both cases and are continuing to make code mode more efficient."
These things that he mentioned were probably the underlying causes and reasons why on arguably lower setups like GPT-5.6-sol-high-normal > GPT-5.5-xhigh-fast the usage utilized was much higher even if it was doing the task inline (that agent by itself with no subagents). The programmatic tool calling is amazing for the model to do things that it needs, as a much more customizable and per task basis tool it is amazing. But the resulted cached input token from more responses is albeit strange that it even happened, as arguably a custom tool would be leaner and more specialized, thus smaller, but not hard to see how something like that would happen.
Also, "- The impact was also very uneven. The median user actually found Sol quite token efficient, while some power users working on harder tasks saw their usage drain much faster. We were very focused on average and median usage before launch and missed some cases where the long tail could use significantly more usage." which is honestly a huge mistake from their part, you cannot look and resolve the median user's concern without even addressing or looking into your outliers, be it on the low or on the high end. Huge mistake on their part, glad they mentioned it and glad they fixed it.
I also mentioned "4. Fix 5.6's token usage problem on a model and inference side." which he indirectly addressed by saying "We’ve been digging into what was happening and have landed several improvements." which I am assuming is just model and inference and architecture side improvements. Though I think it would be much better if he highlighted what they did, remembering the thinking juice change controversy a week+ ago...
Next, what is left to be dealt with and problems I think that they have left behind and not looked at it at all or at least not enough.
One of the main critiques is that usages were not displayed clearly and easily for the user to see.
People had to click on settings, then under a little tab called "Usage Remaining" to see a percentage and a little bar in a small pop out window in the middle of their screen. Codex used to tell you how much usage you had left next to the model selector itself within the chat box, then it moved to the bottom left next to settings, then into settings. This is definitely a mistake and a wrong move. By making it harder to users to see, and thus track their usage, people might end up just unknowingly burning and wasting their usages until a 10% left pop up pops out and then it is already too late. They should make it extremely easy to see at a glance how much usage a user has left, no matter if its at 100% or at 0%. I used to use Claude and my most visited site was https://t.co/2zhGOVFILF, and when I initially switched to Codex around April, I never had to go to any site to just see the usage because it was there in my chat box. But now that they moved it and hid it away for some reason, my most visited site is https://t.co/3cx2vlJ48Q just to see some green bars. Ridiculous, and I honestly have just been coping and not thinking about how dumb and stupid this is compared to what they where doing before.
More suggestions I have are points 1-2 of what solutions I thought of in my post linked below.
Lastly, what they should really do next. (Before the miscellaneous and little things I want to talk about.)
Tibo just said that today would be the last day of no 5h limits, and it would be back tomorrow. Many people are upset, saying things like "Removing the 5h limit was the best thing that Codex did" which is understandable but honestly a bit ridiculous. My thoughts on the 5h limit again in points 1-2 of my initial post, which again is what I think Tibo and team should really implement along side the more upfront and easier to see usage remaining bars.
And of course just talking with users and gathering their honest feedback, from people like @theo to random users on X (me included pls), or adding a little popup within Codex asking if people want to share their opinions for those without X. This would allow them to get the outlier's and less or more than median user's thoughts, so that they do not repeat the same mistake as above again.
Random or little things, I honestly love using Codex as a platform so much better than when I was using Claude, due to how much nicer their PR team and communications have been (though it's most likely just to spite Anthropic, that if Anthropic didn't exist, they probably wouldn't be doing) and the constant resets as well. Also I really hope that Tibo and team implement my 5h limit suggestions, PLEASE, no one would complain and everyone would literally be so much happier and their experience using Codex would be so much better, ALSO MAKE THE USAGE REMAINING VISIBLE FOR GODS SAKE.
Overall, huge kudos to Tibo and team for investigating and being so upfront and communicative with the people of Sol, massive win for us, besides the 5h limit coming back but nonetheless beautiful and another day of using Codex!
I think it is time to talk about the rate limits on Codex now.
Back on gpt-5.5, on the $200/month plan, I could freely run 2-3 gpt-5.5 agents on xhigh on fast mode nearly 24/7. I would never even have to personally think of the usage limits that I had and would even struggle to think of things to do to burn the limits when I was close to a 7 day reset. The 5h limit was not even a consideration at all, it is as if it never existed to be honest.
Now on gpt-5.6 Sol, on the $200/month plan, I run 1-2 High agents on normal speed maybe 12-16h a day. With the same credit count per million tokens as gpt-5.5, that should mean that my usage limits last more because
1. 1-2 agents instead of 2-3
2. high thinking instead of xhigh
3. normal speed instead of fast mode (2.5x more limit usage btw)
I also heavily optimize my local Codex settings and systems, thanks again to @theo for his article and YouTube video on those. So theoretically on paper my limits should last significantly longer, also because OpenAI claims the model is more token efficient than 5.5. However, I still absolutely burn through my limits insanely fast. Back when the 5h limit was still in place (which they said they temporarily removed but haven't added it back) I could burn it in 2-3h with just 2 Sol agents. I am aware that 5.6 loves subagents, and the parallelism of more agents at once could lead to significantly higher token utilization counts, but even then why are my limits still being blasted through so much? I really don't understand how this is happening.
All of the resets @thsottiaux has been giving us is just to tide us over. The removal of the 5h limit alongside the banked resets every week, the global resets every other day for users, is still barely enough for users, people are constantly begging and asking for resets. Users are also highlighting the problems once the 5h limit and mass resets are gone, people will no longer be able to sustain any usage on the plans. I am personally also very scared of what will happen when both of the above get reverted back to normal.
I think that there are a few solutions to this.
1. More detailed and customizable usage tracking. The concept of a 5h limit on paper is fine, a lower and smaller limit to make sure that the user does not burn their entire weekly in a day. But people found it too restrictive.
The solution: Allow users to simply choose the amount of limits they want, provide users with a number of the weekly limit. Users can add on another limit bar and then set how much percentage of their weekly it should be limited to. For example, I can set a 1 day limit to 14.2% of my weekly (around 1/7 of 100%), this would make it so that I manage and ration my limits out well.|
2. Model transparency and displaying of how much tokens and cost they actually do per thread and task in Codex. Claude has /usage which shows the user the API equivalent cost, Codex does not have this feature, I think that this is actually really useful and helpful for users to see costs. Each chat and thread in Codex should also display more detailed usage. I would love to see my cost per per message, the cached message and on my limits, 3 different rate limit estimates which are also customizable. A past day estimate of tokens in the last 24h vs the tokens a user has left to see at the past day's rate how fast the user would run out. A past hour estimate and a 5min estimate. This would allow users to easily just get an estimate of real time use for the past day, hour and 5mins to match their current and existing usage patterns. Some people use it very consistently, some people us it at peak times, some just use it when they use it. This would allow for users of all types to see their estimated time their limits will run out.
3. More transparency with limits. Codex currently operates on a system of "Credits" where 1 credit is $0.04 in API equivalent. I personally think that this is a layer built just for obfuscation and confusion, no one knows the Credits that they have, and let alone the amount of money that is worth. And if possible they should make the limits in terms of Credits, or even better just dollar amounts.
This would allow for users to truly understand how much actual money they are burning instead of just a 100% bar that they don't know how much it communicates and converts to.
4. Fix 5.6's token usage problem on a model and inference side.
This comes with some drawbacks, mainly to OpenAI as a company and not to users, people might end up being able to min-max their subscriptions even more and end up draining OpenAI's money. But us users have already been able to reverse engineer the usage limits. Where the $200 plans give like $14k in equivalent API credits so honestly not much effect on that front.
Is Sol getting smatter? I have used that feature one other time when 5.5 brought it up mid-June, amazing to see it suggesting and being so upfront here, would live for OpenAI to push this more.
Is Sol getting smatter? I have used that feature one other time when 5.5 brought it up mid-June, amazing to see it suggesting and being so upfront here, would live for OpenAI to push this more.
I think it is time to talk about the rate limits on Codex now.
Back on gpt-5.5, on the $200/month plan, I could freely run 2-3 gpt-5.5 agents on xhigh on fast mode nearly 24/7. I would never even have to personally think of the usage limits that I had and would even struggle to think of things to do to burn the limits when I was close to a 7 day reset. The 5h limit was not even a consideration at all, it is as if it never existed to be honest.
Now on gpt-5.6 Sol, on the $200/month plan, I run 1-2 High agents on normal speed maybe 12-16h a day. With the same credit count per million tokens as gpt-5.5, that should mean that my usage limits last more because
1. 1-2 agents instead of 2-3
2. high thinking instead of xhigh
3. normal speed instead of fast mode (2.5x more limit usage btw)
I also heavily optimize my local Codex settings and systems, thanks again to @theo for his article and YouTube video on those. So theoretically on paper my limits should last significantly longer, also because OpenAI claims the model is more token efficient than 5.5. However, I still absolutely burn through my limits insanely fast. Back when the 5h limit was still in place (which they said they temporarily removed but haven't added it back) I could burn it in 2-3h with just 2 Sol agents. I am aware that 5.6 loves subagents, and the parallelism of more agents at once could lead to significantly higher token utilization counts, but even then why are my limits still being blasted through so much? I really don't understand how this is happening.
All of the resets @thsottiaux has been giving us is just to tide us over. The removal of the 5h limit alongside the banked resets every week, the global resets every other day for users, is still barely enough for users, people are constantly begging and asking for resets. Users are also highlighting the problems once the 5h limit and mass resets are gone, people will no longer be able to sustain any usage on the plans. I am personally also very scared of what will happen when both of the above get reverted back to normal.
I think that there are a few solutions to this.
1. More detailed and customizable usage tracking. The concept of a 5h limit on paper is fine, a lower and smaller limit to make sure that the user does not burn their entire weekly in a day. But people found it too restrictive.
The solution: Allow users to simply choose the amount of limits they want, provide users with a number of the weekly limit. Users can add on another limit bar and then set how much percentage of their weekly it should be limited to. For example, I can set a 1 day limit to 14.2% of my weekly (around 1/7 of 100%), this would make it so that I manage and ration my limits out well.|
2. Model transparency and displaying of how much tokens and cost they actually do per thread and task in Codex. Claude has /usage which shows the user the API equivalent cost, Codex does not have this feature, I think that this is actually really useful and helpful for users to see costs. Each chat and thread in Codex should also display more detailed usage. I would love to see my cost per per message, the cached message and on my limits, 3 different rate limit estimates which are also customizable. A past day estimate of tokens in the last 24h vs the tokens a user has left to see at the past day's rate how fast the user would run out. A past hour estimate and a 5min estimate. This would allow users to easily just get an estimate of real time use for the past day, hour and 5mins to match their current and existing usage patterns. Some people use it very consistently, some people us it at peak times, some just use it when they use it. This would allow for users of all types to see their estimated time their limits will run out.
3. More transparency with limits. Codex currently operates on a system of "Credits" where 1 credit is $0.04 in API equivalent. I personally think that this is a layer built just for obfuscation and confusion, no one knows the Credits that they have, and let alone the amount of money that is worth. And if possible they should make the limits in terms of Credits, or even better just dollar amounts.
This would allow for users to truly understand how much actual money they are burning instead of just a 100% bar that they don't know how much it communicates and converts to.
4. Fix 5.6's token usage problem on a model and inference side.
This comes with some drawbacks, mainly to OpenAI as a company and not to users, people might end up being able to min-max their subscriptions even more and end up draining OpenAI's money. But us users have already been able to reverse engineer the usage limits. Where the $200 plans give like $14k in equivalent API credits so honestly not much effect on that front.