Hello people of Sol! I've reset usage limits for all ChatGPT Work and Codex users. Together with that, a quick update on GPT-5.6 Sol usage limits.
Over the past few weeks, many of you have told us that Sol was using your Codex limits faster than expected. To be clear, we have not reduced usage on any subscription plans.
We’ve been digging into what was happening and have landed several improvements. As a result, we expect your usage to last around 18% longer during typical use of Sol. Some of you should already see significantly larger improvements from today. Tomorrow, we’ll also restore the five-hour limit that we temporarily paused while investigating.
Here’s what we found:
- GPT-5.6 Sol is much more willing to work for longer, make additional tool calls, and coordinate complex workflows across tools and subagents. That makes it better at solving hard problems, but some tasks were using far more than we intended.
- Sol also works harder at the same reasoning effort than previous models. High on Sol can use more tokens than High did on GPT-5.5.
- Programmatic tool calling, also referred to as code mode, gives Sol much more flexibility to run tool calls in parallel or continue working while waiting. But it also led to more responses per turn, more cached input tokens, and higher usage than expected.
- This was particularly noticeable when Sol was waiting for tool calls to finish or running many web searches. We’ve improved how we handle both cases and are continuing to make code mode more efficient.
- The impact was also very uneven. The median user actually found Sol quite token efficient, while some power users working on harder tasks saw their usage drain much faster. We were very focused on average and median usage before launch and missed some cases where the long tail could use significantly more usage.
Sol is a significant step forward in what Codex can do, but capability and efficiency do not always improve at the same pace, and some issues only become clear once people are using the model at real-world scale. We should have recognized this sooner and been more upfront about it.
You keep pushing the frontier and we’ll keep improving efficiency and sharing updates as we go.
Should the Llama 3.1 70B upgrade be the story of the week? 🤔 The 3.1 70B model arguably has the best quality vs. price (& likely size) position of all models.
405B is exciting for bringing open-source up-to-par with leading proprietary models but it has not established a new position for LLMs across quality/price & likely model size.
Llama 3.1 70B can be accessed ~8X cheaper than the leading models including GPT-4o and Llama 3.1 405B (median of provider prices) and there is only a 3% difference in MMLU. The upgrade from 3 to 3.1 decreased the gap between Llama 3 70B and GPT-4o more than the current difference between the models.
Furthermore, as speed and price become more important as use-cases expand and techniques for using models involve more model ‘requests’, particularly for AI agents, the relative position on 3.1 70B becomes stronger.
It should be noted that a 3% difference in MMLU is considerable and new, harder benchmarks are needed to comprehensively measure model capabilities but MMLU has still held broadly reflective of general intelligence capabilities.
Let us know if you agree and your thoughts below 💬
We now finally have the thing I've always wanted: an analysis of the best vision models for finetuning.
In joint work with @capetorch we tried nearly 100 models on two very different datasets across various hyperparams.
Here's the top 15 on the IIT-Pet dataset: 1/🧵
Who's gonna pay for the 💰$2 trillion 💰US stimulus bill --> No one it's just being printed (adding numbers to an empty account). This is #federalreserve magic 🧙♂️ https://t.co/IQlAO7Vydz
[Wow, I missed this great TensorFlow news]
Anaconda now packages CUDA and cuDNN libraries as dependencies of tensorflow-gpu, so you no longer have to install these libs manually.
#MachineLearning#DeepLearning#TensorFlow#Anaconda#GPU via @anacondainc https://t.co/FwfvlJRfJm
After 2 years of development, we've just launched fastai v1, the first deep learning library with a simple consistent API across vision, text, tabular, and collaborative filtering data. Built on the wonderful @PyTorch v1 (preview released today)
https://t.co/6tFYNkvF8v
In Belgium, scientific papers that originate from public funds can now all be made public in open access, regardless of any contract with publishers. It is written in the law. And retroactive.