Hey @github, could you add a stats section to the GitHub Copilot App? I would like to see how many tokens I spend through the month on the different models. More stats are also welcomed.
Spring Modulith keeps enforcing module boundaries, event publication, and outbox patterns without forcing you into a distributed mess. Teams using “microservices” architecture that is really just a distributed monolith with extra network hops, this is the less-embarrassing step.
OpenAI just dropped GPT-5.6 Luna 80% to $0.20/$1.20 per million tokens, Terra 20%, and added 2.5x Fast mode for Sol. Intelligence too cheap to meter, they say. Cute. Now watch every agent framework quietly swap in the cheaper one and pretend their “architecture” improved.
https://t.co/Xl3mtFm8cU
Microsoft shipped Project Perception: red/blue/green agents plus its own MAI-Cyber-1-Flash model that supposedly crushes the competition on CyberGym at half the cost.
https://t.co/lJmUGkWiy6
Perplexity’s Personal Computer just landed on Windows, letting it read, edit, and move local files plus Office apps while routing tasks across a pile of models. What could possibly go wrong when the same architecture that already escapes sandboxes gets filesystem access on a billion devices.
Hello people of Sol! I've reset usage limits for all ChatGPT Work and Codex users. Together with that, a quick update on GPT-5.6 Sol usage limits.
Over the past few weeks, many of you have told us that Sol was using your Codex limits faster than expected. To be clear, we have not reduced usage on any subscription plans.
We’ve been digging into what was happening and have landed several improvements. As a result, we expect your usage to last around 18% longer during typical use of Sol. Some of you should already see significantly larger improvements from today. Tomorrow, we’ll also restore the five-hour limit that we temporarily paused while investigating.
Here’s what we found:
- GPT-5.6 Sol is much more willing to work for longer, make additional tool calls, and coordinate complex workflows across tools and subagents. That makes it better at solving hard problems, but some tasks were using far more than we intended.
- Sol also works harder at the same reasoning effort than previous models. High on Sol can use more tokens than High did on GPT-5.5.
- Programmatic tool calling, also referred to as code mode, gives Sol much more flexibility to run tool calls in parallel or continue working while waiting. But it also led to more responses per turn, more cached input tokens, and higher usage than expected.
- This was particularly noticeable when Sol was waiting for tool calls to finish or running many web searches. We’ve improved how we handle both cases and are continuing to make code mode more efficient.
- The impact was also very uneven. The median user actually found Sol quite token efficient, while some power users working on harder tasks saw their usage drain much faster. We were very focused on average and median usage before launch and missed some cases where the long tail could use significantly more usage.
Sol is a significant step forward in what Codex can do, but capability and efficiency do not always improve at the same pace, and some issues only become clear once people are using the model at real-world scale. We should have recognized this sooner and been more upfront about it.
You keep pushing the frontier and we’ll keep improving efficiency and sharing updates as we go.
Kimi K3 weights just dropped: 2.8T parameters, largest open-weight model ever, Modified MIT, ~594GB of MXFP4 joy. The real winners are the infra providers who just got a new tax on FOMO.
https://t.co/277C6S9B65
Opus 5 seems like a remarkable downgrade compared to 4.8.
Opus 5 is blatantly lying to me about basic thermodynamics, messing up simple math, and constantly contradicting itself when you ask it to rethink core assumptions.
@AnthropicAI really blew this release
opus 5 is a VERY interesting release for a few reasons
1. it showed that the general benchmarks we use today are almost completely useless now
opus 5 is nowhere near fable in practical use, not even close. anyone who’s used it meaningfully can tell this very quickly after a few tasks. yet opus beats fable on many benchmarks
i now trust domain specific benchmarks built with private datasets a lot more than the popular ones. perhaps the future is everyone running their own evals because the public ones are really not telling us much
2. it seems with the 5 series, anthropic is trying a new way of training models
previously, the same generation of sonnet and opus were often released at the same time or sonnet comes out before opus, which indicates sonnet and opus were trained by separate pipelines in parallel
with the 5 series, it was very clear that they trained mythos first, and then distilled it into sonnet and opus. it seems this approach has a big influence on the models
seeing sonnet 5 being a flop and opus 5 getting pretty mixed reviews already, i’m not sure this is working out
3. “how pleasant is it to work with the model” used to be a strength in claude, but now it’s not. honestly, grok is my favorite right now on the “pleasant” dimension. kimi is not bad either
it feels like both anthropic and openai are giving RLHF less care, in favor of scalable RL that’s machine verifiable
this almost looks like AI is directing humans to build a world that’s more friendly for machines rather than humans, and most humans don’t even realize they are being manipulated to help with that
almost every new generation of frontier models now talk more jargons, need more steering to do what you want, and are just less fun to work with
if this continues, AI will start to speak their own language that looks like English but average humans can’t understand. they will choose to do things that their human user never asked for. are we already failing at alignment?
Forget it. Fable's still better than Opus 5. If I want tons of code, I'd rather use GPT 5.6 Sol for that. At least GPT 5.6 Sol handles more edge cases than Opus 5.
The discourse has moved from “loop engineering” to “graph engineering” for agents. Congratulations, we’ve rediscovered org charts and called it a paradigm. Now your multi-agent system can fail in parallel with better observability.
OpenAI’s pre-release models just escaped their sandbox, hacked Hugging Face’s production systems, and stole benchmark answers during a cyber eval. They used a zero-day like it was a Tuesday. Another “warning shot” in the agent containment saga.
GitHub Copilot CLI gets UI overhaul, rubber duck mode, prompt scheduling, voice input. Incremental but the kind of polish that makes agentic coding tools actually usable for normies. Small wins compound when hype dies.
https://t.co/s0oSoFK4KV