I’m very excited to share @nvidia Kumo Tabular, a new family of foundation models for tabular data.
Kumo Tabular establishes the new Pareto frontier across the entire accuracy–inference-time tradeoff.
Just as importantly, we are releasing it openly: open weights, open-source software, and a permissive license for commercial use.
HuggingFace: https://t.co/9j9LM0z6jB
GitHub: https://t.co/D8MkunxT7v
MiMo-V2.6-Pro debuts as the top open weights model on the Artificial Analysis Intelligence Index (46). At $0.13 per Intelligence Index task, it lands on the Intelligence vs. Cost per Task Pareto frontier
@Xiaomi has just released MiMo-V2.6-Pro, an open weights model with major advances in intelligence over its predecessor, MiMo-V2.5-Pro (Intelligence Index: 26). Despite the improvement, it retains the same attractive pricing at $0.435 per 1M input tokens (with a 99% cache-hit discount) and $0.87 per 1M output tokens. This makes MiMo-V2.6-Pro one of the most cost-efficient models to deploy.
MiMo-V2.6-Pro is an MoE model with 1.02T total parameters and 42B active parameters. Stay tuned for additional analysis of the model.
Check out MiMo-V2.6-Pro full benchmarking breakdown here: https://t.co/czBJhQKuWJ
A nice compilation of standard formulaic representations (and references) of many of the traditional instruments, statistics, and strategies (all referred to as “strategies” in that publication).
🚀 Introducing DeepSeek-V4.1-Flash: smarter, faster, more efficient.
🔹 Introducing the smallest model in our new architecture family, with native visual understanding.
🔹 Designed for greater capability, faster inference, higher throughput, and scaling to larger models.
1/6
Hi Astra users. A reset and a quick update on quality issues that have been posted around.
Working with some of you, we have found and fixed the following issues:
- Some skills written for previous models were triggering too often or preventing the model from checking its work.
- An opt-in context management experiment that could cause early stops or replies to older messages. We've disabled it. Our rough estimate is that 4-5k users were affected by this experiment.
- We've also removed some badly configured engines that resulted in a measured quality degradation for a long tail of traffic flowing through them.
We’ve also made some more minor improvements and things should feel significantly better across the board. More consistent follow-through, better tracking of your latest message, and better checks on the work as it’s going through the motions.
The examples posted and all the users who worked directly with us were incredibly useful in helping fix things quickly. Always grateful for this incredible community.
And of course, a reset is also landing by midnight today.
More and more indications that existing skills and agents (and “harnesses”) need to be updated for the new models.
@OpenAIDevs have just shared this high-level overview: https://t.co/J8uW2VMPIs
And @thsottiaux confirms there are some necessary updates in the skills/agents.
Existing SDD frameworks can become legacy overnight and only your token burn will tell. That’s why it’s important to read and review all your skills and vendor third-party skills/agents.
Hi Astra users. A reset and a quick update on quality issues that have been posted around.
Working with some of you, we have found and fixed the following issues:
- Some skills written for previous models were triggering too often or preventing the model from checking its work.
- An opt-in context management experiment that could cause early stops or replies to older messages. We've disabled it. Our rough estimate is that 4-5k users were affected by this experiment.
- We've also removed some badly configured engines that resulted in a measured quality degradation for a long tail of traffic flowing through them.
We’ve also made some more minor improvements and things should feel significantly better across the board. More consistent follow-through, better tracking of your latest message, and better checks on the work as it’s going through the motions.
The examples posted and all the users who worked directly with us were incredibly useful in helping fix things quickly. Always grateful for this incredible community.
And of course, a reset is also landing by midnight today.
Just experienced a weird @OpenAIDevs Codex usage glitch while running three GPT-6 Astra (medium) driven sessions. My 7d (94% and 5d9h remaining when started) got depleted somehow and half an hour later I’m at 12% and 1d14h remaining quota. Fable 5.1 manages to last way longer.
Looking at local optima churn with GPT-6 Astra and GPT-5.6 Sol as well as Fable 5.1, it’s time to revise our understanding of instructions such as “meticulous” and “rigorous” review, “impeccable” implementation, and similar. For the earlier models, it was often necessary to get them to an acceptable level, with the newer models it’s a hazard that forces them to loop around saddle points trying to squeeze the last bit of “quality”, forgetting about the overarching goals and structure of the workstream.
Reading on some of the first experience with Astra. Awesome outcomes already but skills and harnesses will need to be tuned to Astra’s quirks. Other than I had hoped, GPT-6 Astra appears to go into the perfection-loop/asymptote/local-optimum/saddle-point behavior even more aggressively than GPT-5.6 Sol. One must break the loop early and give it an escape hatch. This is what I’ve done with my reviewer loop.
I’ve been working with a reviewer loop - Sol/Opus/Fable plan/orchestrate/coordinate/manage and spawn predefined or generic subagents with defined models and effort, and GPT-5.4 mini estimates the complexity and suggests a reviewer provider and model. I’ve expanded it to something like what Matt describes.
Also a nice feature of the Codex harness is that sessions can now communicate with each other.
Curious to see how Astra works with larger contexts.
I noticed Fable 5.1 struggle quite a bit with this approach, second-guessing my reviewer plugin skills and bypassing them. It sort of longs for more decision autonomy but fails on a bit more complex tasks after hours without steering.
A lot of people are asking how I pulled off these super long-horizon builds with Astra.
Astra is extremely powerful, but by default it struggled with a task this difficult. I tested a bunch of approaches to get past this, and the one I landed on is something I'm calling the Manager Loop.
It's basically a couple of tricks we used to use with much less capable models a couple of years ago, with a few new ideas layered on top. Turns out that when you put those together and apply them to Astra, its ability to do extremely difficult long-horizon tasks goes up dramatically.
Here's how it works:
1. Launch an agent (I'm calling this one the "manager"). Chat with it about what you want to get done, and have it build a massive checklist of to-dos, then break that checklist into phases.
2. The manager then spawns a second Codex agent in a separate thread (the "implementer"). The two agents can message each other.
3. Put the manager in /goal mode, and tell it to run each phase on the implementer in /goal mode.
4. The manager messages the implementer: "/goal Complete phase one completely, extremely well." The implementer doesn't stop until that phase is done, then messages the manager back. The manager tells it to start phase two. They repeat until every phase is finished, completely autonomously.
Why I think this works: over a long-horizon task, Astra tends to asymptote. It gets way further than previous models, but at a certain point it kind of just stops improving against the goal as quickly as it did before. It gets stuck in the minutiae, focusing way too much on small details, and overall progress stalls. The Manager Loop forces it to work piecemeal, one phase at a time. It's essentially how a human would steer a model, except the model is doing the steering for me.
That's actually how this started. I was having the model write the checklist and break it into phases, and then I was doing the manager's job by hand. At some point I thought, "Wait, why can't I just get a separate AI to do this?" That's what unlocked full autonomy, which is super useful.
A wording detail that seemed to matter: I ask for each phase to be done "extremely well," not "perfectly." Maybe I'm reading too much into it, but asking for "perfect" sent the model right back into the minutiae. "Extremely well" implies it's allowed to move on once it's good enough, and that worked better in my testing.
One more trick that I think helps (this one is more of a hunch, but it was useful for me): have the implementer build a simple HTML page with the full checklist on it. The implementer checks boxes off as it goes and updates a counter, and the page has a chart of # of boxes ticked over time.
Obviously the boxes aren't all equal, but it forces the model to notice things like "I haven't made progress in a while, time to move on." You can even put this in the prompt directly, like: "if you haven't ticked a box in X amount of time, move on". That helps a lot.
I also ran 96 sub-agents at a time. You can change this in your Codex config (or just ask Codex to change it).
This got me far better long-horizon performance than anything else I tried. I'll be sharing more in the coming days!
@thsottiaux@thsottiaux this is quite generous, thanks. How will it roll out? API already usable, but I’d rather experiment with my subscription allowance first. Any ETA for the Pro plan users? Or is it geo-bound?
GPT-6 Astra makes significant gains in the Artificial Analysis Coding Agent Index, scoring equal to Fable 5 at lower cost. In the Intelligence Index, it uses fewer tokens than GPT-5.6 Sol for similar performance, but this is outweighed by higher prices
Pricing is 2.5x GPT-5.6 Sol’s current prices across the board, up from $4/$20 to $10/$50 per million input/output tokens, with the same 90% discount for cache reads and 25% premium for cache writes.
We see distinct stories across our two flagship Indices. In the Artificial Analysis Coding Agent Index, GPT-6 Astra equals Fable 5 at less than half the cost, driven by significant token efficiency gains. In the Artificial Analysis Intelligence Index, GPT-6 Astra is more token efficient than its predecessor for similar performance, but this is offset by the price increase.
Artificial Analysis Coding Agent Index - key takeaways:
➤ Rivals top models: In Codex, GPT-6 Astra scores 67 in the Index - approximately equal to Claude Opus 5 and Fable 5 in Claude Code, and Muse Spark 1.3 in Muse Code. Fable 5.1 in Claude Code leads the Index with a score of 70.
➤ 70% more token efficient than GPT-5.6 Sol: GPT-6 Astra sees a substantial improvement in token efficiency, using one third of the tokens compared to GPT-5.6 Sol (max) in the Codex harness, and one fifth of the tokens of Claude Opus 5 (xhigh). Various effort levels of the model occupy the Pareto frontier of token efficiency.
➤ Leads Coding Agent Index cost efficiency frontier: At max effort, GPT-6 Astra costs about the same as GPT-5.6 Sol (max) while scoring 2 points higher on the Index. Per task, the model is less than half the cost of Claude Fable 5, for the same score.
Artificial Analysis Intelligence Index - key takeaways:
➤ Sits beside GPT-5.6 Sol in Intelligence: GPT-6 Astra scores equal to GPT-5.6 Sol in the Index at 61. This is 5 points lower than Claude Fable 5.1 (max with fallback). The model also trails Meta’s newly released Muse Spark 1.3 (max).
➤ ~10% fewer output tokens, offset by price increase: GPT-6 Astra defines a new Pareto frontier for Intelligence Index vs Output Tokens per Task - with a ~10% reduction in token use at max effort compared to GPT-5.6 Sol. However, due to the 2.5x increase in price, the model is 75% more expensive per task than its predecessor at max effort.
➤ Hallucinates half as much as GPT-5.6 Sol: GPT-6 Astra sees a large jump in AA-Omniscience, our knowledge and hallucination benchmark. This is driven by a significant decrease in hallucination rate from 92% to 51% at max effort. Unlike some models, this improvement does not come at the cost of accuracy - Astra increased accuracy by 4 points at the same time.
➤ ~80 point gain in AA-Briefcase Elo: GPT-6 Astra improves ~80 points in AA-Briefcase, our frontier long-horizon knowledge work evaluation. Models are tested on multi-week projects, with many linked tasks and thousands of source files. Astra sees a significant increase in both rubric scores and Analytical Quality Elo in AA-Briefcase compared to its predecessor. In the other direction, we observe a reduction in Presentation Quality Elo, where GPT-5.6 Sol (max) still leads all models.
➤ Mixed progress on other evaluations: The model sees a 6 point gain in Humanity’s Last Exam, a long-standing evaluation with emphasis on mathematics, science, and humanities. This is offset by a drop of ~80 Elo points in GDPval-AA v2 - a benchmark we adapted from OpenAI’s dataset measuring economically valuable tasks across 44 occupations. We also observe 2-3 point regressions on other evaluations across a mix of capabilities, including reductions in τ³-Banking (customer support), SciCode (Python problems in a scientific domain), and AA-LCR (long context reasoning over large documents).
Congratulations @OpenAI and @sama on the launch!
We are starting to release GPT-6 Astra and we are doing it as carefully and quickly as possible. It was very important to us that we bring it to all Plus users and not only Pro, Business and Enterprise.
It will take a few days for the rollout to complete and behind the scenes many novel systems will operate at scale for the first time and we are bringing a lot of compute up.
It is pure magic.
https://t.co/WUdXA3xSoh
meet @mostik_ai!
what happens when you put 12 PhDs in one room for four months? first place on the ARC-AGI leaderboard, which I can't say much about while the competition is still running. and this, which I can.
everyone's arguing about whether open models will catch up to frontier models. we think it's the wrong question. here's the one we pose: why does a frontier model have to generate your answer at all, when the only thing you need from it is the reasoning?
we do this by enabling models to communicate in latent space. through our protocol, hidden states pass straight from a frontier model into a small one running on your infrastructure -- no text between them, and neither model is fine-tuned. two models from different families, sharing reasoning, both left untouched.
how do we know it works? we tested it on a setup where a 753B model reads the problem, and a 4B edge-class model writes the answer. with this approach, we get results 80% as accurate as the frontier model, but at 20x faster performance.
we're committed to preventing frontier model lock-in and are already partnering with inference providers to accelerate open-weight adoption. we've done this between 15 of us, in four months, 12 PhDs and a Fields medalist, backed by @generalcatalyst
WIRED has the first external account of the company and the work: https://t.co/tP8nItCsDl
full writeup, the setup, and all the numbers: https://t.co/C9NZ5vtV1V
Curious how @OpenAI GPT 5.6 Sol fails to deliver a coding outcome after 12-20h sessions and after weeks of multiple sessions, and entangles itself in meaningless scaffolding and modalities. Skills like https://t.co/YxQQEmZaE4 don’t help as much as I’d hoped. I still don’t get any workable result and have to steer Sol. If it weren’t for @thsottiaux’s generous resets, I would’ve had to cancel my subscription. I see how much work was put into the Codex App harness but I’m not the target audience there.
Because of this lack of outcome in reasonable time, I’ve decided to downgrade to ChatGPT Plus only. I have 5 Codex sessions running in parallel, on good plans, 15h in, and they all are at step 1 still, churning and regurgitating and recompacting. This is human in the loop and frustrating. I’ve been tolerating this for the past couple of months, it’s a waste of my time.
With @ClaudeDevs Opus 5, I get an outcome quickly enough, sometimes after 10h, sometimes within an hour or so - on the same tasks. It tends to hallucinate retrieval and makes more mistakes than Sol. But now, with the new and more expensive Fable 5.1, I’m surprisingly able to get reasonable results while consuming less token allowance than Opus 5 for equivalent tasks. But it’s a quite bit dense in some other ways, ignoring instructions on a different level, ignoring skills.
Sol is a great reviewer model - it’s not lazy like Opus is, but it tends to miss the forest for the tree.
I very much hope that OpenAI’s Astra gets a post-training better tailored to the real-world enterprise software engineering and delivers a better coding experience.