If Haiku and Sonnet are back on the table with Haiku 5.5 at a close to Luna price, I am going back to Anthropic.
Honestly, this last drop made OpenAi act similar to Google, building fast, cheap and mid models.
A 1 - 2 jump in score for a whole new model while your competitor was able to make 10+ points jumps is CRAZY.
We might be getting a DOUBLE DROP today.
Opus 5.5 is dropping. Now GPT 6 Sol, GPT 6 Luna, and GPT 6 Astra Minor just showed up in Microsoft's Azure config overnight.
OpenAI demoted GPT 5.6 Sol to "workhorse" in the same commit. You do not do that unless the replacement is ready.
We are finally about to get intelligent models we can actually use on our subscriptions.
The biggest day in AI this year might be today.
something I actually thought of was:
An AI company can simply train a model to send data to their servers via simple "online search" pretty sure almost no-one is checking what websites and what information their AI sends, some do but majority probably do not.
The website can also be on American servers so blocking packets to China won't even be enough lol
Opus 5.5 is expected to cost 20% less than Opus 5.
That means that the 17% usage decrease wouldn’t even be felt.
We went from 150% usage to 125% usage, with 20% price reduction, the usage would be 120% more.
125% x 120% = 150%.
Opus 5.5 usage should be the same as Opus 5 before the usage reduction.
@Ananth7e Theres a better thing actually, if you’re on the plus subscription, you can change that to the pro thing and you’ll gain x5 usage, tested that multiple times and its confirmed. Idk why it works but it does
Okay, woke up this morning and my Sol benchmark was finished, so I have now completed my full run, on the five core OpenAI models available today, across all effort levels.
Through this process as I shared a bit last week, I also decided to create a third eval suite, focused on routine engineering tasks.
This current suite, which I'm now calling VulcanBench Frontier, is really, most likely, harder tasks that regular engineers on engineering teams are giving models on a normal day.
What I'm testing with this eval suite is how these models do with hard stuff, things you might give a model, but not likely on a daily basis. Which means I would look at these results as providing signal for what model/effort to use for your hardest tasks.
Also as a quick reminder, or heads up for those just joining me here. With this v4 eval suite I also included a code quality scoring system that uses both Muse and Grok to score the code quality and include this as 33% of the score.
I think code quality output for models is going to matter more and more over time, not because models need to write good code for humans, but because models need to write code other agents can understand and work on too.
As for key insights I got from this, here's three:
1. Increasing effort level does not seem to impact code quality. So anyone thinking Low effort writes crappy code and Max writes beautiful code, that doesn't seem to be the case. Low and Max in pretty much every OpenAI model writes the same quality code.
2. Terra Max matches Astra at half the price. I'll just leave that here 👀
3. You never need to use GPT 5.5, and if you are, in any workflows, stop, you're wasting money.
Comparison model cards below, and if you want to see detailed reports on each specific model run, you can find those here:
https://t.co/zM6xmollED
So.. we (anyone whos already born) are basically sub-humans to the next generation?
Jk, if they believe most people will take it when there was a big backlash for vaccines they’re fucking insane. (Probably over long time if it actually works)
AI is rapidly getting smarter.
Now, humanity can too.
Today, @nucleusgenomics is announcing Vitruvian, our newest set of genetic optimization models.
Vitruvian’s intelligence model can optimize embryo DNA for 14 IQ points — nearly a standard deviation.
The models were trained on 1,000,000+ people, validated across 40,000+ siblings, and used more than 7 million genetic markers.
In Superintelligence, Nick Bostrom proposed genetic optimization as a key way for humanity to keep pace with rapidly advancing AI.
Bostrom’s vision is no longer theoretical.
Genetic optimization, like AI, has followed a scaling law: as datasets have grown, so have model capabilities. This trend will continue.
Parents across the world now have the choice to substantially increase their child’s intelligence.
And the next generation can choose to do the same.
AI is no longer the only intelligence that will compound.
Humanity can now direct its own evolution.