Nerf Bench from @bridgemindai tracks how much of its launch power every model still has. Every model is within normal variance (±10%) of its launch. https://t.co/4IOCb7ulKL
Did Anthropic nerf Claude Opus 5.5?
The first NerfBench results are live.
We retested Opus 5.5 and GPT 6 Astra against their own launch scores.
Claude Opus 5.5: 99.2% (-0.8% vs launch)
GPT 6 Astra: 102.8% (+2.8% vs launch)
Verdict: No nerf detected.
Opus 5.5's small dip and GPT 6's small increase is within normal variance.
More models and more frequent retests are coming.
NerfBench only gets better as we collect more data.
Which models should we retest next?
@almmaasoglu It's just the honey moon period. By now I believe it's just all the same models and we're living between nerf and buff cycles of the same model....
@h_wafeek give it time... software leaped and had to work with peak hardware design of the time, another hardware leap and yeah, this could be the new transistor
@Lon@ArtificialAnlys we need a benchmark called State of the Absurd (SOTA) over time. All providers set the lever to highest so you can measure it at its peak then over time it declines
After Anthropic made Fable 5 permanently available in subscription plans, I noticed a large drop in performance. The model felt dumber, and I couldn't explain why.
Measured five different ways, August delivered dramatically fewer thinking tokens than July.
AI films and videos have gotten insanely good and i want to go deep on how they're made
the prompting, the editing, the whole pipeline
who are the best people to follow for this? tutorials, courses, anything
ISRAELI NGO networks have FLOODED EUROPE’S MIGRATION ROUTES, providing migrants MAPS, MONEY and LOGISTICAL SUPPORT to move deeper across the continent.
Probably nothing.
this is Dario and sama's worst nightmare -
A law firm buying Nvidia servers.
Latham & Watkins is building an in-house AI stack..
this is US’s second-largest law firm with $8.3 Billion in revenue last year
And now it has -
- Nvidia hardware it controls
- open-weight models it can fine tune
- proprietary legal data it is trusted to protect
- infrastructure only Latham employees can access
A law firm has decades of contracts, negotiations, client context, legal reasoning, and institutional knowledge.
They dont want to give all of that away to OpenAI or Anthropic in exchange for expensive tokens..
And on top of that - risk their data being used to train frontier models..
This will happen more and more now..
The big AI labs have no moat.. nothing protecting their largest customers from moving on..
The biggest companies in the world will -
- own the compute
- own the data
- own the workflow
- fine tune the model around their business
- switch providers when the pricing or quality changes
Dario and sama want to make this illegal by bringing in regulation.. and become the AI overlords..