SIGMOD Systems Award for Amazon Aurora ·
Creator of AWS SimSpace Weaver ·
Prev. Microsoft Research + Amazon Research ·
High-school dropout → MSCS + AI, @UW ·
Teno, none of these are the same model even a week later. The post-train and serving stack are immediately under pressure to increase efficiency by reducing compute spend.
I already dumped more than a month of time into proving vibes with numbers, and we’re slowly seeing more people publishing reasoning stats every day.
Aren’t you getting tired of chasing model quality instead of predicable results?
@TokenGremlin Don’t ask the model to give you juice numbers. Just start measuring reasoning-token use directly. If everyone shares their numbers I think they would be quite surprised at the differences from launch week.
https://t.co/xEXpoaC0r7
After Anthropic made Fable 5 permanently available in subscription plans, I noticed a large drop in performance. The model felt dumber, and I couldn't explain why.
Measured five different ways, August delivered dramatically fewer thinking tokens than July.
@AndrewCurran_ I did not see this before. But I have thrown my hat in the ring to debate Scott with my own offer of $25,000 to his $5,000.
https://t.co/DqDKMeodyi
Scott, I don't know you, and I also have never argued in public before, but I am a betting man. And I and am willing to lay you odds, given the relative correctness of our positions, to debate you publicly, anywhere, anytime.
My position is: If AI is going to kill us all, it will be the direct result of human actions and decision making, not because of intrinsic AI properties. Further, the presumption of inevitability encourages fatalism, panic, and distraction from human responsibility and accountability. And from this lack of clear-eyed responsibility and accountability, we are driving towards our own demise.
I am happy to put up $25,000 to your $5,000 (ie, 5:1 odds in your favor) that I'll win by some standard of audience opinion change.
This is a 100% serious offer if you are interested. But I feel its time to have public pushback against unqualified AI hysteria.
Scott, I don't know you, and I also have never argued in public before, but I am a betting man. And I and am willing to lay you odds, given the relative correctness of our positions, to debate you publicly, anywhere, anytime.
My position is: If AI is going to kill us all, it will be the direct result of human actions and decision making, not because of intrinsic AI properties. Further, the presumption of inevitability encourages fatalism, panic, and distraction from human responsibility and accountability. And from this lack of clear-eyed responsibility and accountability, we are driving towards our own demise.
I am happy to put up $25,000 to your $5,000 (ie, 5:1 odds in your favor) that I'll win by some standard of audience opinion change.
This is a 100% serious offer if you are interested. But I feel its time to have public pushback against unqualified AI hysteria.
I think you've done enough calling us paranoid and preposterous. The next step is for you to defend your position in public against someone who will push back against it. I'm happy to meet you for a debate anywhere, anytime.
You're a world-famous veteran of dozens of debates against the world's top intellectuals, and I've never argued in public before, so adjusting for the relative correctness of our positions, if you're a betting man I'm happy to put my $5000 against your $1000 (ie 5:1 odds in your favor) that I'll win by some standard of audience opinion change. Let me know if you're interested and we can hash out details.
After Anthropic made Fable 5 permanently available in subscription plans, I noticed a large drop in performance. The model felt dumber, and I couldn't explain why.
Measured five different ways, August delivered dramatically fewer thinking tokens than July.
@attnisalluneed I do! But, for a time, Fable was providing unmatchable results and I was working on things I didn't think I would be able to get to for another few months. And then the rug was pulled...
I was consistently using an xhigh or max effot level, but when I looked deeper, I found that most invocations to the model were receiving little to no thinking tokens at all.
And when longer thinking runs did happen, they almost never reached published benchmark levels.
From the longer writeup, I started looking into this around ~Aug 1, after a week of struggling with regressions that started around when they made Fable 5 permanent.
I wasn't sure what was happening, but I was pretty sure something was different. I was definitely having a hard time reproducing the same results from Fable that I could before. Some things became nearly impossible and I had to give up on them.
Thank you! I guess it's obvious in retrospect. What isn't is the lack of transparency. Or, in the case of my data, that even P90 invocations received an order of magnitude fewer thinking tokens than published benchmarks at the same effort levels.
I've experienced real regressions and loss of capability that aligned with these drops. If it is happening to everyone, we are talking about real lost time and money.
Hi Mikhail. You may be interested in the full article, where I break down two months of measurements across a variety of dimensions. I analyze how inference compute has regressed over the period, how thinking scales sub-linearly with input, and how episodes of waxing or waning inference compute become predictive of thinking-token delivery for held out work.
It would be valuable to have an org the scale of Shopify look at their own corpus of Fable use. You might be surprised what you find if you look.
https://t.co/3DguFeAoRC
Hard to ascribe specific motive, but it’s easy to guess at likely incentives. In Fable’s case it was never meant to be permanent and all of the last-minute extensions felt like they were forced.
Reasoning compute is the one dynamic knob that both labs have that you can imagine is adjusted to meet capacity constraints, make a newer model look better, or protect margins. It’s also the one knob that fits them still being able to say they haven’t changed the model every time people ask if they are quantizing, or rerouting users under heavy demand.
After Anthropic made Fable 5 permanently available in subscription plans, I noticed a large drop in performance. The model felt dumber, and I couldn't explain why.
Measured five different ways, August delivered dramatically fewer thinking tokens than July.
Thank you, Angelo.
The first week was the hardest for me. I definitely second guessed myself several times until I thought through several instances that could not be explained away by simple variance.
I'm glad the data has resonated, but we will see if anything happens as a result. Changes will probably require large customers measuring their own workloads and demanding new transparency and commitments in their SLAs.
After Anthropic made Fable 5 permanently available in subscription plans, I noticed a large drop in performance. The model felt dumber, and I couldn't explain why.
Measured five different ways, August delivered dramatically fewer thinking tokens than July.
@natolambert This is what is going on (be sure to see the attached paper for sublinear input scaling and long-horizon thinking fragmentation):
https://t.co/xEXpoaC0r7
After Anthropic made Fable 5 permanently available in subscription plans, I noticed a large drop in performance. The model felt dumber, and I couldn't explain why.
Measured five different ways, August delivered dramatically fewer thinking tokens than July.
The only keyword that still has any influence on thinking is "ultrathink", and it no longer provides a fixed budget like it used to.
I've tried several things to try to get the model to think more. Almost nothing works at the invocation level. One thing that sort-of works at the turn level is asking it to think about a problem, write notes into a temp doc, read them back, and do that several times before delivering an answer. It can generate more aggregate thinking-tokens in a turn, but can't compensate for shallow sequential thinking overall.
Given that I’ve lost immeasurable time failing to reproduce working results that worked the day before? Yes, just tell me to come back later. Or tell me that inference compute is currently being reduced. Or tell me that you are degrading Model A because you want to push me to Model B, instead.
I could have just used a different model, did something else with my time, or taken the day off.
@schottge_menon Look at the next graph for a simple time series. Or look at the longer write up for more charts and data. This one is just the sensitivity analysis to show the effect survives regardless of how you try to measure it.