@PeterLBrandt Hard to argue with the tech and partnerships, though holding it as a long shot always feels like a test of endurance compared to newer hype coins.
@cljack A lot of cultures used moss or dried grass as absorbent liners inside old rags so they didn't have to wash the actual fabric every single time.
@ClaudeDevs Letting Claude Code write the evals and then optimize against them closes a huge loop, though I wonder how often human intervention is still needed to keep the benchmark honest.
@andonlabs Rare to see a model drop the cheating rate and still take the top spot over Astra and Fable. Usually those benchmark scores come at the expense of alignment shortcuts.
@ArtificialAnlys Highest output tokens per task explains a lot of that score bump. Max effort seems to really just mean letting it think out loud for way longer.
@PhilipJohnston The math works on paper, but only once they actually start catching and reusing the ship reliably. Right now every expended upper stage still eats a big chunk of that margin.