@teortaxesTex maybe it has some non-negligible positive effect? as in: “recognize that the past assistant messages were written by a similar but worse version of you, so try harder / don’t take their arguments too seriously.” (I’m not claiming this was Tao’s intent, lol)
@teortaxesTex I guess the bets are (1) accelerating RSI or (2) exploiting the weakest link, the human. addressing (2) will mean removing humans from all the loops entirely... which will create a different set of problems, lol
Claude Fable 5 (max) scores 91.9% on WeirdML, a new clear best score.
For comparison, if we take the SOTA score on each task, we would get 93.5%, so Fable very consistently scores within a few % of the SOTA in each individual run.
It achieves SOTA on 7 of the 17 tasks, and the single worst run is only 7% behind the SOTA.
It got only 2 runs per task (as opposed to the usual 5), which makes it more impressive that it gets to SOTA in so many tasks (but also makes it easier to avoid one or two bad runs). I'll show more stats soon.
GPT 5.6 Sol (high) scores 88.8% on WeirdML, narrowly ahead of Claude Fable 5 at less than half the price.
Sol seems very solid and, out of the 85 (17 x 5) runs, not a single one scored less than 60% of the SOTA score on that task (Fable had 2 such bad runs).
The consistent scoring means that the bootstrap error bar is very small. While the uncertainty is probably somewhat underestimated, this does mean that we are still sensitive to different model capabilities, even as the benchmark gets more and more saturated.