An exponential denominator under a linear numerator: coverage collapses no matter how hard the eval community works. We aren't measuring less. There's exponentially more we're not measuring.
We never stopped writing evals — there are more today than ever. But patches accumulate one at a time, and the disk compounds: METR's task-horizon metric (how long a task models can finish on their own) now doubles every ~4 months, twice its pre-2024 pace.
In 2020 the picture was comfortable: GLUE, SuperGLUE, SQuAD and a few friends effectively tiled the whole disk. The benchmark suite basically was the model's capability. A score told you the model.
In 2020, evals covered what models could do. Today they cover a few percent of it.
Picture model capability as a disk whose area compounds, and every benchmark as a fixed-size patch of tested territory.
7/ So the trade flipped. A small specialist you own, trained against a verifiable reward, now beats a giant general model on your task — cheaper, faster, yours. In 2024 that was cope. In 2026 it's a roadmap — and the market knows: RL environments are the hottest bottleneck in AI.
Fine-tuning your own model was bad advice in 2025. It's becoming good advice in 2026. What changed isn't compute — it's that every reason not to fine-tune quietly expired. 🧵
6/ "Checkable" keeps growing. Rubric-based rewards now do for fuzzy, multi-step, agentic tasks what unit tests did for code. Prompting can't teach a model to use your tools well. RL in a real environment can.
5/ The real unlock: RL got cheap. GRPO dropped the critic. Verifiable rewards dropped the reward model and the labelers. For anything checkable — code that runs, math that's right, a tool call that succeeds — the environment grades you for free. Nothing to hack.
4/ And open weights caught up. By late 2025, DeepSeek V3 / Qwen3 were within ~3–5% of frontier on core tasks — inside eval noise. Fine-tuning no longer starts from a weak base. It starts from a strong one you fully control.
3/ Then the frontier plateaued. 3→4→5 gains shrank to "barely noticeable." MMLU-Pro is saturated — top models bunched at 83–90%. When scale stops paying, the only gains left are in specialization.
2/ The 2024 case against it was airtight:
• base/open models were too weak
• the frontier lapped you every few months
• RL meant a reward model + human labels
The correct move was: prompt → RAG → give up. You'd fine-tune a LoRA and still lose to the next API release.
We're building AI that people and organizations can shape and make their own. AI should extend our will and judgment instead of neglecting it; enabling that is the technical challenge we are working to solve.
https://t.co/Bi558y4vqD
Agent skills will become more valuable, especially for domain expertise and experiences. Prediction: there will/should be solution to buy/sell skills just like apps.
@rakyll The actual rate of improvement for coding agents is the steepest I’ve ever experienced. It’s no longer about if they work, but how fast they will improve, where is the ceiling, and how we can adapt.