Dollars Matter, Scores Flatter. 💰
Can LLM agents really do high-valued experts work end-to-end — not just “solve a question”?
Introducing $OneMillion-Bench benchmarks real expert work, priced in dollars 💵
Why it matters + resources 🔗
If we want agents in finance/law/medicine/engineering, we need evaluation that matches real workflows + real cost of mistakes.
What moves the needle 🔧 Tools and evaluation design matter a lot.
• Official scaffolds > OpenRouter > no search (consistent gains)
• Rankings are stable across different judge models (robustness check)
• We also stress-test temporal sensitivity and test-time scaling behavior.
How Does We Grade? ✅ ❌
We only count value when the output is deliverable, not “partially correct.”
• 15–35 expert rubrics per task + negative rubrics for hallucination/unsafe/non-compliance
• Asymmetric weights (-20 to +10) to reduce reward hacking
From high-economic-value tasks to verifiable rewards, we help frontier teams push model capability beyond benchmark optimization and toward real-world cognitive performance.
We design expert-level tasks, build rigorous verification systems, and turn high-density human knowledge into scalable training signals for advanced models.