FrontierBench is live! Excited to have been a task author and reviewer on this effort, with @boolean_ai as a data partner.
The team was rigorous in keeping the quality bar high at scale, across a wide range of domains. Congrats to @ryanmart3n and everyone involved on the launch!
We ran a frontend eval from an in-progress internal benchmark. Kimi K3 is not at the same level as current frontier models like GPT 5.6 Sol or Fable 5. It's closest to Opus 4.7 on this eval so it's 3 months behind frontier. An impressive result nonetheless.
Evaluating frontier models is going to increasingly require very high taste and in-depth domain expertise.