GPT-6 Astra and Claude's 5 series just landed on the AI Frontier - https://t.co/6ZqljIhKBh
People expect a new frontier model to make the old ones obsolete.
Instead, something much more interesting happens:
GPT-6 made the case for using multiple models stronger.
Here’s what we found ↓
We just added GPT-5.6 Luna, Terra and Sol to the AI Frontier: https://t.co/6ZqljIhKBh
The most interesting result isn't that Sol is the strongest model.
It's this:
2× Luna > 1× Terra
4× Luna > 1× Sol
Here's what happens when you stop benchmarking models one run at a time 🧵
I'm excited to share some research I've been working on for the last few months called The AI Frontier!
It allows you to get more out of LLMs with less compute. I studied 44 models x 16 benchmarks and found that dynamic model selection can reduce errors by an average of 45% and save 95% of token costs 🤯 🚀
Interactive Site: https://t.co/NikTDGfRRh
Academic Paper: https://t.co/a5LEM7fvhY
LLM providers charge you for input, thinking, and output tokens yet you only control one of them. Some LLMs may have higher quoted token costs but are more efficient, answering the prompt with less thinking and fewer output tokens, resulting in a lower real cost. See real LLM costs with our `Cost Insights` view.
👋 Hey everyone! I figured it was finally time to join the research conversation here. I'll be posting about all things #ML. I'm looking forward to connecting with other ML researchers and practitioners!