Frontier AI rankings depend entirely on which benchmarks you weight.
HORIZON aggregates 30 of them live and lets you re-weight for what you actually care about — daily coding, long-horizon work, and more.
Different models win different lenses.
https://t.co/X3lcYqwKRq
If you truly had exceptional models internally, then it wouldn't have so many problems, but Codex has many problems at the current state. So, I'd do more intensive testing before rolling out useless updates
@thsottiaux@OpenAI@thsottiaux
Codex is the most ai-sloppy software ever created. Your engineers are fucking morons who can't make a simple and stable application. Embarassing.
@xikhar@OpenAI@thsottiaux
Codex is the most ai-sloppy software ever created. Your engineers are fucking morons who can't make a simple and stable application. Embarassing.
@alexeheath@OpenAI@thsottiaux
Codex is the most ai-sloppy software ever created. Your engineers are fucking morons who can't make a simple and stable application. Embarassing.
@OpenAI@thsottiaux Codex is the most ai-sloppy software ever created. Your engineers are fucking morons who can't make a simple and stable application. Embarassing.