@scaling01@tszzl Its because simple bench doesn’t have public available answer and its not influential enough for benchmaxing.
5.2 is what happens if we benchmax. Looks good on paper, Terrible in practice.
@quantumaidev@ai_for_success Worked with 5.1 Codex, 4.5 Sonnet and 3.0 Pro in git worktrees .
In my experience 3.0 Pro is the perfect middle ground for speed and intelligence. 5.1 Codex-High is Too much Intelligence too Little speed and 4.5 Sonnet is Too much Speed and Too little Intelligence.