π₯New blog post: Are LLM agents time-aware? Can they predict wall-clock time of tasks? Can they estimate time they spent?
We study this on a range of tasks, including long-horizon ones (ProgramBench, PaperBench, DeepSWE, etc).
Joint work with my MATS mentee @MOfengenden!
@eliebakouch@maksym_andr Thank you! In our experiments, Opus 5.0 and Fable 5 also ran for more turns in Claude Code than in Codex (1.5x more). Testing them for clock-time wasnβt fair as GPT had much slower latency in Claude Code than the other way around.