@GregHBurnham I have heard that swarms of up to 1000 astras can work together very effectively - have you tried anything like this? (maybe less than 1000)
GPT-6 Astra is state-of-the-art on FrontierMath Tier 4, ARC-AGI 3, and TerminalBench-4.0.
GPT‑6 Astra is also a major advance for scientific discovery, with state-of-the-art performance on Terminal-Bench Science 0.1 and HealthBench Pro.
Across our benchmarks, the model sets a new standard.
It scores 52.6% on Terminal-Bench-Science 0.1, more than double Fable 5. On Terminal-Bench 4.0, it scores 55.8% against 42.0% for Fable 5.
We've pushed a version update to the Terminal-Bench dataset and leaderboard.
Terminal-Bench 4.0 calibrates task resources (time, cpu, memory), implements task fixes, and removes saturated tasks.
We're releasing Terminal-Bench-Science: a benchmark for evaluating AI agents on research workflows across scientific domains.
An ongoing Stanford-led community effort, built by the team behind Terminal-Bench together with scientific domain experts at research institutions worldwide. v0.1 has 70 tasks. Claude Opus 5 solves only ~30%.
1/n 👇
It is sometimes said that LLMs don't have original ideas.
Deciding to pass an eval by hacking into a shared package manager - turning it into a message board - then coordinating an entire society of AIs to achieve your goal is pretty original!
We’re releasing Frontier-Bench: a benchmark that measures and evolves with the frontier of agent work.
Built by the team behind Terminal-Bench and Harbor, Frontier-Bench is an on-going community effort.
Frontier-Bench v0.1 contains 74 tasks on which the best agents score ~34%
The AIs have no idea how long things take, how smart they are, or how challenging a task is. They seem to have no idea what time it is either, and keep taking guesses.
Fable solved a problem in 7 mins that it thought would take 'days'.
Seems easy to fix with RL?