@OpenAI You can code with OpenAI models on @Replit for free.
Our router also makes high-end models incredibly cost-efficient.
If you’re a business looking for an independent, multi-model alternative to Cursor, we’d be happy to fund your transition.
@jcyhc_ai@janhkirchner@jiaxinwen22 Prompt injection is one of major concerns right now and this much improvement with AAR tells us that what ai can actually improve in security.
We've pushed a version update to the Terminal-Bench dataset and leaderboard.
Terminal-Bench 4.0 calibrates task resources (time, cpu, memory), implements task fixes, and removes saturated tasks.
very good benchmark , we are also working on benchmarks for the AI ecosystem , we are doing research part for that and going to come up with something useful next month !
We're releasing Terminal-Bench-Science: a benchmark for evaluating AI agents on research workflows across scientific domains.
An ongoing Stanford-led community effort, built by the team behind Terminal-Bench together with scientific domain experts at research institutions worldwide. v0.1 has 70 tasks. Claude Opus 5 solves only ~30%.
1/n 👇