PHYBOT is using robot soccer as a full-contact locomotion test.
And having the referee casually kick the robot to test its recovery is hilarious. Half stress test, half trolling.
It's crazy how far AI animation has come 🤯
I made this almost one-minute animation with just two Seedance 2.5 generations. That's it haha.
Prompt below ⤵️
Demis Hassabis is the greatest scientist on earth. Whatever he does next will be towards the end of eventually being able to explore Alpha Centauri aboard his own personal starship. Goodspeed Demis, the world is behind you.
@demishassabis
Everyone says the latest AI agents will be "job-ready" soon, especially after the release of Fable 5 this week. But is that really the case?
Over the past many months, my group and collaborators have been building Agents' Last Exam (ALE), a benchmark designed to test exactly that claim on real digital labor-market work.
My group and collaborators previously have created many of the benchmarks the field runs on, including MMLU, MATH, CyberGym, and ExploitGym. Today, I'm excited to share Agents' Last Exam (ALE): a rolling benchmark that measures whether AI agents can actually perform economically valuable work across a broad range of real-world domains.
With ALE, we evaluated Fable 5, GPT-5.5, Composer 2.5, and other frontier agent systems across more than 1,500 expert-sourced tasks spanning 55 occupations.
The result is both impressive and sobering.
Today's agents can solve a meaningful fraction of professional tasks. But when we look at the hardest tasks, the ones requiring sustained reasoning, deep domain expertise, and reliable execution over long horizons, they are still far from human-level performance.
On ALE's hardest tier, every frontier agent we tested, including Fable 5, achieved a 0% success rate.
The age of useful agents is here.
The age of truly job-ready agents is not.
We hope Agents' Last Exam (ALE) will serve as a new guidepost and north star for developing agents capable of reliably performing economically valuable work across a broad range of domains.
🧵
@demishassabis the news of Demis Hassabis stepping down from the head honcho position at DeepMind literally staggered me. I wonder if this raises or lowers the chances of Demis becoming the dictator of AGA? We might've lost our best chance for the good ending.
@demishassabis knowing Demis, he made the call a year ago for DeepMind to completely neglect the product side and focus 99% of their efforts on pure, frontier-pushing research
It might seem like DeepMind is lagging, but they're actually accelerating at at max velocity towards RSI