I disagree with @elonmusk on this one. Problem is not only generalization, but also the creative part of surgery. What happens when the robot is faced with a scenario it wasn’t expecting? And to be clear, I think robots would eventually be able to work well within medical sciences and other sciences, but high-risk fields such as surgery and family medicine will require more time than 3-5 years. Theoretical capabilities do not translate to deployable capabilities and Elon knows this.
ELON MUSK'S JAW-DROPPING PREDICTION:
"DON'T GO INTO MEDICAL SCHOOL."
ELON MUSK: "YES. POINTLESS."
IN 3 YEARS (2029), OPTIMUS ROBOTS WILL BE BETTER SURGEONS THAN ANY HUMAN ON EARTH — AT SCALE.
BY 4–5 YEARS? NOT EVEN CLOSE.
THE BEST MEDICINE IN THE WORLD WILL BE FREE — BETTER THAN WHAT THE PRESIDENT GETS TODAY.
THIS IS WILD.
@kimmonismus I can agree to this and throw Opus 5 in the mixed. Yet I don’t think staying with sonnet 4.6 was feasible either. So I guess I would have preferred Anthropic had spent more time “fixing” sonnet 5 than releasing Opus 5.
Is there any data/paper to back these claims? And I don’t mean just numbers in the vacuum. I mean task benchmarks quantifying cost per task on both context window (258k vs 1M) where once compact say 3 times and the other just fills its own up to 750k. I’d be curious to know the trade-offs.
@gdb This is awesome! @ClaudeDevs you should definitely do a usage limit reset. It would allow us to stop reading news and keep building more. Can we get a reset? 😃
@thsottiaux Good to know! I recently launched a task with my own agentic framework that guides GPT-sol and other models to run for days if necessary to accomplish their goals. That reset was appreciated—thanks.