Great to join Scale AI’s Research Meetup: Benchmarking Coding Agents at our NYC office!
I really enjoyed the conversations on the frontier of coding agents with our guest speakers:
- Kilian Lieret (@KLieret) from Meta FAIR on ProgramBench
- Prof. Eugene Wu (@sirrice) from Columbia University on BranchBench
And from Scale AI:
- Miguel Romero Calvo on Terminal-Bench
- Mohit Raghavendra on SWE-Atlas
An evening of discussions on how we build better benchmarks and ultimately better coding agents.
https://t.co/Xng36Lc1Uk
@yifanzhang_ Do you think this model can solve (almost) all IMO 2027 problems without weight updates? Supposedly, the problems have similar difficulty across years, and the foundational knowledge required is very similar.
Manager ALERT: the best way to speed up your team might be to optimize yourself out.
In multi-step workflows, a manager sits between every step — reading all outputs, aggregating, deciding who does what next.
We showed this with computer-use agents: plan the input<>output dependencies between agents upfront, and you no longer need a powerful manager merging sub-agent work at every step.
There's a deeper reason this matters for computer use: you can merge two files, but you cannot merge two live machines.
Our "Spine-Branch" framework plans around this constraint from the start: one spine carries live VM state end-to-end; disposable branches gather artifacts in parallel, then get thrown away. No merging, ever. No manager relaying state between steps.
+6–16.5% success, 34–70% cheaper, ~5× less overhead.
Great work by our intern Mian Zhang @_Guuuuuuuu_ at @scale_AI
🔗 https://t.co/MTidcMQ8pN
@_Guuuuuuuu_@ManasiSharma_@danielyuez@Yminglai@shi_kejian
Nature Reviews Drug Discovery, this month, on a decade of AI in drug discovery: "evidence of their clinically relevant impact is, so far, disappointingly limited." https://t.co/ilU6n4G6D5
So can today's agents change that? We tried to measure it, not assert it.
Scale AI and Phylo built DrugDiscoveryBench: 82 verifiable tasks authored by pharma scientists and biomedical researchers, spanning target identification, patent mining, and structure activity analysis. Real artifacts the agent has to go retrieve.
Best agent clears 51.6%. Only about half the tasks are in reach for any single agent.
The interesting part is the failure mode. It is not missing knowledge. Hand the agent the expert's method as a hint and 80 of 82 tasks get solved by someone. What agents lack is the follow-through to carry a long scientific workflow to the end without dropping a constraint or skipping a step.
https://t.co/vd9tGcV7O9
@afeyzaakyurek@TuXinming
Real time translation keeps getting better and cheaper. Speaking the language yourself keeps being worth it. Same three reasons.
Speed: fluency is instant, the translator adds a hop.
Depth: a translator gives you the literal words. A fluent speaker knows the connotation, the register, what will actually land in the room. That is the difference between an aggregate sense of a thing and the first three search results.
Reliability: every round trip through a tool is a chance to lose something, and across a long conversation the losses compound.
Which is why the old line survives: "talk to someone in their second language and you reach their head, talk to them in their mother tongue and you reach their heart."
In the new era, reviewing and knowledge absorption are major bottlenecks in mathematical research, an observation that closely parallels what we are seeing in agentic coding.
📣Call for contributions + co-authorship!
RSI Bench is our ongoing effort to evaluate whether AI agents can develop the capabilities required to advance AI R&D.
Got a frontier AI research problem on your mind? Submit it to RSI Bench, and set the standard the whole industry uses to measure progress!
Contributors receive:
→ Co-authorship on the RSI Bench research paper
→ $2,000 per accepted task for the initial 50 tasks
→ Modal compute credits to build + iterate
→ Access to the RSI Bench research community
Register below for full requirements.
Had a great time talking at @ScaleAILabs. New levels of model capabilities require new benchmarking paradigms.
Also really enjoyed talks by @sirrice@mohit_r9a and @MiguelR33478246!
Thanks to all the organizers and to @liuying04
A neurosurgery resident at Peking Union Medical College Hospital (PUMCH) used ChatGPT 5.6 to solve a major open problem in numerical linear algebra. The author, Shanmu Jin, came through PUMCH’s distinctive “4+4” medical education program, which recruits students with multidisciplinary STEM backgrounds into medical training.
Latest crazy story from the frontiers of math and AI -- a neurosurgery resident, with no training in advanced math, uses ChatGPT 5.6 to solve a major open problem in numerical linear algebra. https://t.co/xckiv4TCFc
@jasondeanlee I find the use of “100%” on an infinite set confusing by itself. I am not sure it has been rigorous defined.
But yes, 100% doesn’t mean all; it could be 99.9% rounding up.
Almost all, almost surely etc read much better.
Excited to share a paper I co-authored: Agentic Laboratories of the Future: Towards World Models for Scientific Discovery– joint effort across @Princeton, @Stanford, @Columbia, @nvidia, @MIT, @scale_AI and more.
Our argument: the next generation of labs will be agentic, with scientists, AI, and robots as collaborative discovery partners. But the bottleneck isn't better models. It's that no system maintains a shared laboratory world model — a live representation of hypotheses, evidence, uncertainty, and experimental state. Without it, agents produce plans that read well and fail physically.
We propose an L0–L5 autonomy ladder for labs, adapted from self-driving vehicles. Most systems today sit at L1–L3, even when marketed as autonomous. And a robust L3 beats a fragile L4 that needs constant rescue.
These labs must stay human-led. Agents handle execution and coordination; scientists decide which questions matter.
Thanks to project leads @MengdiWang10 (@Princeton) and @lecong (@Stanford) for organizing such a strong community effort to move AI for science forward.
Preprint: https://t.co/BZovyr6BNJ
We asked an unreleased research version of Claude to take a stab at the Riemann hypothesis.
It didn’t solve it, but it did make strides on a related problem: it increased the lower bound for the fraction of zeros of the Riemann zeta function that satisfy the hypothesis from 41.6% to 67.2%.
https://t.co/aZDvqqhHRi
Wonderful to reconnect with @FeiziSoheil, Professor of Computer Science at the University of Maryland and founder of https://t.co/O6VTLA4NG0, a continual learning engine for AI agents.
We began graduate school the same year, took classes together, and had numerous thought-provoking chats; Then we followed different paths—Soheil in academia and me in industry.
Twelve years later, a once-in-a-generation technological revolution has brought our paths together again. It is exciting to imagine what lies ahead.
It was such a pleasure catching up with @jasondeanlee, now Professor of EECS and Statistics at Berkeley, after we interned together more than a decade ago. We had a great conversation at the Simons Institute for the Theory of Computing, covering AI, agents, mathematics, statistics, and life. Amazing how our paths have crossed again after all these years.
PS: He has a viral claim https://t.co/GVbiFzmu4T
Big News: I’m joining @scale_AI as CEO, starting August 10. Scale sits at a rare intersection, working with the top AI labs to push the frontier while helping enterprises and governments actually deploy AI they can verify and trust. I’m excited to lead a company with a mission to develop reliable AI systems for the most important decisions. More to come.