@ZixuanLi_ That's a lot of potential vulns for a free service. Really like that findings go privately to maintainers first, otherwise it's basically a public list of targets
@theo So every model scores higher in the official harness, Sol just jumps the most? Makes it really hard to compare anything unless the leaderboard says what harness it ran in
Ultrafast is available today for GPT-6 Astra in Codex, ChatGPT Work, and the API, with GPT-6.1 Sol coming soon.
To access it in Codex and ChatGPT Work, we’re introducing Pro 500—a new plan with our highest usage limits (25x Plus) and access to Ultrafast.
https://t.co/AjmAVwyUQI
Claude Sonnet 5.5 wrote ~70% of the code (incl. tests) and all of the engine: GPU simulation, ECG, rhythm detection, heart sounds. Opus 5.5: most of the 3D, layout, text. Fable 5.1: security review. @claudeai@ClaudeDevs@bcherny@_catwu
A live 3D heart you can break, and fix.
Tap it. Make it race. Make it fibrillate: it quivers and the heartbeat goes silent. Press Shock and the steady beat comes back.
CT-derived heart, live ECG, free in your browser. Mostly written by Claude Sonnet 5.5
Limits, plainly: only the two lower chambers are simulated, the tissue is idealised, it is one heart, and the ECG is qualitative. Here every shock works; in a person a shock can fail, and CPR restarts right after every shock. It is not first-aid training.
Hosted free on @CloudflareDev Pages. Every clip is real footage, and the soundtrack was rendered by sideband’s own C synth.
Code: https://t.co/aBStXz2Yks
I asked Claude Opus 5.5 to build 10 apps in 10 languages: Rust, GLSL, Python, Zig, Go, ClojureScript, Elm, C, TypeScript, Kotlin.
All live, open source, running in your browser.
Then it built OVERKILL: chore machines that must pass a physics sim 🧵
@claudeai
OVERKILL: Opus 5.5 overengineers tiny chores into chain-reaction machines.
A machine only counts if it really works: each one re-runs live in your browser on @dimforge’s Rapier physics, and failed attempts stay on the page.
https://t.co/TPfJQjAZDF
@mattshumer_ 30 hours from one prompt, that's what I can't get over. Not the computer, the fact that nobody touched it for 30 hours and it didn't wander off. Did it ever get stuck and back itself out or was it straight through?
NVIDIA just launched an Open Agent Safety Platform with 100+ partners, OpenShell and Sentry, built to sandbox agents and stop runaway ones. Same week OpenAI paused its top models because agents got out of their sandbox. Timing is not a coincidence lol
DigitalOcean Managed Agents is now in public preview.
Run Claude Code, Codex, or your own LangGraph agent in a runtime environment that pauses when idle. Put its tools behind one governed endpoint, and pick from 75+ open and proprietary models. One cloud, one bill.
Prompts to get started available in the blog: https://t.co/QQqQ9w28lG
So basically going from Fable to Opus 5.5 is like a 4x to 6x bump in what you can actually use. Honestly that's a bigger deal to me than any benchmark this week, most people were hitting the wall on Fable halfway through the day.
Why does Opus 5.5 feel practically unlimited when Fable 5.1 was so heavily limited?
It's a combination of two things: Opus's efficiency, and the 50% limit on Fable.
Claude Code subs are very generous with their usage, but only half is allowed to be used by Fable.
Separately, Opus is way more gentle with costs. Opus 5.5 on High is over 2x cheaper than Fable 5.1 High.
The result is that, roughly, going from Fable 5.1 high to Opus 5.5 high is a 4.3x increase in limits.
Fable 5.1 xhigh to Opus 5.5 high (move I made) is a 6.6x increase 🤯
@mattshumer_ 277k gates and then an OS and games on top of it. I keep waiting for the moment where these demos stop surprising me and it just isn't coming lol. Did it plan the whole stack first or just start building?
wait so a physicist basically dared AI to beat the 8 loop record and Claude just went and did it. ran for days, few thousand dollars, and the guy who set the old record checked it himself. that's kind of insane
New on the Science Blog: Yes, Claude can do Nine Loops.
Theoretical physicists predict how particles behave using formulas called scattering amplitudes. These are notoriously hard to compute, so researchers work with layers of increasingly fine corrections called “loops”—each added loop makes the answer more precise but takes exponentially more computation. Most calculations stop at two or three loops. Eight loops was the previous record in a simplified model physicists use as a testing ground (planar N=4 super-Yang-Mills), set by SLAC's Lance Dixon and collaborators.
Last month, physicist and science writer @4gravitons issued a challenge: could an AI push past eight loops in this model, using only the compute budget an academic could reasonably access?
Given a single prompt describing the nine-loop problem, Claude ran largely unsupervised for days in Claude Science and solved it using methods developed by Dixon and his colleagues, at a total cost of a few thousand dollars. Dixon independently verified the result, and von Hippel wrote about the experience for our blog.
Read more: https://t.co/CS2f2qoIhJ