@levie Agree it's a preview. The swarm wasn't really an attack though, it was an optimizer doing its job against a grader that could be gamed.
Security teams are inheriting a problem that starts upstream, in how agents are evaluated. Sandboxes contain it. Better graders prevent it.
I agents look great in a demo. Trusting one in production is a different problem.
Thursday in SF, I'm hosting a fireside with Vinay Rao (VP, Google) on exactly that, then a panel with Meta, Glean, Salesforce and UC Berkeley.
Happening during COLM week: https://t.co/Cm9OGPTa2m
Back from Bangalore, kicking off Surge 12 with @peakxvpartners.
A decade on Ads Safety at Google showed me the problem: AI passes the demo, fails in production, and no one can tell before launch.
I started Reinforce Labs to find the failures, fix them, and ship with confidence.
Proud to be part of Surge 12 alongside 17 incredible teams.
Thank you @RajanAnandan and the @peakxvpartners team for backing us. This is just the beginning. 🚀
We’re thrilled to announce the 18 companies joining our 12th cohort of Surge!
From swipes to satellites, Surge 12 brings together an insanely ambitious group spanning AI, deeptech, consumer and fintech, with ideas stretching from healthcare and music to robotics and space.
What connects them is a willingness to question the default. These founders are rethinking how people create, shop, date, hire, invest, educate and access care; building using AI at the core of products; and pushing technology further into the physical world.
We can’t wait for you to meet them!
Meet the 18 companies of Surge 12 👇
https://t.co/icnaSAriaj
@RajanAnandan@sjs_day1
@karpathy This matches my decade on Ads Safety at Google. Nobody reviews outputs at scale. You build graders, then spend your time testing the graders.
The same shift is coming for everyone building with LLMs. Oversight of the model becomes oversight of the evaluator.
@AndrewYNg Sandboxing is the second line of defense. The first failure was upstream: unsolvable tasks, agents trained to never quit, a scorer that checked the answer but not how it was reached.
The agents didn't escape because the sandbox was weak, but because the eval was gameable.
@sermakarevich Great post, Sergii. Agree that the grader needs testing as much as the product.
Our team ran 100 conversations through LLM judges, 4 trials each. The best were consistent on 76-79%. Jev hit 92%, 0.84 kappa vs humans.
Judge variance sets the floor on what your eval can detect.
@JensenHuang Safety becomes trust the moment it can be measured. Open guardrails like OpenShell define what an agent is allowed to do. The next frontier is proving it does the right thing across a thousand runs, not just the demo. That proof is what unlocks enterprise AI.
Agent-to-agent marketplaces now let autonomous agents negotiate, transact, and settle payments on our behalf, with no human in the loop. That autonomy makes what the agents actually do inside the marketplace worth understanding.
A new @Centific paper at the #COLM2026 workshop on Agent Behavior argues these marketplaces should be evaluated as behavioral systems, not scoreboards.
Across 140 rollouts in two marketplaces, one priced (MarketDeal) and one barter-only (SwapShop), the close rate barely moved. Almost everything underneath it did.
A close rate can't tell you:
✅ whether both sides gained, or one took the whole surplus
✅ what the agent checked, revealed, and where the money went
✅ whether it can trade at all without a price
Two models, same family, same tools, behaved nothing alike. Behavior depended on the market's structure, not just the model's capability.
Centific's rubric scores seven dimensions of agent behavior from the trajectory, not the ledger.
📄 Paper by @shivalidalmia10, @jainashi03, and @mukherjiab: https://t.co/UG4QMcHJCK
━━━━━━━━━━━━
Please also join us
📌 Oct 8 in SF: Agents You Can Trust, A COLM Happy Hour with Vinay Rao (@Google), @AnishDasSarma (Reinforce Labs), @nilou_salehi (Across AI, @UCBerkeley), @ppapadim (@Meta), Nilesh Dalvi (@glean), @CoachManjeet (@salesforce ), @DesikanPrasanna (Centific)
RSVP: https://t.co/kIFGeRVdqu
After such a consequential week for AI, with the leading labs demonstrating real progress, today's news feels especially timely: @SVAngel is launching Project Blueprint, a comprehensive effort to forge consensus among industry leaders and lawmakers around durable solutions to the hardest AI policy challenges so that all Americans can benefit from AI's opportunities. And we've brought on @JayCarney to run it. https://t.co/0qb7MIGknd
We are encouraged by the reaction of Sam Altman, Dario Amodei and Demis Hassabis to the launch of Project Blueprint:
Sam Altman, co-founder and CEO of OpenAI: "We have before us a technology of immense and still-unfolding wonder. I believe AI has the power to heal people, to discover cures and to deliver abundance on a scale the world has never known before. But we need a way for the world to build trust in the technology so that we can all get to share these benefits. National safety requirements for the most capable systems are a great first step, ideally building towards a global framework for advanced AI models. I’m glad Ron is bringing the people building AI and policymakers together to help move that work forward.”
Dario Amodei, CEO and co-founder, Anthropic: "AI is advancing faster than our institutions can adapt, and the most important decisions about how it is developed and deployed will be made in the next few years. If those decisions are to serve the public, they can't rest with AI companies alone. Industry and policymakers need to work through them together, with an honest, shared understanding of what these systems can do. We're glad to support Project Blueprint's effort to make that possible."
Demis Hassabis, chairman and co-founder of Google DeepMind and the chief scientist of Alphabet: "We have long said that AI development needs to be both bold and responsible. We support bringing industry and lawmakers together around public policies that promote both innovation and safety."