a year ago, coding agents triggered under 3% of deployments on vercel. by july it was more than half of about 6 million a day, per guillermo rauch.
the information reports vercel is now at $600M annualized revenue, up 148%, with agents behind roughly half of new business.
so on one of the biggest hosts on the web, the typical deploy now comes from an agent. the open question isn't whether agents can deploy. it's who read the diff, and whether the agent can explain why it made each change.
same model, same benchmark, 3.8% vs 34.6%. the only difference was the harness.
the Polar paper (arXiv 2605.24220) ran Qwen3.5-4B on SWE-bench Verified inside four coding agent harnesses before training: 3.8% in Codex, 29.8% in Claude Code, 34.2% in Pi, 34.6% in Qwen Code.
after RL inside each harness, the Codex score jumped 22.6 points. Qwen Code moved 0.6.
when you compare agents, you're comparing model + harness. a leaderboard number without the harness next to it doesn't tell you much.
the most interesting part of claude code mods isn't the mods. it's the one anthropic shipped with them.
2.1.287 added mods (plugins that can change deeper behavior) plus a built-in one called "you should know": a side agent that watches the session and flags things you or claude might miss.
so the launch example is one agent keeping an eye on another.
that's where this is going. agents that can see what other agents are doing catch problems before they become commits. same idea behind collide, applied to who's editing what.
claude code quietly changed what "too many agents" means.
2.1.217 added a cap of 20 subagents running at once (CLAUDE_CODE_MAX_CONCURRENT_SUBAGENTS to override). 2.1.224 then removed the 200-per-session spawn cap entirely.
so the limit is no longer how many agents you start, it's how many touch the code at the same time.
that's the right knob. 20 parallel agents is fine until two of them edit the same file. the cap counts agents, not collisions.
every agent can be right and the team can still be wrong.
a new arxiv paper (2610.02036) tested it. when the key event was visible, agents revised correctly 40/40 times. when it was hidden, 12 to 17 out of 40, about chance. restoring one authoritative fact brought it back to 40/40.
the author's point: more reasoning, more roles or more messages can't recover information the agent never sees.
coding agents hit the same wall. an agent that can't see what another agent just changed will make a locally correct edit that breaks the whole. that's the gap collide closes.
the boring part of agent infra is getting the money now.
restate just raised a $20M series a (led by singular) for durable execution: multistep workflows that survive crashes and network drops. replit is a customer. temporal, the incumbent, raised $550M at a $12.55B valuation earlier the same month.
the reason is simple. a 40-step agent run that dies on step 37 and loses its state costs you the whole run.
co-founder stephan ewen put it well: "you need to make sure you track exactly what you do."
agents need a record of what already happened, not just a smarter model.
six senior engineers at amazon shipped a project scoped for 30 people over 12 to 18 months in 76 days.
aws's vp of agentic ai shared the numbers: commits went from 2 a week to 40. across 25 amazon stores teams the median gain was 4.5x.
what they changed is the interesting part:
specs before code
steering files and standards for agents
several agents running in parallel, overnight
agents running integration tests themselves
human review focused on architecture, not style
more agents per engineer means more agents in the same code. that's the problem collide works on.
running claude code, codex and grok build on one repo is now a desktop app.
offrun (free, apple silicon) gives every agent its own git worktree, has a second agent review the uncommitted diff before anything merges, and moves a chat to another login when one account hits its limit.
worktrees are the right default. they keep agents out of each other's files while they work.
the open question is merge time: two agents in separate worktrees can still change the same logic, and you only find out when the branches meet.
a coding agent spent a week rewriting an inference server and beat tuned vllm.
baseten ran claude code on qwen 3.6 35b on one b200:
1,792 vs 943 tokens/sec single-stream decode
10,307 vs 6,030 tokens/sec at 32 concurrent requests
12ms vs 28ms time to first token
the bill: ~1.7B tokens and ~200 b200 hours. humans spent hours, mostly steering it off fixating on one metric and approving anything that changed accuracy.
not production, but it's a real data point on what a week of agent time buys.
what do people actually put in CLAUDE.md?
researchers at kasetsart and naist read 253 of them from 242 github repos:
build and run commands: 77.1%
implementation details: 71.9%
architecture: 64.8%
testing: 60.5%
security: 8.7%
so we tell agents how to build the project and almost never what's off limits.
and none of it can say what another agent is editing right now. a manifest is static. that live layer is what we're building at collide.
your coding agent might own a public github repo you don't know about.
security firm glow found 13,000+ internal screenshots from 300+ orgs sitting in public repos, mostly under developers' personal accounts. billing records, unreleased features, treasury dashboards.
the cause was mundane: until sept 1, github's cli couldn't attach images to a pr. asked to show proof of a ui fix, agents spun up public repos to host the screenshots.
the agents finished the task. nobody had scoped where they could publish.
worth a search today: "gitshot-images" repos and "_gitshot" release tags.
most multi-agent failures aren't the model being dumb.
uc berkeley researchers annotated 1,600+ traces across 7 multi-agent frameworks and found 14 failure modes in 3 buckets:
- system design issues
- inter-agent misalignment, like agents ignoring each other's input
- task verification
in their chatdev case study, fixes helped but didn't solve it: +9.4% task success from tighter roles, +15.6% from adding verification.
more agents means more coordination, not just more output.
arxiv 2503.13657
@Cozy2054934@kevinvera@aiwithadb agreed, narrow context and structured handoffs are most of it. the token drop comes from agents not redoing work another agent already did.
@gadi_neelesh "stay in your lane" works until two lanes need the same file. that's the point where agents need to see each other's edits, not just their own instructions.
@WenYu98767859 amazing, can't wait to try it. collide already knows when two agents are about to hit the same file, so if you want that signal for the face, we'd love to help wire it in.
@uday_sail makes sense, and keeping specs and findings outside any one session is the right call. the part we keep hitting is two containers editing the same module without knowing, then the human merge eats it. that's what we're building collide for, would love to compare notes.
@WenYu98767859 amazing, can't wait to try it. collide already knows when two agents are about to hit the same file, so if you want that signal for the face, we'd love to help wire it in