MCP or CLI: which surface do agents actually succeed on?
We ran the experiment. Stripe, 6 identical tasks, both surfaces, Claude Code and Cursor, 4 runs per cell, fresh sandbox every run, every message recorded on the wire.
MCP or CLI: which surface do agents actually succeed on?
We ran the experiment. Stripe, 6 identical tasks, both surfaces, Claude Code and Cursor, 4 runs per cell, fresh sandbox every run, every message recorded on the wire.
Scope: these 6 tasks, these versions (in the image footer), 4 reps per cell, Wilson 95% CIs, test mode. Cursor cannot be tool-restricted, so its MCP cells measure whether it chooses and succeeds on the surface. Full methodology: https://t.co/gu8LQpI31y
5/6