1,720 agents had the right answer. by the end, 644 were left.
not one of them was argued out of it.
I ran one question through 46,000 agents for 2,140 rounds and logged why every single agent changed its mind:
a neighbour did - 27,330
new evidence - 0
re-read the source - 0
nobody was persuaded. they were surrounded.
here is the number that should bother you most. as the swarm got more wrong, it argued more, not less.
round 96: 1,755 agents arguing, 46.9% still holding the correct answer
round 2,140: 2,750 agents arguing, 2.6% still holding it
debate volume went up 57%. correctness collapsed by 94%. those two lines have nothing to do with each other, and that is the entire problem with measuring a system by how much deliberation it produces.
the flip wave only ever moved in one direction. across 2,140 rounds, not a single agent that abandoned the correct answer ever came back to it. a loop with no reversals is not reasoning. it is settling.
the holdouts did not fade either. they died in blocks. 1,670 agents in one round. 1,330 in another. the log has a name for it: cluster lost.
and underneath all of it, the default scoring rule almost every multi-agent framework ships with:
whichever answer more agents hold.
so if you are running agents that talk to each other, six things worth checking tonight:
→ keep a control group that never sees the other agents' output
→ require a flip to cite a source, not a peer
→ make flips reversible. if nothing ever flips back, you have no error correction
→ weight positions by evidence behind them, never by headcount holding them
→ pay one agent to hold the minority line and score it on being right, not on agreeing
→ alert on block flips. a whole colony changing its mind in one round is not thinking
more agents is not more evidence. it is the same evidence, counted again, louder.
the desk in the article below is built the other way on purpose. the watchers never see each other, and one agent exists only to say no.
consensus is not evidence.
everyone showing off their six agent company is showing off one login.
I counted what mine actually did in twenty minutes. 23,999 actions. Peak rate 38 per second. Six bots on the board.
One profile.
That is the whole thing. Six names in a sidebar, one browser session underneath, and every one of them inherits it.
The header on my own dashboard, which I built and then did not want to look at:
PROFILES 1. SHARED LOGIN yes. SPEND CAP none. AUDIT pending.
While I watched, the access ledger filled in on its own. Mail 624 actions. Drive 362. CRM 245. Ad manager 148. Billing 113. Crash reports 149. Eight systems reachable from one login, and I authorized exactly one of them, once, in an evening.
The blast radius counter went from 1 to 8 in the time it takes to make coffee.
Then the log line that ended it for me.
session: still open. no expiry.
I deleted a bot to test it. The session it was using did not die with it. Separate bots are not a security boundary. That is not my opinion, it is in the docs nobody opens.
So before you add a seventh agent, six things worth twenty minutes:
→ count your profiles, not your bots
→ list every system reachable from that one login. write the number down
→ set a spend cap before you set a schedule
→ delete a test bot, then check whether its session actually died
→ find your peak actions per second, then ask what twenty minutes at that rate touches
→ turn audit on before you need it, not after
None of this makes your fleet slower. It makes it something you can still switch off.
Twenty four thousand actions is not a flex. It is a number you should be able to explain.
I shut down my research desk on a Tuesday.
Six AI agents worked its first night shift. Payroll for the night: $4.72.
By 06:00 they had read 912 filings, 68 earnings calls, 171 pages of SEC text and covered 97 tickers. The humans it replaced cost $294,000 a year to do less.
I did not write a strategy. I wrote six job descriptions, one rule each.
→ SCOUT never interprets. It only brings the document, timestamped, the second it hits the wire.
→ SNIPER never summarizes. Every output must contain an entry price or it is thrown away.
→ BUBBLE never counts mentions. It counts money moving behind the story.
→ PULSE never predicts. It only answers one question: does the market already know?
→ RISK never explains itself. It holds a hard cap and overrules the other five without a vote.
→ EXIT never lets an entry exist before the way out is written down.
One chief of staff sits above them. It never trades. Its only job is deciding which of the six is allowed to wake me up.
Here is the part I did not expect.
The header on the desk reads NO CHATTER = RISK. When the agents go quiet, the system flags it as a failure. A human desk goes silent when it is scared and everyone calls that focus.
Morning brief at 04:47. Twenty five flags, ranked, each carrying the filing it came from.
The models are public. The filings are public. The APIs have been public for years.
Almost nobody has bothered to give six of them job titles and a night shift.
The data is free. The org chart is the entire edge.