A 19-YEAR-OLD PUT GROK, OPUS 5.5 AND DOTS ON ONE TRADING DESK. 487 BAD TRADES DIED BEFORE TOUCHING A REAL DOLLAR
he built it in his dorm over one weekend, and every trade has to get past all four agents before it goes live
01 / scan
grokbot reads 14,280 markets at the same time and throws out 32 trade ideas a second. most of them should never ship
02 / plan
opus 5.5 stress-tests every single one, sizes the bet and caps drawdown at 1.2% before a cent moves
03 / fill
jev fires the trade in a sealed sandbox first. if it breaks there, it dies there, and the real book never sees it
04 / settle
dots signs off last. only trades that already proved themselves go live
after one run:
> win rate up 20.3%
> 487 bad trades blocked
> 0 leaks
> 4,456 ideas through one desk
grok does 30% of the work, opus 28%, jev 22%, dots 20%
the full setup for hiring all four is in my article below:
this is pure f*cking treasure
a research team at the Harvard Institute put Jev, Dots, Grokbot and Opus 5.5 on one search tree, and it kills the retry tax
most teams still spend test-time compute on 32 blind resamples of one model. the same wrong answers come back in the same clusters, and you pay for every one of them
their paper, "Refining Over Resampling", gives each model one job:
> breadth: Grokbot samples 32 candidate paths at T=0.8 and pushes step entropy to 0.92
> critic: Opus 5.5 walks every step at T=0.0 and flags the exact one that breaks
> local repair: Jev rewrites only the broken step at T=0.2 and keeps the rest of the chain
> consensus: Dots votes across all 32 paths and locks the answer at 27/32 with 0 reward models
the numbers against plain resampling:
> MATH500 hits 64.4%, up 27.2 points
> AMC climbs 7.5 points to 42.2%
> resampling alone gains 6.2% on the same token budget
> 604 paths explored, 0 PRMs
Grokbot explores. Opus finds the flaw. Jev fixes it. Dots signs off
you stop paying for the same mistake 32 times
the full setup for hiring all four is in my article below:
OpenAI co-founder Andrej Karpathy built a way to make GPT, Claude and Grok work as one team. I turned it into a prompt that builds the whole app for you
It's f*cking unreal...
Send it to Claude Code with Opus 5.5 and thank me later. Then read the full guide below to see what hiring all of them actually costs
I PUT ANTHROPIC, OPENAI AND ELON'S GROK ON ONE PAYROLL. THEIR BOSS IS A SINGLE AGENT THAT RUNS MY MONEY, HEALTH AND BUSINESS AND CLOSES 90% OF THE WORK
three AI labs that fight each other every day now take orders from the same agent
it's a concept build, and it already juggles 12,480 boards while the counter in the corner climbs from 18 to 79
-> 18 processes
money wakes up first. $24,860 this month, up 12.4%
rent 38%, food 21%, subs 6%, and 31% goes to savings
> income target is $10,000+ a month and tracked. invoices and taxes wait in the queue
-> 29 processes
health joins. 7h 40m average sleep, in bed by 23:00
labs and checkups on reminders
> food and steps get their own boards
-> 38 processes
business. it scans 214 emails and scores every one
the weekly report goes out monday at 9:00
> crm, mail and a 5-agent crew share one desk. then the agent links business to health
-> 45 processes
the payroll shows up. opus 5.5, openai's dots and grok haul data into the core
80% of the work is already closed
> every lab feeds the same core. none of them sees the whole board
-> 54 processes
jev takes the wheel. it flags a task as urgent with 96% confidence in 81ms
if jev is sure, the step runs in code. if it isn't, opus 5.5 gets it
> route and audit run right behind every decision
-> 79 processes
1 web, 11 agent networks, 12,480 boards, and 90% of the work closed
it closed 90% of my life in one night and left me a to-do list with 3 items. i haven't opened it οΏ½οΏ½
CLAUDE BUILT A MOTION DESIGN SPIDER THAT WALKS ACROSS 7 STYLES IN 24 SECONDS - 720 FRAMES, 8 LEGS, 0 KEYFRAMES SET BY HAND
a reel like this is 7 clips and a week of editing - here it is 1 scene and 1 camera
7 styles β 1 floor β 1 spider β 720 frames β 1 web β back to the floor
nobody moved a single leg in this clip
the first 5 minutes are yours: 7 styles to show, 1 guide to carry them, and 24 seconds
Claude lays them out as tiles on one floor - kinetic type, voxel wave, spring physics and 4 more
each leg is a rule, not a drawing - a foot stays put until the body gets too far
the camera is 1 path with 7 stops, about 2.7 seconds a stop, and a label lands with the feet
a browser films 30 frames a second, and ffmpeg glues 720 shots into 1 MP4 at 1080 Γ 1350
but 7 clean clips in a row is a folder, not a reel - people leave at clip 2
so the spider is the cut - 1 body crosses all 7 tiles, then spins a web in the last 5 seconds
every reel needs a guide before it needs more clips - remove the spider and it is 7 loose tabs
it shows 7 things Claude can draw in code - it does not show 7 finished films
check the article below, it is the 10-step studio behind this reel β
NVIDIA and OpenAI researchers just exposed where the next $100B AI infrastructure race is heading, AI agents that never stop working
"the chatbot era sold intelligence, the next AI boom could sell machines that work 24/7, with companies paying $1,200 to $6,000 a year for an AI worker that never clocks out"
1 β $100/month gets you a persistent Dot, roughly $1,200/year for an AI agent that keeps working while you sleep
2 β Dots can monitor Gmail, Slack, GitHub and 4,000+ apps, maintain context, and hand heavy work to Codex and ChatGPT Work
3 β OpenAI tested 16,600 prompt injection emails and reported 0 successful compromises, while explicit permission removals worked 17/17 times
4 β The catch, long running agents still slipped 8.6% of the time after 5 chained tasks and 19.7% after 10
5 β The real opportunity is not another chatbot, it is turning $1,200 to $6,000 a year into an AI worker that runs for days, weeks, and eventually 24/7
follow & bookmark this before everyone starts talking about autonomous agents
i mutated a dumb web crawler into a spider with Jev, and now it builds one giant web out of the most popular AI posts
the old crawler read everything and remembered nothing. Jev gave it one question per post: keep or drop, answered in 81ms
every post it keeps becomes a strand, tied to the posts it agrees or argues with
ask the web anything and it answers in ONE reply, with every claim pinned to the post it came from
full build, prompts and CODEBASE in the article below:
SAM ALTMAN HAS X. ANTHROPIC HAS RELEASE NOTES. RESEARCHERS HAVE PAPERS.
MY SPIDER CRAWLS ALL OF THEM AND FINDS THE ONE STORY THEYβRE ALL ACCIDENTALLY TELLING.
I fed it 128 AI sources.
It came back with a web.
Not a summary.
A map of what is actually connected.
One branch finds a benchmark:
Vals Index #1 β 68.9%
another hits an ecosystem:
4,000+ AI apps
another catches a research signal:
Claude β 26% of R&D
then it starts connecting the dots.
model release β benchmark jump
paper β product launch
research result β new agent
company announcement β second-order effect
Thatβs when the feed stops looking random.
OpenAI ships something.
Anthropic answers.
A benchmark moves.
Three startups suddenly build the same feature.
Most people read those as 4 different stories.
The spider asks:
what changed upstream that made all 4 happen?
128 sources go in.
One connected story comes out.
The next edge in AI probably isnβt getting news faster.
Itβs understanding the pattern before everyone else realizes there is one.
ANTHROPIC, OPENAI AND AN 8 GB SPIDER WALK INTO MY MACBOOK. THE SPIDER WOVE 2,418 FILES INTO ONE WEB
it's one laptop, and it still says more about how builders really use AI than any lab report
the spider has 8 legs, 8 eyes and runs at temperature 0.2. here's what it found:
folder: agents
jev, the decision vault
1,204 decisions today. 81ms each. 0 regrets
invoice β code. urgent? yes. delete dupes? yes
reply to ana? β human
it decides everything except the one thing that's personal
folder: code
claude_archive. est. 2025. do not delete. ever
1,882 sessions
12,406 drafts
31 hard cases
940 shipped
2,113 never shipped
48 love letters to tests
oldest file: the first hello world
newest: 3 seconds ago. status: still writing
folder: archive/vault
the dots crew
find β read β judge β size β place
the one thing he actually searched for. 1 match, 0.4 seconds
2,418 files. 5 dots. 1 web
and right in the middle of it, the three things that do the work
the labs show you benchmarks. a laptop shows you the 2,113 drafts nobody posts
if a spider wove your laptop, what would end up in the middle? β
@leopardracer step 7 is the one nobody does. everyone tests their agent with a perfect prompt and then acts surprised when real users break it in 5 minutes