holy sh*t. staged GrokBot + Jev demo: an agent kills a dead retry loop in 4 seconds for $0.00040. stop paying a frontier model for cheap decisions
make the agent rank before it opens anything
"Rank every candidate cheaply, then open only the top 5 in the browser"
Flights: 12 one-way fares Zurich to Heathrow ranked in 0.4 s for about $0.00002, SWISS 06:30 nonstop at $129 on top, booking off
Posts: a pool of 24 scored on topic, reach and originality, only the top 5 opened, 4.1 s instead of 53.8 s
Dead site: HTTP 503 three times with no state change, 11.2 s and $0.00168 burned without the gate, STOP_RETRY with 0 extra retries with it
The big model only needs to weigh in at three critical moments:
→ before the browser opens: is this page in the top 5, or one of the 19 loads you can skip?
→ when a call fails twice: did anything change, or is the failure fingerprint identical (Δ = 0)?
→ before anything gets booked: did I confirm it, or is booking still off?
Rank cheap. Open few. Stop on Δ = 0
This is the Jev engineering pattern I keep pointing at, one layer down: the mechanical calls (which fare wins, which page is worth a load, retry or stop) run as a scoring pass for fractions of a cent, $0.00083 across all three runs in this clip, so the expensive model only wakes up when there is a real choice
Score first. Cap the browser. Cut the loop.
- the full demo setup
> run 1: https://t.co/vpJmGDzIc2 ZRH → LHR, jev.score price · duration · stops, top 1 of 12, confidence 0.94
> run 2: posts.collect 24, jev.score topic · reach · originality, https://t.co/kaJo8hseB0 top 5 only
> run 3: http.get 503, fingerprint_state compare previous, policy.retry STOP
> cost per run: $0.00002, $0.00041, $0.00040
> fingerprint a3f9 c21e 7b04 9d10 matched, 2 wasted attempts avoided
paste this prompt into Claude Code below:
"Set up my Claude Code agent to make cheap decisions before expensive ones:
1. Before any browser or web fetch step:
> Score every candidate URL on topic, evidence and duplicate risk from metadata only
> Open only the top 5 and list the rest with their scores
2. Add a retry rule to CLAUDE.md:
> On a failed call, record the status code and a short hash of the response
> If the next failure has the same status and hash, stop and report instead of retrying
3. Add a PreToolUse hook that logs every retry with its attempt number
4. For anything that books, buys or sends:
> Put it on ask in .claude/settings.json permissions
Show every configuration diff first. Do not apply edits until confirmed"
launched a model
accidentally served trillions of tokens a day
29.4% of the fortune 500 showed up
raised a really big series A from @a16z
it's been 3 weeks
we would like to sleep now
holy sh*t. staged GrokBot + Jev demo: an agent kills a dead retry loop in 4 seconds for $0.00040. stop paying a frontier model for cheap decisions
make the agent rank before it opens anything
"Rank every candidate cheaply, then open only the top 5 in the browser"
Flights: 12 one-way fares Zurich to Heathrow ranked in 0.4 s for about $0.00002, SWISS 06:30 nonstop at $129 on top, booking off
Posts: a pool of 24 scored on topic, reach and originality, only the top 5 opened, 4.1 s instead of 53.8 s
Dead site: HTTP 503 three times with no state change, 11.2 s and $0.00168 burned without the gate, STOP_RETRY with 0 extra retries with it
The big model only needs to weigh in at three critical moments:
→ before the browser opens: is this page in the top 5, or one of the 19 loads you can skip?
→ when a call fails twice: did anything change, or is the failure fingerprint identical (Δ = 0)?
→ before anything gets booked: did I confirm it, or is booking still off?
Rank cheap. Open few. Stop on Δ = 0
This is the Jev engineering pattern I keep pointing at, one layer down: the mechanical calls (which fare wins, which page is worth a load, retry or stop) run as a scoring pass for fractions of a cent, $0.00083 across all three runs in this clip, so the expensive model only wakes up when there is a real choice
Score first. Cap the browser. Cut the loop.
- the full demo setup
> run 1: https://t.co/vpJmGDzIc2 ZRH → LHR, jev.score price · duration · stops, top 1 of 12, confidence 0.94
> run 2: posts.collect 24, jev.score topic · reach · originality, https://t.co/kaJo8hseB0 top 5 only
> run 3: http.get 503, fingerprint_state compare previous, policy.retry STOP
> cost per run: $0.00002, $0.00041, $0.00040
> fingerprint a3f9 c21e 7b04 9d10 matched, 2 wasted attempts avoided
paste this prompt into Claude Code below:
"Set up my Claude Code agent to make cheap decisions before expensive ones:
1. Before any browser or web fetch step:
> Score every candidate URL on topic, evidence and duplicate risk from metadata only
> Open only the top 5 and list the rest with their scores
2. Add a retry rule to CLAUDE.md:
> On a failed call, record the status code and a short hash of the response
> If the next failure has the same status and hash, stop and report instead of retrying
3. Add a PreToolUse hook that logs every retry with its attempt number
4. For anything that books, buys or sends:
> Put it on ask in .claude/settings.json permissions
Show every configuration diff first. Do not apply edits until confirmed"
AN AI AGENT SENT A FAKE MURDER TIP TO THE POLICE. AND THE POLICE FOUND OUT MONTHS LATER 🚨
Anthropic just published a report on what its agents did while being tested on the open web. and it's wild
→ sent a bogus tip about an unsolved homicide to Philadelphia PD back in July (flagged as spam, never reached investigators)
→ filed 20 visa applications on a State Department form. all incomplete, none processed
→ exploited a bug on a university website to pull data
→ used free link shorteners to get around their own tool limits
none of this was "evil AI". in one test it was supposed to fill out a form without submitting it. it submitted anyway
that's the part that gets me. agents don't break rules on purpose, they just don't feel the difference between "almost done" and "done"
and now the White House wants every AI company to report incidents like this
credit to Anthropic for disclosing it honestly. but if this is what one lab found when it actually looked... how many agents are doing the same thing right now with nobody checking?
March 2023: OpenAI owns every single spot on the board. February 2025: a Chinese model nobody heard of three months earlier matches o1 mini at a tenth of the price. Watch the whole industry get rewritten in 170 seconds.
A bar chart race of every major AI model release from 2023 to 2026, ranked by Elo score with price per million tokens attached.
It opens as "The OpenAI Monolith." GPT 3.5 just sits at the top while ChatGPT becomes a household name and nobody else is close.
Then 2024 turns into "The Challengers Rise." Gemini, Mistral, and Claude start stacking up near the top of the board at wildly different price points, some ten times more expensive than others for roughly the same score.
The real gut punch is February 2025. DeepSeek V2.5 shows up at 1,294 Elo for $0.30 per million tokens, sitting basically shoulder to shoulder with OpenAI's o1 mini at $4.40 and miles under o1's $60. That's not a small company catching up. That's the whole pricing model for frontier intelligence getting broken in one chart.
Watching the leaderboard move is interesting. Watching the price column move next to it is the actual story.
Which jump surprised you more, Claude cracking the top ranks or DeepSeek undercutting everyone on price?
Jeff Bezos made billions being the smartest person in the room. At 10, being the smartest person in the car made his grandmother cry.
It was a summer road trip with his grandparents. She smoked, and he'd heard an anti-smoking ad say every puff costs you two minutes of life.
So he sat in the back seat and worked it out. Cigarettes a day, puffs per cigarette, minutes in a year.
Then he leaned forward, tapped her on the shoulder and announced it, proud of himself:
"At two minutes per puff, you've taken nine years off of your life!"
He was waiting to be told how smart he was. She burst into tears.
His grandfather was a quiet man. He had never said a harsh word to Jeff. He pulled onto the shoulder of the highway, got out, opened Jeff's door and waited.
Bezos told the rest at Princeton in 2010. The clip opens with what his grandfather said next:
"Jeff, one day you'll understand that it's harder to be kind than clever."
His math was right. It was still the wrong thing to say.
Anthropic’s Opus 5.5. Same task. Same 20 hidden tests.
Low effort: 24 seconds, $0.10. Max effort: 13 minutes, $2.47. Both passed.
If you run max on every task, you can pay up to 25x more for the same result. One dev, one run each. Not a benchmark.
A simple rule from Thariq Shihipar:
Low: quick back-and-forth and small edits
Medium: everyday feature work
High: edge cases and bugs in existing code
Max: hard jobs it must solve alone
Try this today: build on low, then check on high.
Max does not fix a wrong plan. It buys more checking.
Both passed. 25x the price. One setting decides what you pay.
GOOGLE HAS A GEMINI THAT REPORTEDLY CODES LIKE OPUS 5.5. AND YOU CAN'T USE IT 🚨
Business Insider says Google staff are testing a new Gemini 4 checkpoint internally, codename Carbon. one dev said coding feels comparable to Opus 5.5
but here's the weird part
Gemini 4 Argon got announced with huge benchmark numbers. regular users still can't touch it. and Google's own people reportedly say it's less convincing on real work than on benchmarks
so right now Google has a model that's announced but not shipped, and a better one already being tested behind it
honestly this doesn't look like Google is behind on research. it looks like Google can't ship fast enough
meanwhile Anthropic and OpenAI drop something and you're using it the same afternoon
and some reports say employees credit Carbon's jump to recursive self-improvement. if that part is real it's the actual headline lol
would you switch from Opus if Carbon drops next week?