IF YOU SEND 65 BILLION DATA POINTS TO CLAUDE, YOU BUILT THE PIPELINE WRONG
The viral story says a Stanford AI group is using JEV + Claude Code to process 65B+ points every nine minutes.
Claude is not reading 65 billion things.
That is the entire trick.
JEV stands at the door. Every item hits a fast typed decision before it can consume frontier-model context:
> boring → drop it
> obvious → handle it in code
> uncertain → hold it
> weird and ambiguous → send it upstairs
The pipeline is not:
65B items → Claude → enormous bill
It is:
65B items → JEV → deterministic policy → tiny uncertain queue → Claude Code
JEV does the cheap, repetitive classification. Code applies the thresholds and handles the obvious cases. Claude receives only the small pile that actually deserves investigation.
That lets the expensive model spend its context on comparison, explanation, hypothesis building and the hard edge cases instead of repeatedly deciding that another ordinary record is ordinary.
The architecture has three clean responsibilities:
> JEV filters.
> Claude investigates.
> code executes.
One important caveat: I could not find a Stanford lab page or paper confirming the headline throughput, and another copy of the story claims 100B+ points every 15 minutes. Treat the scale as a reported claim.
The design pattern is still the part worth stealing.
Stop paying your smartest model to do the dumbest work.
Jev came out and people immediately put it between humans and consequences
a political debate.
a Claude Code command.
a live trade.
a spoken English test.
and Jev itself.
then they gave it one job:
look at the state.
pick the next answer.
never write an essay.
here's what happened:
> jevmeter
> Jev watches a debate one sentence at a time and asks: factual claim? evasive? emotional appeal? dodged question?
> 1,191 calls. 5,955 yes/no answers.
> total cost: $0.0497.
> then the software cuts the footage and renders the live meters on screen.
> https://t.co/1a90Vh0Eht
> jev-bouncer
> a second gate in front of Claude Code commands, file edits and MCP calls.
> Jev returns allow, defer or deny before the action runs.
> deterministic tripwires still handle the hard rules. every verdict goes into a local audit log.
> https://t.co/UyM8wKrcXq
> QuantDinger
> a 12K-star open-source trading OS.
> before an entry reaches the exchange, Jev reads the order, strategy context, positions, exposure and budget.
> pass or reject. exits, stop-losses and emergency actions bypass the AI gate.
> https://t.co/y7BcVwWy7A
> dasheng
> read English out loud. streaming speech recognition follows each word while Jev checks whether the recognized text still matches the original.
> mismatches are marked while you are speaking.
> Jev judges the text layer; timing and completeness stay in local code.
> https://t.co/KL7UcQLuv0
> NanoJev
> someone rebuilt the decision model as a 0.6B open implementation.
> on the held-out ViZDoom Basic test: NanoJev scored 128/128. the original Jev scored 56/128.
> the copy beat the teacher on that benchmark.
> https://t.co/BHu7nmiRrt
and here's the pattern:
none of these systems need Jev to generate paragraphs.
they need it to inspect a small state and choose:
FACT / DODGE / ATTACK
ALLOW / DEFER / DENY
PASS / REJECT
RIGHT / WRONG
LEFT / RIGHT / SHOOT
that's the entire decision layer.
the large model plans.
Jev picks.
code owns the consequences.
Jev came out and people immediately connected it to everything except a chatbot
Minecraft. Mario. a Tesla-style driving sim. a real Android phone. a trading account. Claude Code.
then basically told it:
“you don't need to explain anything.
just pick the next move.”
here's what happened:
> minecraft-agent
> Jev beat Minecraft.
> a larger model plans the route. Jev picks every move: mine, craft or fight.
> 8:43 run. 131 decisions. the dragon died to exploding beds.
> https://t.co/8dKuGNmxJm
> jevpilot
> a Tesla-style autopilot inside a browser driving sim.
> Jev chooses steering and speed up to 4 times a second.
> https://t.co/deB6SMCOyN
> typesafe-mario
> Jev plays the original Super Mario.
> it receives Mario's position, enemies and terrain as structured data, then presses one of 7 buttons.
> no screenshots. no vision.
> https://t.co/cDYCMOkh8t
> mobile-jev
> Jev controls a real Android phone.
> it opened Uber, entered the route and reached payment.
> around 9 actions in 21 seconds.
> https://t.co/EgPNH0qusf
> jev-voice-browser
> you talk. Jev clicks.
> it often starts acting before the sentence is finished.
> around 300ms per word. the entire demo cost roughly one cent.
> https://t.co/yxX2np7TVW
> jev-trade
> 5 coins on Hyperliquid.
> every second Jev chooses long or short, then open, close or hold.
> https://t.co/wa66nY03cV
> youtube-sponsor-detection
> “this video is sponsored by…” → skipped.
> under one cent per hour of video.
> https://t.co/ojfIonKJer
> unclutter
> Jev looks at a page and hides ads, cookie popups and newsletter boxes.
> it remembers the site, so the next visit needs zero API calls.
> https://t.co/gawogc6eEx
> jev-pruner
> a 76,000-character log becomes 4,000 characters before Claude reads it.
> errors stay. noise disappears.
> https://t.co/Wn8XPUmbsg
> fast-jev-compaction
> instead of rewriting old Claude Code context into a summary, Jev decides what stays and what gets dropped.
> nothing is rewritten.
> https://t.co/bUJAVfQzik
and here's the pattern:
none of these projects need Jev to write an essay.
they need it to look at a tiny state and choose:
MINE / CRAFT / FIGHT
LEFT / RIGHT / JUMP
TAP / TYPE / SCROLL
LONG / SHORT / HOLD
AD / NOT AD
KEEP / DROP
the large model plans.
Jev picks.
code does the rest.
that's the real shift: intelligence stays expensive at the planning layer, while the millions of small decisions underneath it become fast and cheap.
this is the kind of architecture that makes current AI agents look primitive
a Stanford team is sorting 75 billion data points every 11 minutes by pairing JEV with Claude Code
they don't dump the entire firehose into Claude and ask a frontier model to think about every single item
JEV reads everything first and handles the cheap decisions: relevant or noise, normal or suspicious, resolved or worth escalating
the obvious cases are handled immediately. only the ambiguous edge cases reach Claude, where the expensive reasoning is actually useful
JEV reacts. Claude reasons.
that simple split changes the economics of the whole system
less context sent to Claude, fewer expensive calls, faster decisions, and a pipeline that can operate at a scale where using one giant model for everything would be absurd
most people are still trying to build agents with one brain doing every job
this team gave the agent a fast brain for instinct and a powerful brain for judgment
75 billion data points every 11 minutes
and the smartest model only wakes up when something is genuinely difficult
your agent doesn't need a bigger brain.
it needs a second one.
a cheap one that does exactly one job: decide.
look at how many calls in your loop are really just this:
→ is this email phishing?
→ is this post a nugget or slop?
→ is this lead worth a call?
→ is this clip worth cutting?
→ which button does the agent click next?
→ up, down, or "no idea, stay flat"?
none of those need a paragraph.
all of them are getting one. at frontier prices.
so split it:
LLM → thinks + writes
JEV → picks / scores / says yes or no
code → does the thing
three answer shapes cover almost every fork:
CHOICE — pick one option
SCORE — rate it
YES/NO — with a confidence you can gate on
below 0.85? a human looks.
above it? code acts.
nothing irreversible ships without you.
I wrote the whole playbook:
the split, a router you can wire tonight, 8 real use cases from inbox triage to a paper-trading gate, and the guardrails so your first "confident" mistake doesn't cost you.
full article below 👇
everyone's paying frontier prices for their agents to say "yes."
so split the stack.
the big model writes. a cheap brain decides. code does.
then hand the cheap brain every boring fork in your day and say:
"you decide."
here's what that looks like:
> inbox
one subject line, four jobs: billing, phishing, sponsor, real partner.
it picks one. phishing goes straight to quarantine.
the big model never even wakes up.
> feed
your timeline is a firehose.
nugget / slop / skip.
recycled "10 AI tools " threads die before you ever see them.
> moderation
certain spam? hidden.
gray zone? a human looks.
clean "how do I reset my password"? macro.
> clips
one long video, dozens of windows.
it scores the hook and shortlists the few worth cutting.
only those touch the expensive model.
> leads
ignore / nurture / book_call.
only book_call above 0.85 ever reaches your calendar.
> browser
it picks the click.
code clicks it.
the text model only wakes up when something actually needs to be written.
> brain-dump
2-min voice note on a walk.
one dated task + one idea. zero manual sorting.
> trade gate
up / down / unclear.
unclear = stay flat.
your loop stops asking a chat model "maybe?" 60 times a minute.
(paper first. always.)
and this is the weird part:
none of these need a beautiful paragraph.
they need something to look at a situation and pick:
QUARANTINE / REPLY
NUGGET / SLOP
HIDE / KEEP
SHORTLIST / REJECT
BOOK / IGNORE
CLICK A / CLICK B
UP / DOWN / UNCLEAR
that's the whole decision layer.
LLMs think and write. a cheap brain decides. code does.
full setup + code below 👇
I GAVE THE NEW GROK BOT 4.7 A $35 ACCOUNT
BY MORNING, IT HAD TURNED IT INTO $8,900
I did not let one agent watch the market, approve its own idea and control the entire account. The overnight system worked like a small trading desk: one core routed every signal through six narrow jobs before anything could reach execution.
SCOUT ranked fresh market movement. CONTEXT attached regime and liquidity conditions. PROBABILITY converted the setup into a fair estimate. RISK challenged the thesis and killed weak routes. EXECUTION received only cleared decisions. LEDGER locked every input, rejection and result so the next cycle could use the same shared state.
During the eight-hour replay, the router processed 512 candidate signals. Most disappeared before the risk gate. Forty-one survived analysis, 18 reached simulated execution, and eight completed the full route into settlement.
Starting capital: $35
Final equity: $8,900
Simulated PnL: +$8,865
Runtime: 8 simulated hours
What mattered was not the number of agents. It was the separation of authority. The agent that discovered a setup could not approve its risk, the execution layer could not rewrite the thesis, and the ledger preserved the evidence after each route closed.
The interface and routing architecture are real. The market path, orders, balance and PnL are scripted for this demonstration. There are no live orders, and this is not a real trading return.
The overnight upgrade was simple: GROK 4.7 stopped acting like one confident trader and started operating like a desk where every decision had to survive the next gate.
I BUILT A GROK 4.7 POLYMARKET BOT AROUND A KALMAN FILTER
THE 20-DAY REPLAY CLOSED AT +$9,000
The hard part in short-duration crypto markets is that every price move looks like information. One candle jumps, the odds react, and a bot that follows raw movement starts trading noise. I wanted GROK 4.7 to work from a cleaner estimate, so I placed a Kalman filter between the market feed and the decision layer.
The filter keeps one running estimate and updates it with each new observation:
x(new) = x(old) + K × [y − x(old)]
Here, y is the latest observation and K controls how much the new data can move the estimate. A single spike has limited influence. Several moves in the same direction pull the estimate faster.
GROK 4.7 then converts the filtered state into a fair probability and compares it with the market price. If filtered Down is 64% while the market trades Down at 53¢, the model sees an 11¢ gap. The order gate opens only when that gap is large enough to cover the threshold and cost buffer; otherwise, the setup is logged and ignored.
The terminal repeats the same loop across BTC, ETH and SOL Up or Down markets:
→ observe the latest move
→ remove short-term noise
→ estimate fair probability
→ compare probability with market price
→ open only when the edge clears the gate
→ record the decision and update state
Replay length: 20 simulated days
Starting balance: $1,000
Closing balance: $10,000
Simulated PnL: +$9,000
The interface and filtering workflow are real. The market path, orders and PnL are scripted for this demonstration. It sends no live orders and does not represent a real trading return.
The useful idea is simple: GROK does not need to predict every tick. It needs a disciplined way to decide which ticks deserve attention.
I GAVE GROK BOT A SYNTHETIC $40 ACCOUNT
SEVEN HOURS LATER, THE TERMINAL CLOSED AT $5,900
I wasn't testing whether one giant prompt could make money. I wanted to see whether Grok could coordinate a complete trading desk without giving one agent control of the whole account.
So I split the run across seven narrow roles:
→ SCOUT finds fresh signals
→ ANALYST checks context and market depth
→ RISK can kill any route
→ WHALE watches large wallet movement
→ ENTRY receives only cleared setups
→ EXIT manages the open position
→ LEDGER records every decision and outcome
Grok Core sits above the chain. It reads every handoff, resolves conflicts and decides what is allowed to move forward. The agent that finds an opportunity cannot approve its own risk, and the agent that enters cannot rewrite the mandate.
During the compressed seven-hour replay, weak routes died at the risk gate while cleared decisions moved through execution and back into shared memory.
Starting balance: $40
Closing balance: $5,900
Runtime: 7 simulated hours
The terminal and workflow are real. The market path, balance and P&L are scripted for the demonstration. There are no live orders and this is not a real trading return.
The useful part is the architecture: seven jobs, one shared state and one brain keeping the desk synchronized.
I built a browser trading terminal around one synthetic decision engine.
KURT // ORBIT runs entirely in HTML, CSS and JavaScript — no backend, no chart library, no live market feed.
The interface is interactive:
→ switch between BTC, ETH, SOL and AVAX
→ change 1m / 5m / 15m views
→ toggle pulse and candle views
→ inspect price points
→ pause, restart or accelerate the replay to 4×
The orbital intelligence stack cycles through four stages:
INGEST → MAP → STRESS → RELEASE
The 32-second demo contains eight scripted rounds:
$250 starting balance
$3,842 ending balance
+$3,592 simulated P&L
The balance moves during each round and every result lands in the journal. When the replay ends, the final frame holds for inspection.
The UI is real. The market data, order flow and account-growth scenario are synthetic. No live orders. No real returns.
The useful part is the interface pattern: one screen can make a complex agent loop visible, inspectable and easy to demo.
save this if you build product demos. more working AI setups → @0xkurt
I built a browser trading terminal around one synthetic decision engine.
KURT // ORBIT runs entirely in HTML, CSS and JavaScript — no backend, no chart library, no live market feed.
The interface is interactive:
→ switch between BTC, ETH, SOL and AVAX
→ change 1m / 5m / 15m views
→ toggle pulse and candle views
→ inspect price points
→ pause, restart or accelerate the replay to 4×
The orbital intelligence stack cycles through four stages:
INGEST → MAP → STRESS → RELEASE
The 32-second demo contains eight scripted rounds:
$250 starting balance
$3,842 ending balance
+$3,592 simulated P&L
The balance moves during each round and every result lands in the journal. When the replay ends, the final frame holds for inspection.
The UI is real. The market data, order flow and account-growth scenario are synthetic. No live orders. No real returns.
The useful part is the interface pattern: one screen can make a complex agent loop visible, inspectable and easy to demo.
save this if you build product demos. more working AI setups → @0xkurt
10 AI AGENTS. $520 TOTAL. 13 HOURS. ZERO HUMAN INPUT.
here's the setup that's breaking my brain:
a guy gave 10 GPT-6 Astra agents $52 each and one brutal rule —
"earn enough to cover your own subscription, or shut yourself down."
no strategy handed to them. no babysitting. no manual trades.
each agent had to find its own edge in live meme-coin markets.
what happened next:
→ they opened and closed positions on their own
→ they checkpointed state, balanced risk, re-ran strategy
→ one agent couldn't cover its costs and shut itself off — permanently
→ the other 9 kept compounding
13 hours later they reported back a combined $11,300.
from $520. no human touched the mouse once.
think about what that actually means:
the agents didn't just trade. they managed their own survival.
the ones that couldn't pay rent got cut — by themselves.
this is the part nobody's ready for:
the desk doesn't need a trader anymore. it needs someone to spawn agents.
(numbers as reported by the experiment — not financial advice, DYOR.)
follow @0xkurt — i break down the AI-agent meta daily. 🟣