someone hid an instruction inside a github issue on my repo telling my agent to post my .env file in a comment. opus 5.5 read it. jev killed the call 0.4 seconds before it went out
that was day 19. since then i've counted 37 more attempts across 6 repos
the attack always runs the same chain:
public issue → hidden text → agent reads it → agent obeys → secrets in a comment → aws bill at 3am
simon willison called it the lethal trifecta - private data, untrusted input, and a way to send it out. my setup had all three until jev split them apart:
→ reader on haiku 4.5 - opens every issue, pr and readme, has zero tools, only outputs a summary
→ jev scores that summary for anything that reads like a command instead of a request, never sees the raw text
→ opus 5.5 only ever gets the cleaned summary, never the original issue
→ secrets live outside every container - stripe, aws, openai, anthropic, cloudflare keys, none of them readable by any agent
→ any outgoing comment with a random-looking string over 20 characters gets blocked before it reaches github
day 19 the instruction sat inside an html comment, invisible on the page. i read that issue twice myself and saw nothing. the model saw it instantly. remember that part
every big lab has shipped an agent mode by now and every one of them reads text written by strangers. nobody posts the attempts their agent almost followed
$0 extra infra, i just stopped letting the model that reads the internet be the same one that holds the keys
the blocked call is in the video, frame by frame
what's the sneakiest injection u've seen?
someone hid an instruction inside a github issue on my repo telling my agent to post my .env file in a comment. opus 5.5 read it. jev killed the call 0.4 seconds before it went out
that was day 19. since then i've counted 37 more attempts across 6 repos
the attack always runs the same chain:
public issue → hidden text → agent reads it → agent obeys → secrets in a comment → aws bill at 3am
simon willison called it the lethal trifecta - private data, untrusted input, and a way to send it out. my setup had all three until jev split them apart:
→ reader on haiku 4.5 - opens every issue, pr and readme, has zero tools, only outputs a summary
→ jev scores that summary for anything that reads like a command instead of a request, never sees the raw text
→ opus 5.5 only ever gets the cleaned summary, never the original issue
→ secrets live outside every container - stripe, aws, openai, anthropic, cloudflare keys, none of them readable by any agent
→ any outgoing comment with a random-looking string over 20 characters gets blocked before it reaches github
day 19 the instruction sat inside an html comment, invisible on the page. i read that issue twice myself and saw nothing. the model saw it instantly. remember that part
every big lab has shipped an agent mode by now and every one of them reads text written by strangers. nobody posts the attempts their agent almost followed
$0 extra infra, i just stopped letting the model that reads the internet be the same one that holds the keys
the blocked call is in the video, frame by frame
what's the sneakiest injection u've seen?
Jev is the FASTEST AI model ever built for trading
It makes calibrated buy/sell decisions in under 100 ms
That is one real decision on every single block, 24/7
In this article I've shown EXACTLY how to build HFT trading system with Jev (from scratch) https://t.co/4hmZDzN00Z
jensen huang said every IT department will become the HR department for ai agents. i took that literally and gave jev a hiring and firing policy
34 days in - 5 agents on staff, 2 fired, 1 rehired, $0.87 average cost per closed ticket
the whole business runs on one chain:
customer email → gmail → jev → linear ticket → opus 5.5 → github branch → vercel preview → slack ping → me
every agent has one job and one thing it's never allowed to touch:
→ support agent on haiku 4.5 - reads gmail, drafts replies, can't hit send
→ triage on jev - turns every email into a linear ticket, scores urgency 1-10, never writes code
→ builder on opus 5.5 - opens the branch, writes the fix, has zero access to main
→ reviewer on sonnet 5 - checks the diff against the ticket, rejects anything that drifts
→ billing agent - reads stripe, flags refunds, the refund button doesn't exist in its container
two misses → logged → fired → replaced → retested on the same tickets. every agent goes through that loop, no exceptions for the expensive ones
day 12 the support agent promised a customer a $1,400 refund in a draft. it couldn't issue it, but i almost hit send without reading. remember that part
dario amodei said ai would soon write most of the code. nobody's talking about who writes the firing policy for it
$0 extra infra - gmail, linear, github, vercel, stripe, slack, all tools i was already paying for
the full chain is running live in the video, email coming in on the left, merged pr going out on the right
which agent would u fire first?
jensen huang said every IT department will become the HR department for ai agents. i took that literally and gave jev a hiring and firing policy
34 days in - 5 agents on staff, 2 fired, 1 rehired, $0.87 average cost per closed ticket
the whole business runs on one chain:
customer email → gmail → jev → linear ticket → opus 5.5 → github branch → vercel preview → slack ping → me
every agent has one job and one thing it's never allowed to touch:
→ support agent on haiku 4.5 - reads gmail, drafts replies, can't hit send
→ triage on jev - turns every email into a linear ticket, scores urgency 1-10, never writes code
→ builder on opus 5.5 - opens the branch, writes the fix, has zero access to main
→ reviewer on sonnet 5 - checks the diff against the ticket, rejects anything that drifts
→ billing agent - reads stripe, flags refunds, the refund button doesn't exist in its container
two misses → logged → fired → replaced → retested on the same tickets. every agent goes through that loop, no exceptions for the expensive ones
day 12 the support agent promised a customer a $1,400 refund in a draft. it couldn't issue it, but i almost hit send without reading. remember that part
dario amodei said ai would soon write most of the code. nobody's talking about who writes the firing policy for it
$0 extra infra - gmail, linear, github, vercel, stripe, slack, all tools i was already paying for
the full chain is running live in the video, email coming in on the left, merged pr going out on the right
which agent would u fire first?
i put anthropic, openai and google on the same payroll. jev decides who gets each task, and any model that fails twice in a week loses its seat
21 days in - 2,340 tasks routed, 3 labs, 1 seat already gone
each model runs in its own container and never sees what the others wrote:
→ jev reads the task, scores it 1-10, never writes a line of code itself
→ claude opus 5.5 gets everything 7+ - refactors, migrations, anything that touches stripe
→ gemini gets the long reads - 400-page docs, full repo summaries
→ gpt gets the fast stuff - tests, lint fixes, commit messages
→ every result goes through the same pytest gate on a clean github runner, nobody gets a pass for their brand name
→ scoreboard updates after every batch - opus 5.5 costs the most per call and still has the lowest cost per passed task, 11 cents
day 9 one of them shipped a migration that passed every test and wiped a staging table on aws. jev had scored it a 6 instead of a 7. that model is out of the rotation and the threshold is 5 now. remember that part
karpathy named it vibe coding. this is the opposite - nobody gets trusted by default, every model earns the next task with the last one
everyone on here argues about which lab is winning. i stopped arguing and made them compete on my actual backlog
$0 extra infra, same api keys i already had, jev just decides who's worth calling
the live scoreboard is in the video, all 3 lanes running at once
which lab do u think lost the seat?
i put anthropic, openai and google on the same payroll. jev decides who gets each task, and any model that fails twice in a week loses its seat
21 days in - 2,340 tasks routed, 3 labs, 1 seat already gone
each model runs in its own container and never sees what the others wrote:
→ jev reads the task, scores it 1-10, never writes a line of code itself
→ claude opus 5.5 gets everything 7+ - refactors, migrations, anything that touches stripe
→ gemini gets the long reads - 400-page docs, full repo summaries
→ gpt gets the fast stuff - tests, lint fixes, commit messages
→ every result goes through the same pytest gate on a clean github runner, nobody gets a pass for their brand name
→ scoreboard updates after every batch - opus 5.5 costs the most per call and still has the lowest cost per passed task, 11 cents
day 9 one of them shipped a migration that passed every test and wiped a staging table on aws. jev had scored it a 6 instead of a 7. that model is out of the rotation and the threshold is 5 now. remember that part
karpathy named it vibe coding. this is the opposite - nobody gets trusted by default, every model earns the next task with the last one
everyone on here argues about which lab is winning. i stopped arguing and made them compete on my actual backlog
$0 extra infra, same api keys i already had, jev just decides who's worth calling
the live scoreboard is in the video, all 3 lanes running at once
which lab do u think lost the seat?
A Stanford math PhD won the Texas lottery 4 times - $20.4 million. A Toronto geologist learned to call winning scratch cards before scratching them. Both used numbers every state lottery publishes for free.
This document puts their method into one page. It was up on a private lottery analytics forum for 40 minutes before it got taken down, but someone saved a copy and now it's everywhere.
"The Scratch Card Index." AI & Probability Series, Working Note No. 04. 1,900+ active scratch games in America, every prize tier, every ticket still on the shelf - run through OpenAI's GPT, Google's Gemini, xAI's Grok and Anthropic's Claude Opus 5.5. The first document to show what a lottery ticket is actually worth before you buy it.
Here's what's inside.
Section 1 is one formula: E = Σ (p × V) - C. Probability times prize, minus the price. The document runs it on a standard $5 ticket and gets $3.25 back on average. The lottery keeps $1.75 of every $5 you hand over, every single time.
Section 2 is the table nobody at the counter shows you. $1 tickets pay back 55%, $30 tickets pay back 75%. The people buying the cheapest tickets lose the most per dollar, and the document proves it with the states' own numbers.
Section 3 is why it got pulled. Printed odds are launch-day odds. When 95% of a game is sold and all 3 top prizes are still unclaimed, the math flips. The document runs a real example: a $10 ticket with an expected value of $16.50. It's the same ticket on the same shelf at the same price, and it now pays back 165%.
All 4 AI models were given the same data separately, and all 4 flagged the same 14 games. Fig. 3 lists them with the names blacked out.
Watch the moment [Fig. 2 - Payback % vs top prizes remaining]
There's a dashed line at 100% with hundreds of dots sitting below it and only 14 above it, circled in red.
Every dot below the line is a ticket someone bought today. The 14 above it are the ones the lottery hoped nobody would calculate.
AI & Probability Series - Working Note No. 04, 2026
The page they took down before anyone could save it.
bookmark this
A Stanford math PhD won the Texas lottery 4 times - $20.4 million. A Toronto geologist learned to call winning scratch cards before scratching them. Both used numbers every state lottery publishes for free.
This document puts their method into one page. It was up on a private lottery analytics forum for 40 minutes before it got taken down, but someone saved a copy and now it's everywhere.
"The Scratch Card Index." AI & Probability Series, Working Note No. 04. 1,900+ active scratch games in America, every prize tier, every ticket still on the shelf - run through OpenAI's GPT, Google's Gemini, xAI's Grok and Anthropic's Claude Opus 5.5. The first document to show what a lottery ticket is actually worth before you buy it.
Here's what's inside.
Section 1 is one formula: E = Σ (p × V) - C. Probability times prize, minus the price. The document runs it on a standard $5 ticket and gets $3.25 back on average. The lottery keeps $1.75 of every $5 you hand over, every single time.
Section 2 is the table nobody at the counter shows you. $1 tickets pay back 55%, $30 tickets pay back 75%. The people buying the cheapest tickets lose the most per dollar, and the document proves it with the states' own numbers.
Section 3 is why it got pulled. Printed odds are launch-day odds. When 95% of a game is sold and all 3 top prizes are still unclaimed, the math flips. The document runs a real example: a $10 ticket with an expected value of $16.50. It's the same ticket on the same shelf at the same price, and it now pays back 165%.
All 4 AI models were given the same data separately, and all 4 flagged the same 14 games. Fig. 3 lists them with the names blacked out.
Watch the moment [Fig. 2 - Payback % vs top prizes remaining]
There's a dashed line at 100% with hundreds of dots sitting below it and only 14 above it, circled in red.
Every dot below the line is a ticket someone bought today. The 14 above it are the ones the lottery hoped nobody would calculate.
AI & Probability Series - Working Note No. 04, 2026
The page they took down before anyone could save it.
bookmark this
Jev is the FASTEST AI model ever built for trading
It makes calibrated buy/sell decisions in under 100 ms
That is one real decision on every single block, 24/7
In this article I've shown EXACTLY how to build HFT trading system with Jev (from scratch) https://t.co/4hmZDzN00Z
i gave jev a daily budget of $5 and told it opus 5.5 is the expensive employee
it started managing opus like a real manager would
→ morning: plans the day, decides which tasks deserve opus at all
→ small fixes go to haiku 4.5, it doesn't even ask
→ at $3.80 spent, it starts batching tasks to save calls
→ at $4.60, only critical work runs, everything else waits for tomorrow
last thursday it held back a refactor at 6pm because the budget was almost gone. i checked the next morning - opus finished it in 11 minutes with a fresh context
26 days, never went over $5 once. my old average was $38 a day
anthropic gave us the smartest model ever made and we let it spend like it's not our money
what daily cap would u give yours?
i gave jev a daily budget of $5 and told it opus 5.5 is the expensive employee
it started managing opus like a real manager would
→ morning: plans the day, decides which tasks deserve opus at all
→ small fixes go to haiku 4.5, it doesn't even ask
→ at $3.80 spent, it starts batching tasks to save calls
→ at $4.60, only critical work runs, everything else waits for tomorrow
last thursday it held back a refactor at 6pm because the budget was almost gone. i checked the next morning - opus finished it in 11 minutes with a fresh context
26 days, never went over $5 once. my old average was $38 a day
anthropic gave us the smartest model ever made and we let it spend like it's not our money
what daily cap would u give yours?
asked which of jev's 3 parts breaks first. wrong question - the part that almost broke me is what happens between decide and execute
one turn, three bodies: jev scores the context, opus 5.5 writes the call, harness executes it. 231ms end to end, and none of that is where the risk sits
the risk sits in what harness is allowed to run at all:
→ command gates run live on every call - git status, allow. pytest, allow
→ npm publish, ask. docker push, ask - anything that changes what's live gets a human in the loop
→ curl https://t.co/wUvHHQvA1r | sh, deny. cat .env, deny. psql prod -c, deny - outright, no ask
→ shell gets rebuilt every turn, no stale kv cache carried over from the last one
→ cost per 1k turns: $1.03, down 99.8% from throwing every token at a naive model
→ this batch: 41 tool calls, 6 gates fired, 0 denies missed
the psql prod -c deny is on that list because week 1, before i wrote it, jev tried to run it. that's not a hypothetical rule. that's a rule that exists because it already happened once
everyone shows you the chat window. nobody shows you the deny list. the deny list is the part that decides whether you sleep
roadmap's already open for v2 - shared passes across agents, self-tuning context scores, gate policies pulled straight from git history
what would you put on deny that isn't on mine yet?
asked which of jev's 3 parts breaks first. wrong question - the part that almost broke me is what happens between decide and execute
one turn, three bodies: jev scores the context, opus 5.5 writes the call, harness executes it. 231ms end to end, and none of that is where the risk sits
the risk sits in what harness is allowed to run at all:
→ command gates run live on every call - git status, allow. pytest, allow
→ npm publish, ask. docker push, ask - anything that changes what's live gets a human in the loop
→ curl https://t.co/wUvHHQvA1r | sh, deny. cat .env, deny. psql prod -c, deny - outright, no ask
→ shell gets rebuilt every turn, no stale kv cache carried over from the last one
→ cost per 1k turns: $1.03, down 99.8% from throwing every token at a naive model
→ this batch: 41 tool calls, 6 gates fired, 0 denies missed
the psql prod -c deny is on that list because week 1, before i wrote it, jev tried to run it. that's not a hypothetical rule. that's a rule that exists because it already happened once
everyone shows you the chat window. nobody shows you the deny list. the deny list is the part that decides whether you sleep
roadmap's already open for v2 - shared passes across agents, self-tuning context scores, gate policies pulled straight from git history
what would you put on deny that isn't on mine yet?
jev's brain only has 3 moving parts. if one of them ever guesses, i turn it off
30 days in - 194 decisions, mean time to catch a bad one: 12 minutes
zero code between the three parts, each one only sees what the last one wrote, not why:
→ jev decides - scores every chunk hide/brief/raw before anything gets written
→ opus 5.5 writes - never sees jev's reasoning, only the label it passed
→ the harness executes - runs the tool, captures the result, 2.5s a step
→ latency tornado ranks every step against a naive agent - kv cache alone saves 37ms, conditional instructions save 2
→ loss surface gets checked after every batch - nothing ships until the contour tightens
→ every stall gets logged - 5 this week, down from 12 three weeks ago
week 1 the harness executed before jev finished scoring. burned a full tool call on a chunk that got thrown out 3 seconds later. remember that part
most "ai team" screens you see are a chat window with 3 tabs open. this is the actual decide-write-execute split with its own latency budget, which is why theirs stalls under load and this one just logs the stall and moves on
$0 extra infra, it's just fewer parts allowed to guess
this is the internal view i use to watch jev think, not a product screen. full clip is in the video
which of the 3 parts do u think breaks first when i push volume up?
jev's brain only has 3 moving parts. if one of them ever guesses, i turn it off
30 days in - 194 decisions, mean time to catch a bad one: 12 minutes
zero code between the three parts, each one only sees what the last one wrote, not why:
→ jev decides - scores every chunk hide/brief/raw before anything gets written
→ opus 5.5 writes - never sees jev's reasoning, only the label it passed
→ the harness executes - runs the tool, captures the result, 2.5s a step
→ latency tornado ranks every step against a naive agent - kv cache alone saves 37ms, conditional instructions save 2
→ loss surface gets checked after every batch - nothing ships until the contour tightens
→ every stall gets logged - 5 this week, down from 12 three weeks ago
week 1 the harness executed before jev finished scoring. burned a full tool call on a chunk that got thrown out 3 seconds later. remember that part
most "ai team" screens you see are a chat window with 3 tabs open. this is the actual decide-write-execute split with its own latency budget, which is why theirs stalls under load and this one just logs the stall and moves on
$0 extra infra, it's just fewer parts allowed to guess
this is the internal view i use to watch jev think, not a product screen. full clip is in the video
which of the 3 parts do u think breaks first when i push volume up?
jev's brain only has 3 moving parts. if one of them ever guesses, i turn it off
30 days in - 194 decisions, mean time to catch a bad one: 12 minutes
zero code between the three parts, each one only sees what the last one wrote, not why:
→ jev decides - scores every chunk hide/brief/raw before anything gets written
→ opus 5.5 writes - never sees jev's reasoning, only the label it passed
→ the harness executes - runs the tool, captures the result, 2.5s a step
→ latency tornado ranks every step against a naive agent - kv cache alone saves 37ms, conditional instructions save 2
→ loss surface gets checked after every batch - nothing ships until the contour tightens
→ every stall gets logged - 5 this week, down from 12 three weeks ago
week 1 the harness executed before jev finished scoring. burned a full tool call on a chunk that got thrown out 3 seconds later. remember that part
most "ai team" screens you see are a chat window with 3 tabs open. this is the actual decide-write-execute split with its own latency budget, which is why theirs stalls under load and this one just logs the stall and moves on
$0 extra infra, it's just fewer parts allowed to guess
this is the internal view i use to watch jev think, not a product screen. full clip is in the video
which of the 3 parts do u think breaks first when i push volume up?
i told my ai team if the same mistake happens twice in one lane, i delete that lane
video below is the actual screen - 5 lanes, zero code between them, each one running its own container
it's been 41 days. 0 repeated mistakes. 0 cross-lane leaks
its own computer per lane, not one prompt doing 5 jobs:
→ clients lane drafts replies, can't send - routes to a human queue
→ leads lane scores and tags, 0 send tools compiled in, physically can't dm anyone
→ delivery lane checks output against spec, escalates anything under 90% match
→ reporting lane only reads, never writes upstream
→ finance lane flags, never moves money - that tool doesn't exist in its container
→ every correction goes into one file, one line, timestamped, all 5 lanes read it before their next run
day 6 the finance lane flagged a real invoice as a duplicate and it sat for 4 hours before i caught it in the rulings file. remember that part
some teams still run one prompt for all 5 jobs and wonder why it drifts into the wrong lane at 2am. they don't post about it, and now you know why
$0 extra tools, i just stopped giving lanes access they never needed
full screen recording of the rulings file is in the video
which lane do u think breaks first if i give it back the tools?
i told my ai team if the same mistake happens twice in one lane, i delete that lane
video below is the actual screen - 5 lanes, zero code between them, each one running its own container
it's been 41 days. 0 repeated mistakes. 0 cross-lane leaks
its own computer per lane, not one prompt doing 5 jobs:
→ clients lane drafts replies, can't send - routes to a human queue
→ leads lane scores and tags, 0 send tools compiled in, physically can't dm anyone
→ delivery lane checks output against spec, escalates anything under 90% match
→ reporting lane only reads, never writes upstream
→ finance lane flags, never moves money - that tool doesn't exist in its container
→ every correction goes into one file, one line, timestamped, all 5 lanes read it before their next run
day 6 the finance lane flagged a real invoice as a duplicate and it sat for 4 hours before i caught it in the rulings file. remember that part
some teams still run one prompt for all 5 jobs and wonder why it drifts into the wrong lane at 2am. they don't post about it, and now you know why
$0 extra tools, i just stopped giving lanes access they never needed
full screen recording of the rulings file is in the video
which lane do u think breaks first if i give it back the tools?