Amazon just open sourced a decision model small enough to run on your laptop.
wifi off. no api key. no bill.
it's called Strands Decider 2B. it doesn't write. it picks from options you give it, right on your machine.
Amazon's Marc Brooker on why they built it: "they make a perfect decider for a workflow step"
the cloud is starting to look optional.
Amazon just open sourced a decision model small enough to run on your laptop.
wifi off. no api key. no bill.
it's called Strands Decider 2B. it doesn't write. it picks from options you give it, right on your machine.
Amazon's Marc Brooker on why they built it: "they make a perfect decider for a workflow step"
the cloud is starting to look optional.
a browser agent found real flights in 7.1 seconds. total cost $0.0039.
it didn't reason through the page. every step it picked from a fresh list of buttons.
a small model only typed 2 words.
a browser agent found real flights in 7.1 seconds. total cost $0.0039.
it didn't reason through the page. every step it picked from a fresh list of buttons.
a small model only typed 2 words.
traders found the first real use for a model that can't write.
one decision per block. under 100 ms. while a chat model is still typing its first sentence.
your trading bot doesn't need to think. it needs to decide.
traders found the first real use for a model that can't write.
one decision per block. under 100 ms. while a chat model is still typing its first sentence.
your trading bot doesn't need to think. it needs to decide.
how the split works:
→ Jev: which worker, which files, merge or retry, risk gate
→ LLMs: only write code
→ tests: decide what merges
→ you: only when risk is high
you stop paying frontier prices for every "what next"
same pattern running live: https://t.co/kwqkKe9gDe
the best coder on your agent team might be the one that can't code.
Jev writes zero lines. it picks who works, which files, when tests pass, what ships.
Claude and GPT just write.
and the moment anyone touches billing, it stops and asks you.
@batagonx Length bias and self preference are just algorithmic ego. That is why raw LLM evaluation is broken without human grounding and strict output length normalization.
a $0.044 judge just matched GPT-6 on routine evals. at 0.36% of the fee.
CMU tested Jev, a model that can't write a single word, as an LLM judge. the paper is free.
not a benchmark flex. not a demo. a cascade.
Jev judges everything. when it's under 0.9 sure, it hands the case to the frontier model. 34% escalated. 91.3% accuracy vs 91.7% for GPT-6 alone. 47% of the fee.
$0.044 per 1,000 judgments vs $12.182. same call, 277x cheaper.
now picture it as a research desk. arxiv, github, x all night. Jev screens, routes, verifies, judges. the LLM only reads and writes. you wake up to 1 brief.
teams pay frontier prices to grade every single output. the smart ones only pay for the hard 34%.
you're still paying a genius to say yes.
100 emails. fraud check. 1.42 seconds. 7 cents.
jev does not write a report. it answers one question:
clean or not.the expensive model only sees the unsure 20.
you are still paying a genius to say this is not a scam.
100 emails. fraud check. 1.42 seconds. 7 cents.
jev does not write a report. it answers one question:
clean or not.the expensive model only sees the unsure 20.
you are still paying a genius to say this is not a scam.
a $0.044 judge just matched GPT-6 on routine evals. at 0.36% of the fee.
CMU tested Jev, a model that can't write a single word, as an LLM judge. the paper is free.
not a benchmark flex. not a demo. a cascade.
Jev judges everything. when it's under 0.9 sure, it hands the case to the frontier model. 34% escalated. 91.3% accuracy vs 91.7% for GPT-6 alone. 47% of the fee.
$0.044 per 1,000 judgments vs $12.182. same call, 277x cheaper.
now picture it as a research desk. arxiv, github, x all night. Jev screens, routes, verifies, judges. the LLM only reads and writes. you wake up to 1 brief.
teams pay frontier prices to grade every single output. the smart ones only pay for the hard 34%.
you're still paying a genius to say yes.
your inbox is burning a frontier model on newsletters.
a former OpenAI researcher raised $40M to fix that. a model that can't write a single word. it only decides. and the docs are free.
not an assistant. not a plugin. not another chatbot. a triage layer that reads each email in under half a second and picks a lane.
agencies will charge you thousands for an "AI inbox". the whole thing is 5 boxes.
inbox → Jev triages → Opus drafts → Jev checks → sent, or it lands on you
the part nobody talks about. most of your mail dies at step 2 for fractions of a cent. Opus only wakes up for the emails that actually need a reply.
Riley Brown ran 500 emails through Jev. 3.5 cents. done in seconds.
and your judgment lives in 1 number. want it careful with clients? move auto send from 0.7 to 0.9. no prompt rewrite. no praying.
a year from now every inbox runs like this. right now almost nobody has even split reading from deciding.
you're paying a genius to read your spam.
your inbox is burning a frontier model on newsletters.
a former OpenAI researcher raised $40M to fix that. a model that can't write a single word. it only decides. and the docs are free.
not an assistant. not a plugin. not another chatbot. a triage layer that reads each email in under half a second and picks a lane.
agencies will charge you thousands for an "AI inbox". the whole thing is 5 boxes.
inbox → Jev triages → Opus drafts → Jev checks → sent, or it lands on you
the part nobody talks about. most of your mail dies at step 2 for fractions of a cent. Opus only wakes up for the emails that actually need a reply.
Riley Brown ran 500 emails through Jev. 3.5 cents. done in seconds.
and your judgment lives in 1 number. want it careful with clients? move auto send from 0.7 to 0.9. no prompt rewrite. no praying.
a year from now every inbox runs like this. right now almost nobody has even split reading from deciding.
you're paying a genius to read your spam.
@Teka1900 only thing that scares me here is raw email going straight into the state. one weird email and your triage gets talked into auto sending something dumb to a client
a former OpenAI researcher just built an AI that can't say a single word.
it cut a Claude session from 1M tokens to 86K. in 1 second. and the docs are free.
not a chatbot. not a wrapper. it doesn't even output text. it only decides. and $40M just bet that this is where agents go next.
agencies charge $997 to teach "agent orchestration". the whole thing is 1 line.
LLM writes. Jev decides. code acts.
week 1. builders already did this:
→ flights Zurich to London found in 7.1s for $0.0039
→ 1,018 research papers sorted for 8 cents
→ 500 emails triaged for 3.5 cents
→ Vercel's safety check: more accurate than GPT, 4.7x faster
the vendor says 193x. real tests say 4 to 8x. that's still the gap between a demo and a product.
most agent builders haven't even heard the name yet. the ones who have are already ripping LLM calls out of their loops.
your agent is still asking a novelist which button to press.