this Jev stack is f*cking crazy
5 open-source repos that let your agent make decisions for cents instead of burning frontier tokens.
each one handles a different part of the job:
ROUTE
> jev-for-all - Jev picks which skill to load and which tools your coding agent gets. 64 requests, 0 wrong picks, ~$0.0001 per decision
https://t.co/XreSBsVgTy
GUARD
> agent-jev - a 0.6B model that reads diffs, logs and traces and answers typed questions in ~60 ms. zero generated tokens. ships with a Claude Code hook that gates every tool call
https://t.co/wdcviYegzg
RUN LOCAL
> jev-rs - a Jev-compatible engine in Rust. plug in any local model, state gets prefilled once, every question is answered from probabilities. 75 ms per question on a Mac
https://t.co/Sjhaai0xcE
TRAIN
> typed-decisions - build your own Jev-style model. DeBERTa version hits 85.5% accuracy at 42 ms, and training cost $0.26
https://t.co/0KyH0OmIBE
SHIP
> jev-usecases - 36 production harnesses: support triage, fraud, invoices, guardrails, SOC agents. every decision lands in auto, confirm, human or block by confidence
https://t.co/C2yl2gDEeC
the cool part isn't any single repo.
it's that your agent can now:
pick the right skill -> gate its own tool calls -> run offline -> learn your decisions -> ship with confidence bands
your LLM doesn't need to think about every yes or no anymore.
the decision layer is already on GitHub. you just plug it in.
bookmark this before you build your next agent.
Jev is brand new and there's no tutorial for it yet.
so I wrote the prompt that makes Claude read the docs, pick one real use case, and build the agent. it's a runnable code, added failure handling, and a rule for when Claude takes over.
paste it in and try it.
Opus 5.5 makes this AI stack look f…cking illegal
10 GitHub repos for building everything around the model
01 zen-mcp-server
▸ https://t.co/hANULvRJWL
→ your agent debates Gemini, GPT and local models mid-task
02 claude-context
▸ https://t.co/2lj3Mf1h1Q
→ semantic code search so the agent stops grepping blind
03 wshobson/agents
▸ https://t.co/WjiOo7Jfr9
→ a roster of subagents, each with one narrow job
04 mcp-agent
▸ https://t.co/VUFhc2ZRsZ
→ workflow patterns built around MCP tools, not around chat
05 Backlog.md
▸ https://t.co/eMfdP75BjH
→ tasks live in git, so the agent and you read the same board
06 Claude Code Development Kit
▸ https://t.co/dgiFcIxEFl
→ hooks, subagents and MCP wired for context at scale
07 agent-rules
▸ https://t.co/2xPnhTwIUE
→ shared guardrails the agent inherits in every repo
08 Agent Skills
▸ https://t.co/5H2D5yaXbw
→ one file that teaches the agent a job, portable across tools
09 agent-scan
▸ https://t.co/JHEsgSfU9l
→ scans your agents, MCP servers and skills for what leaks
10 awesome-claude-code-toolkit
▸ https://t.co/lNRtLsQhFc
→ the map: agents, skills, commands, hooks, MCP configs
the architecture:
rules → context → subagents → tools → state → review → scan
I'd split the stack like this:
context:
claude-context → Development Kit
work:
wshobson/agents → mcp-agent → zen-mcp-server
state:
Backlog.md
safety:
agent-rules → Agent Skills → agent-scan
everyone posts the SDK and the CLI
the layer that decides whether your agent survives a week is this one
the model is one folder in the stack ⭣
Your JEV agent knows the answer. It doesn't know whether it should act on it
every call returns two things: the answer, and the confidence behind it. most setups only read the first one.
confidence comes from the spread. same 4 options, one distribution concentrated at 0.91, one flat at 0.22. identical answer, completely different decision.
the fix is a bar that rises with the stakes:
read-only 0.5 → refund 0.7 → escalation 0.8 → destructive 0.9
below 0.5 the ticket goes to a human. no guessing. two traps most people miss. a noul of 0.5 means "as likely yes as no", not "medium". and a confident "done" doesn't prove the file was saved. confidence is not permission, check the outcome.
pin the version. aliases move, and thresholds tuned on one version quietly break on the next.
check out the full article on Jev below ↓ ↓
building a second brain with jev is insane...
once Jev is wired to your vault, paste these rules into AGENTS.md
they tell Astra when to propose a new memory, when to route a decision through Jev, when to skip a stale page, and when to say “not in the vault”
the copy-ready rules are in the image
the full Jev + GPT-6 Astra build is in the article below👇🏻
Send this Graph Engineering prompt to your coding agent before you build another multi-step workflow.
It stops your agent from grading its own homework, and catches the bugs walking back in twice.
Steal it before your bill finds out how many times it paid for the same mistake.
JEV + OPUS 5.5 IS INSANE FOR BUILDING A COMPANY BRAIN
I pulled the whole architecture out of the TypeSafe and Anthropic docs and packed it into a 14-page PDF
the 10 steps:
1. meet the pair
> Opus 5.5 thinks, Jev decides, your code holds the branch
2. stop asking a text generator for a yes or no
> Jev returns a typed answer with a calibrated probability in 0.44s for $0.00035
3. ask everything at once
> Choice, Score and Noul run in parallel, so the fourth question costs almost nothing
4. branch on the number
> 0.999 goes straight into the if statement. ~99% of turns end right here
5. stop routing blind
> Opus 5.5 to Sonnet and back costs 5.84 against 3.32 for staying on 5.5
6. keep one context warm
> cache reads at $0.20 per Mtok are 20x cheaper than a fresh load
7. escalate the hard part
> the toughest 1% goes to Opus 5.5 with 1M context and 66.4% on Terminal-Bench 4.0
8. score every chunk on every query
> keep whole, summarize or drop. the context gets rebuilt each turn
9. gate the actual command
> every bash call gets classified before it runs, inside your own code
10. judge 100% of runs
> $3.50 a day for 10,000 traces, and it matched the human label on all 500 decisions
the result: a while loop that paid a frontier model for every tiny call turns into a brain that spends a fraction of a cent to notice and pays properly only when it has to think
the person who brings this into their team walks into the budget meeting with the AI bill cut and the output up
the PDF maps the company brain. the loop side of it - how Jev takes a Claude bill from $765 to $3 a month - is in the article below ↓
Researchers from Anthropic, OpenAI, and SpaceX just built an autonomous Meta-Agent Orchestrator on top of Jev, created by Jev Founder, Diogo Almeida (tested across 18,430 runs, arXiv:2606.04455)
It is more useful than most paid AI courses:
this is a 10-step blueprint on how to build a faster, cheaper and more controllable AI system around Claude Opus, GPT-6, Grok or any other LLM:
step 1 → split the responsibilities: LLMs generate, Jev handles bounded semantic decisions, and deterministic code keeps authority
step 2 → build the state: give Jev the active task node, relevant test evidence, and proposed action instead of sending the entire conversation
step 3 → choose the right primitive: Choice selects the branch, Score evaluates rubric thresholds, and Noul computes verification probability
step 4 → replace giant evaluation prompts with atomic questions: intent, urgency, evidence, risk and scope become separate typed decisions
step 5 → put Jev before the LLM: filter context, select agent pools (Opus, Sol, Luna), and set permissions before paying for expensive calls
step 6 → give the model a bounded job: once Jev selects the route, the LLM receives only the minimal AST slices and tools required for that branch
step 7 → put Jev after the LLM: run objective assertions (Luna) to verify outputs stay inside security bounds before committing
step 8 → route by confidence & risk: high-confidence cases proceed automatically, while consequential mutations go to human review
step 9 → batch independent decisions: evaluate multiple Choice, Score, and Noul questions over one shared state instead of cascading LLM calls
step 10 → record complete decision receipts: state hash, branching depth, tool telemetry, latency, and prune status for full auditability
most AI courses teach you how to write a bigger prompt
this architecture teaches you how to build the control system around frontier models
the result: 94.2% benchmark pass rate, $0.18 inference cost per task, 31-level deep reasoning, and zero recursive self-hallucination
The full 31-deep task tree, live policy field telemetry, and 60 FPS Meta-Orchestrator in the clip below ↓
Top 12 agentic use cases for Jev:
(bookmark this)
Jev handles semantic decisions that ordinary code cannot express reliably. It returns typed answers and probabilities, while code continues to cover the workflow.
Here are 12 practical use cases for Jev:
1. Browser next action
> Convert the current DOM state into a bounded action such as click, type, or stop. Code executes only valid operation-target pairs. There are already several open-source Jev web agents.
2. Context compaction
> Decide which events from a long agent trace should remain. The selected text stays verbatim instead of being replaced with a generated summary.
3. Skill and context loading
> Compare the current user turn against the available skills. Load only the instructions needed for that turn instead of filling the context window with every skill.
4. Typed tool-call compilation
> Map a natural-language request to a function and fill its typed arguments. Each argument is evaluated separately before code allows execution.
5. Citation verification
> Check whether a quoted passage exists and whether the surrounding evidence supports the claim. The output can be supported, unsupported, or contradicted.
6. Extraction verification
> Run a cheap extractor first, then use Jev to verify questionable fields. Clean records stay on the fast path while uncertain ones reach a reasoning model.
7. Agent trace evaluation
> Turn raw trajectories into queryable labels such as progress and repetition. This avoids asking another LLM to write a full review of every run.
8. Semantic regression tests
> Replay a trace suite against a new agent build. Semantic checks can then pass or block prompt, model, tool, and policy changes in CI.
9. Jevgrep code search
> Search a codebase by what the code does rather than its exact words. Jev scores candidate snippets and returns the most relevant code first.
10. Entity alignment
> Compare two candidate records and decide whether to merge, review, or keep them separate. Candidate generation remains deterministic while Jev handles semantic identity.
11. Retrieval reranking
> Let embeddings retrieve a broad candidate set, then use Jev to reorder passages by relevance. The generation model receives the most useful evidence first.
12. Memory promotion gate
> Capture a completed agent trace, then judge whether its corrections contain a reusable lesson. Trace-backed lessons can be promoted while task-specific noise is discarded.
If you want to see the final pattern in practice, it is already implemented in the Beacon open-source project.
Beacon captures full sessions across Claude Code, Codex, Cursor, OpenCode, and 20+ agent harnesses, and then Jev identifies which workflows and corrections are worth learning from, so that a lesson discovered by one agent can become available to the others.
GitHub repo: https://t.co/8ee3KmP6th
(don’t forget to star it ⭐)
If you want to dive deeper, I also wrote about a similar mechanism in a hands-on guide. It covers building a Jev-style decision path with open models, entirely locally.
Read it below.
This is f*ckin awesome.
A Stanford team pairs JEV with Claude Code to sort 75 billion data points every 11 minutes.
JEV runs a cheap first pass on everything. Claude only gets the hard cases.
Faster, cheaper, way less compute burned.
I still don't understand why people let the same thing that did the work also decide if the work was good. I split mine three weeks ago and found a mistake that had been passing clean for months
nothing about it looked broken. every result checked out fine, because the thing checking it had no reason to disagree with the thing that made it
two separate studies measured exactly why this fails every time. a model recognizes its own writing well above chance, and that recognition alone tilts it toward liking what it wrote. put a model in a judge seat scoring its own answer against others, and it favors itself every single time, no exceptions
a real check starts from nothing. it never read how the answer got made, never saw the reasoning, never sat in the room while the decision happened. the second it shares even a sentence of that history, it stops checking and starts rubber-stamping instead
here's the part almost nobody looks for. confidence and correctness hand back the exact same number. a model that's sure and a model that's right look identical from the outside, and only one of them ever deserved the trust it got
one thing about the fix holds true at any scale: whoever grades the work has to be far enough away from it that agreeing was never the easy option
think about whatever got marked done today, in your own work, not just an agent's. ask who actually looked at it, and how close they'd been standing to it before they said yes
full guide below
Send this Jev prompt to any LLM or AI agent
It sets up Jev → analyzes you → finds where you waste time and money → upgrades your AI setup beyond 95% of people...
Copy this now. Thank me later:
JEV engineering is not what you think it is.
Most people hear "200x faster, $400 x cheaper" and think: cheaper routing.
That's the wrong frame.
What TypeSafe AI actually built is a split in how AI systems work. Not a cheaper alternative to the LLM. A separate layer that does the one thing LLMs are terrible at: deciding.
Your LLM is a generative engine. It thinks, reasons, creates. But inside every agent loop, it's also doing something else, choosing which tool to run, rating whether a result is good enough, deciding whether to go again. Binary questions. No generation needed. And you're paying frontier model prices for every single one.
Jev reads the system state and returns a typed answer with probability:
Choice - which agent should act next?
Score - is this result good enough to keep?
Noul - is the goal met?
No reasoning. No text. Just the answer.
That's a completely different layer of the stack.
The architecture it creates:
LLM creates work. Jev decides what happens next. Agents execute.
When you build it that way, the economics change completely. 1,018 papers classified for $0.08. 500 emails triaged for 3.5 cents. 300-agent swarms that run on autopilot without a human watching the loop.
This is Jevons Paradox in AI infrastructure. Make decisions cheap enough and you stop rationing them. Systems that were financially impossible become routine.
One model to think. One to decide.
The diagram below maps how the split actually works in a running system.
I still don't understand why everyone is still running agents in a line. I switched to graphs three weeks ago and my fleet finished in the time my single agent used to spend on step two.
what slows every agent system I have seen is not intelligence. it is geometry. and almost nobody is talking about it.
one engineer used this to rewrite 535,000 lines of code in 11 days. a manual rewrite of that scale could take close to a year. it cost $165,000 in tokens. the graph was not cheap. it was just faster than a human year.
a node is one agent with one job. research one competitor. review one file. check one claim. the moment a node owns two independent jobs you lose the ability to parallelize them cleanly, verify them independently, and debug them in isolation. an edge is a dependency. it only exists when data actually moves across it. everything else is a fake edge. a wait you invented that costs time and produces nothing.
find the fake edges and the line collapses into something wider. jobs that can run at the same time run at the same time. what used to take the sum of forty steps now finishes in the time of the slowest layer.
the pattern behind every serious agent system looks like a diamond. fan out to gather breadth, one agent per angle, all at once. reduce with plain code, no model tokens spent. verify with a fresh skeptic on every finding. synthesize once from what survived. Claude's own research feature uses a very similar pattern in production.
the part nobody warns you about: the verifier needs clean context. give it the same conversation the worker had and it is not checking anything. it is nodding along to itself in a different window. a graph of agents sharing one context is a single loop in a costume. it breaks the same way, just later and more expensively.
one rule that holds at every scale. a worker and its verifier must never share a context.
your agents are not too slow. they are waiting in a line that did not need to exist.
full guide in the article. save it before you build your next agent from scratch.
this is free f*cking gold for anyone still gluing agents together with loops
Andrew Ng mapped 4 agentic patterns into one graph instead of four separate tricks
Reflection, Tool Use, Planning, Multi-Agent Collaboration, one flow
User → Architect Agent → Tech Lead Agent → Developer Agent, typed handoffs
feedback loops back to the user at every stage, not just at the end
GPT-3.5 wrapped in this style of workflow: 95.1% on HumanEval. GPT-4 zero-shot: 67%.
give this to your next agent build before you touch the model card again.
Inference Engineering is becoming one of the most important skill sets in AI, and I think a lot of engineers are still underestimating it.
Training gets most of the attention because that is where the model is created. But once a model has to serve real traffic, a completely different class of problems appears. You are suddenly dealing with queueing, batching, KV-cache pressure, GPU memory, scheduling, routing, kernel efficiency, tail latency, and the uncomfortable fact that two requests hitting the same model can have completely different cost profiles.
The first thing worth understanding is that inference itself is not one homogeneous workload. Prefill and decode behave very differently. Prefill processes the prompt and builds the KV cache, with work across prompt positions that can be parallelized efficiently. Decode is autoregressive, each sequence advances token by token while repeatedly reading model weights and accumulated KV state. That is why prefill is often compute-heavy, while decode at practical serving batch sizes can become strongly constrained by memory bandwidth.
Once you understand that split, a lot of modern inference work stops looking like a random collection of tricks. FlashAttention reduces memory traffic inside attention. PagedAttention improves how KV-cache memory is allocated and managed. GQA and MQA reduce the amount of KV state stored per token. Continuous batching lets the scheduler keep mixing active sequences instead of waiting for a fixed batch to finish. Chunked prefill prevents a very long prompt from monopolizing execution. Prefix caching avoids recomputing context the system has already processed.
Scale makes the problem harder. Tensor parallelism reduces the per-device model memory footprint at the cost of communication. Pipeline parallelism introduces scheduling complexity and pipeline bubbles. MoE serving adds expert routing, communication, and load imbalance. Some systems even separate prefill and decode onto different workers because they have sufficiently different resource profiles that optimizing them together can become difficult.
This is also why tokens per second is a dangerously incomplete metric. A production system has to care about TTFT, inter-token latency, throughput, queueing delay, KV-cache utilization, batch composition, GPU utilization, p95/p99 latency, and goodput, meaning how much useful request throughput actually meets the latency objective. A server can show excellent aggregate throughput while individual users are still waiting far too long for their first token.
The interesting part is that inference engineering sits at the intersection of ML systems, distributed systems, compilers, GPU architecture, networking, scheduling, and performance engineering. As models get larger and serving volumes grow, improving the model is only half the problem. Making every GPU deliver more useful work becomes its own engineering discipline.
The model decides what token should come next. Inference engineering determines how much serving that token costs, how quickly it reaches the user, and how many users your hardware can support at once.
That is why I think inference engineering is going to matter far more over the next few years than most people currently realize.
300 agents shouldn’t leave you with 300 answers. they should leave you with one structure.
that’s the whole idea behind this K3 context graph setup.
the swarm starts messy:
300 agents
→ hundreds of research paths
→ 140 linked sources
→ 9 contradictions surfaced
→ 100% traceable claims
then the graph starts compressing everything.
same entity found twice → linked
3–4 sources agree → confidence rises
2 sources disagree → contradiction flagged
1 source only → claim stays weak
so the useful output isn’t the swarm itself.
it’s what survives after the swarm disappears.
one graph → shared context.
every claim connected back to where it came from.
300 agents do the searching. the graph is what turns the search into memory.
full breakdown in the article below ↓
Four caches in LLM serving, clearly explained:
Every LLM request reads the whole prompt and computes attention state for every token in it.
This step is called prefill, and it impacts both the input bill and the time before the first token appears.
In an agent loop, most of the prompt comprises text that the model already processed in the previous turn.
There are four cache layers that prevent paying for the processed tokens at each turn.
↳ The KV cache holds the key and value tensors for every token at every layer, for one active request.
↳ Prefix caching keeps those tensors on the server instead. vLLM stores them in 16-token blocks and identifies each block by a hash that chains in the previous block's hash, so a block only matches if everything before it matched too. The scheduler stops at the first miss and prefills the suffix from there.
↳ Prompt caching is the same reuse that runs on a provider's hardware, with a price sheet attached. Anthropic charges 1.25x the base input rate to write an entry and 0.1x to read it.
↳ Semantic caching works differently. It embeds the incoming prompt, runs a similarity search over stored prompts, and returns a stored answer outright when the score is above a threshold.
That's why it saves output tokens as well as input tokens. It's also why every request pays for an embedding round trip, including every miss.
The first three match on exact tokens and cannot change what the model produces.
This technique matches on similarity, which means it is quite susceptible to generating a wrong response since embeddings may match to a wrong prompt.
The diagram below depicts all these techniques.
To use these techniques, you don't need to build a custom serving stack.
The transformers library already implements the cache as an object of KV vectors that you can preserve, so you can prefill a corpus once, retain the returned tensors, and reuse them across queries in about ten lines.
And this KV cache is only one of four separate caching layers in an LLM stack.
The other three are prefix caching on the server, prompt caching billed by a provider, and a semantic cache that skips the model entirely.
I wrote a full breakdown of all four caches in LLM serving that you should know as an AI engineer, with code for each.
Read it below.
CPU vs GPU vs TPU vs NPU vs LPU, explained visually:
(bookmark this)
5 hardware architectures power AI today.
Each one makes a fundamentally different tradeoff between flexibility, parallelism, and memory access.
> CPU
It is built for general-purpose computing. A few powerful cores handle complex logic, branching, and system-level tasks.
It has deep cache hierarchies and off-chip main memory (DRAM). It's great for operating systems, databases, and decision-heavy code, but not that great for repetitive math like matrix multiplications.
> GPU
Instead of a few powerful cores, GPUs spread work across thousands of smaller cores that all execute the same instruction on different data.
This is why GPUs dominate AI training. The parallelism maps directly to the kind of math neural networks need.
> TPU
They go one step further with specialization.
The core compute unit is a grid of multiply-accumulate (MAC) units where data flows through in a wave pattern.
Weights enter from one side, activations from the other, and partial results propagate without going back to memory each time.
The entire execution is compiler-controlled, not hardware-scheduled. Google designed TPUs specifically for neural network workloads.
> NPU
This is an edge-optimized variant.
The architecture is built around a Neural Compute Engine packed with MAC arrays and on-chip SRAM, but instead of high-bandwidth memory (HBM), NPUs use low-power system memory.
The design goal is to run inference at single-digit watt power budgets, like smartphones, wearables, and IoT devices.
Apple Neural Engine and Intel's NPU follow this pattern.
> LPU (Language Processing Unit)
This is the newest entrant, by Groq.
The architecture removes off-chip memory from the critical path entirely. All weight storage lives in on-chip SRAM.
Execution is fully deterministic and compiler-scheduled, which means zero cache misses and zero runtime scheduling overhead.
The tradeoff is that it provides limited memory per chip, which means you need hundreds of chips linked together to serve a single large model. But the latency advantage is real.
AI compute has evolved from general-purpose flexibility (CPU) to extreme specialization (LPU). Each step trades some level of generality for efficiency.
The visual below maps the internal architecture of all five side by side.
To dive deeper into GPU specifically, Akshay wrote a detailed article on it.
It builds up from first principles why memory and compute compete, why that gap exists in the hardware, and what makes a workload memory-bound in the first place.
Read it below.