Jev for RAG, clearly explained!
Hybrid search gives you a shortlist. It does not decide which passages contain evidence, which are merely adjacent, or whether the evidence is strong enough to answer.
That missing judgment is where Jev fits.
Jev sits between retrieval and generation. It does not replace BM25, embeddings, or the LLM. It evaluates the candidates they produce before those candidates enter the context window.
→ Retrieve wide
Combine dense and keyword search, then merge the results with reciprocal rank fusion. This gives Jev a broad candidate pool, such as the top 20 passages shown in the visual.
Retrieval still sets the ceiling. If the right passage is missing from this shortlist, Jev cannot recover it.
→ Judge every candidate together
Send the query as Jev’s state. For each candidate, ask a typed yes-or-no question such as “Does C7 help answer this query?”
Jev evaluates every question in one packed request and returns a calibrated probability for each passage. This avoids making a separate model call for every query-passage pair.
→ Let code apply the threshold
Your application compares each probability with a threshold. Candidates above it continue to the LLM. Everything below it is removed before generation.
Jev makes the fuzzy judgment. Code remains responsible for the actual decision.
→ Gate the entire answer
The same request can check whether the retained passages make the query answerable. It can also flag signs of prompt injection inside a candidate.
If answerability falls below the threshold, the application can skip the LLM and return “not in the documents.” The injection score should remain a filtering signal, not a security boundary.
The result is a cleaner division of work.
Hybrid search retrieves broadly. Jev reranks, filters, and decides whether sufficient evidence exists. The LLM writes only from the passages that survive.
Jev’s value here is not simply moving passages up or down a list. It turns relevance into an explicit probability that your application can inspect, threshold, and act on.
To summarise:
- Retrieval finds the candidates.
- Jev decides what deserves context.
- The LLM writes the grounded answer.
----
I also built an open-source project showing how to use Jev as a judge for AI observability with Comet Opik.
It evaluates support traces for groundedness, request coverage, action honesty, and helpfulness, then records the results as an auditable experiment.
You can explore the project here: https://t.co/XVmpg6ldHT
My article on how Jev works is quoted below.
Anthropic's cost optimization skill is worth reading line by line
Run /claude-api cost-optimize on any repo calling the API
It profiles where your tokens go, ranks levers by savings, and applies free wins first (caching, input hygiene, batch) before tradeoffs (effort, model)
One diff per lever, measured against your eval
Read the sections, It'll help you optimize any agent app you're building
https://t.co/BQEvBrXi90
The Data Engineering Professional Certificate is now available on DeepLearningAI.
Across four courses, design and build the systems that generate, ingest, store, transform, and serve data, including batch and streaming pipelines on AWS and open-source tools.
Built in partnership with @awscloud and taught by Joe Reis, co-author of Fundamentals of Data Engineering.
Enroll for free: https://t.co/idIf3m2Ff1
Meet the winners of The WebMCP Challenge.
These 10 projects show what people and agents can build together when websites expose structured tools agents can use.
🧵 See the winning projects and the builders behind them:
Today, with over 100 industry partners, we introduced the NVIDIA Open Agent Safety Platform, bringing together OpenShell and Sentry.
Artificial intelligence is extraordinary technology that will advance discovery, productivity, security, health, and prosperity for generations to come.
But its full promise can only be realized when people have confidence that AI is being built to be safe and deployed with wisdom and responsibility.
This is bigger than a single product. It's the beginning of an open ecosystem to build the trust layer for safe agent systems.
Together, we are building the foundation of the AI economy.
Trust and innovation are not in conflict. Safety is how trust is earned. We must build not only the most capable AI, but the most trusted AI, so that this extraordinary technology can realize its enormous promise for the world. https://t.co/ugYWQ1MyRi
Grow your AI skills with the new Microsoft Copilot 🚀
Did you hear the news? Microsoft is reimagining Copilot, bringing familiar tools together with new capabilities that help people and organizations build, customize, and scale AI across work.
And we’ve got just the playlist for you. Explore the AI at Work playlist to learn when to use Chat, Cowork, Autopilot, and Code for different outcomes—and how these modes can work together to take an idea, question, or task through to completion. Plus, get up to speed on AI FinOps and prepare for what’s next with agents.
Start the playlist today.
🔗 https://t.co/c5UU77N5OM
Introducing typesafe/jev-router: a cache-aware model router powered by Jev and @typesafeai
The Jev Router picks the best model and reasoning effort for each request, balancing quality, speed, and cost.
Here's how it works 👇🏻
Trending repository of the day 📈
paperclip
The open-source app everyone uses to manage agents at work
Last 24h: 1,853 ⭐
Total: 83,390 ⭐️
https://t.co/MHHbhHpU2g
Today we announce our new unified multi-agent framework that provides creators with a system for generating temporally consistent, long-form video narratives while mitigating visual drift and pipeline error propagation. Learn more: https://t.co/FWbKBvca3s
This is the most valuable free resource we've created on AI Evals 🎉 (not exaggerating!)
I organized our public materials into this guide. It allows you to find answers to your eval problems w/o searching aimlessly.
Humans: pick the row that sounds like you.
Agents: point it at the post
The material draw on 60+ hours of office hours from our Evals course, where @sh_reya and I have taught 5k engineers and PMs Evals.
We add new material often, with 15 FAQs added in the last two weeks. Recent additions are marked with a "New" badge. You can find all of these and more here:
https://t.co/yVp03nPmbY
"I aspire to be an incredibly lazy prompter." —@_lopopolo
Ryan joins The Agent Factory to make the case for not writing any code manually, and break down how harness engineering can help you get there → https://t.co/tC8Z9I73nx
10 GitHub repositories every developer should know for System Design 👇
1- System Design Primer
https://t.co/pCptYob0Md
2- System Design 101
https://t.co/BNsXmyGv06
3- System Design by Karan Pratap Singh
https://t.co/6Dmu3dZxVj
4- Awesome System Design Resources
https://t.co/t2dCo3YEOG
5- Awesome Scalability
https://t.co/wTWsl8Oiyc
6- Awesome System Design
https://t.co/2kBZmfzidZ
7- System Design Interview
https://t.co/G2lzMrHMh9
8- Machine Learning Systems Design
https://t.co/A7ZTXQsXRT
9- System Design Academy
https://t.co/i29k4FiAhc
10- Agentic Design Patterns
https://t.co/oNwcdrOLCd
✅ Save this for your System Design preparation.
Follow for more developer resources, interview prep, jobs & opportunities.