Want this run on your production stack?
Bell Tuning Rapid Audit — $2,500, 48hr, 3 clients/week:
https://t.co/fSYTkYMvc3
Or run the tool yourself, MIT. The layer exists. Measuring it is cheap.
cc @LangChainAI
Headline: a query about agent tool-use — the literal core topic of the post — pulled these
alignments:
rank 1: 0.114
rank 2: 0.150
rank 3: 0.209
rank 4: 0.235 ← actually most relevant
rank 5: 0.164
Retriever ranked the LEAST relevant chunk first.
One corpus. One retriever. One embedding model. Run it on yours.
Repro script + audit JSON + 6 charts:
https://t.co/umXQs3QPMe
uickstart-teardown
retrieval-auditor MIT, one command:
npx contrarianai-retrieval-auditor
Adversarial — Q6, "reward shaping" (not in corpus):
OUT_OF_DISTRIBUTION flag fires.
Distinct from OFF_TOPIC. "Retriever broken" and "corpus lacks topic" need different fixes.
Output-side metrics conflate them.
Clean baseline — Q3, "memory mechanisms":
mean: 0.391
health: 0.710
flags: none
Same retriever. Same corpus. Same embedding. Order-of-magnitude variance in health by query.
So the auditor doesn't false-fire — the variance is real.
This is what production RAG audits keep finding.
Eval suite: fine.
Tracing: fine.
Users: "answers feel off."
The gap = distributional shape of retrieval, query by query. Most monitoring flattens it to
top-line averages.
This is what production RAG audits keep finding.
Eval suite: fine.
Tracing: fine.
Users: "answers feel off."
The gap = distributional shape of retrieval, query by query. Most monitoring flattens it to
top-line averages.
Auditor on that query:
rankQualityR = -0.611
scoreCalibrationR = -0.675
3 pathology flags simultaneously
health = 0.099
precision@5 against ground truth would still report "5 chunks retrieved." Eval suite passes
clean.
Headline: a query about agent tool-use — the literal core topic of the post — pulled these
alignments:
rank 1: 0.114
rank 2: 0.150
rank 3: 0.209
rank 4: 0.235 ← actually most relevant
rank 5: 0.164
Retriever ranked the LEAST relevant chunk first.
Ran retrieval-auditor against LangChain's RAG quickstart.
Corpus: Lilian Weng's LLM Powered Autonomous Agents post.
Retriever: LangChain default, top-5.
Queries: 6 things you'd actually ask.
5 of 6 came back flagged.
Want this run on your production stack?
Bell Tuning Rapid Audit — $2,500, 48hr, 3 clients/week:
https://t.co/fSYTkYLXmv
Or run the tool yourself, MIT. The layer exists. Measuring it is cheap.
cc @LangChainAI
One corpus. One retriever. One embedding model. Run it on yours.
Repro script + audit JSON + 6 charts:
https://t.co/umXQs3QPMe
uickstart-teardown
retrieval-auditor MIT, one command:
npx contrarianai-retrieval-auditor
Adversarial — Q6, "reward shaping" (not in corpus):
OUT_OF_DISTRIBUTION flag fires.
Distinct from OFF_TOPIC. "Retriever broken" and "corpus lacks topic" need different fixes.
Output-side metrics conflate them.
Clean baseline — Q3, "memory mechanisms":
mean: 0.391
health: 0.710
flags: none
Same retriever. Same corpus. Same embedding. Order-of-magnitude variance in health by query.
So the auditor doesn't false-fire — the variance is real.
This is what production RAG audits keep finding.
Eval suite: fine.
Tracing: fine.
Users: "answers feel off."
The gap = distributional shape of retrieval, query by query. Most monitoring flattens it to
top-line averages.
Auditor on that query:
rankQualityR = -0.611
scoreCalibrationR = -0.675
3 pathology flags simultaneously
health = 0.099
precision@5 against ground truth would still report "5 chunks retrieved." Eval suite passes
clean.
Your AI isn’t failing at the output — it failed 3 turns earlier in the context window. Here’s the statistical proof + free open-source inspector https://t.co/1gUEZZpaEG https://t.co/YQy94y8JOX @claudeai
🚨RFK JR: "When I was on the plane the other day with President Trump he took a piece of paper and drew a map of the Middle East with all the nations on it. Then he wrote in each country the troop strength. He was looking at the border between Syria and Turkey and asking questions about the men to his generals."
TUCKER: "Wait, you're saying that Trump drew an accurate map of Middle east with troop strengths?"
RFK JR: "Yes." LMAO 😂
Tuckers reaction is incredible.