7/ Have you watched your AI return an opposite-meaning passage in production, and did a cross-encoder reranker fix it cleanly or did you need additional checks?
RAG fundamentals: Cross-encoders catch negation
1/ Why does your AI retrieval return passages that mean the exact opposite of what the user asked about? Bi-encoders smooth across negation. "Not deprecated" and "deprecated" land near each other in embedding space.
6/ The cost - Reranker runs only over top 50 from the bi-encoder. Added latency around 100ms on CrossEncoder or 30ms on FlashRank. Cheap insurance against returning opposite-meaning passages.
7/ Cost: FlashRank and CrossEncoder free at inference. Cohere around $1 per 1k searches. Which reranker turned out to be the right choice for your latency budget, and did you measure or guess?
RAG fundamentals:
1/ Why does choosing between FlashRank, sentence-transformers CrossEncoder, and Cohere's reranker matter more than which embedding model you pick? The reranker is where your latency budget actually goes.
RAG fundamentals:
1/ Why does your retrieval slow to a crawl when you try to use a cross-encoder as your only retriever? Math reasons: Skipping the bi-encoder breaks latency budgets that the architecture was designed around.