Most people don’t know how accurate their LLM judge actually is.
At @forus we are building AI agents to automate medication access. Accuracy is critical as errors can delay when someone gets medicine they need.
Awesome blog post where @IsabelChien_ breaks down how we evaluated our LLM judge - highly recommend for anyone building evals today.
Most evals can tell you when an LLM judge is right, but not when it's silently wrong.
At Forus, our LLM judge reads prior authorization forms and decides whether a patient's medical record supports them. A missed error delays someone's medication.
To catch those failures, we built a way to generate realistic wrong answers and test the judge against them.
Here's how 🧵
@Austen Core tech was centralized but it was built to be extremely configurable. Everything from pricing, driver onboarding, safety/fraud rules, etc could be configured at the city level.
We had rule engines galore!