In agentic pipelines, one LLM can play many roles: writer, reviewer, judge, editor, and more. Does it use the same understanding of a concept across them?
In our new ICML paper, we measure conceptual consistency and test whether more consistent models make fewer mistakes.
We call this the consistency dilemma: when choosing among models, consistency is operationally useful, but doesn’t guarantee reliability.
Much more to do with this metric! We’re excited to study consistency in other settings, and disentangle it from other model-level qualities.