New research from Meta.
(bookmark it)
Most factuality work checks whether the claims in an answer are correct. GAMUT goes after the harder question of whether the answer covers everything it should.
It involves a two-level meta-rubric:
It encodes content organization and importance, open-ended sets, ordered processes, and relationships between facts, then compiles mechanically into a flat checklist of binary rubrics an LLM judge can grade reliably.
The benchmark itself is 1,813 questions grounded in real wearable imagery across 10 domains, each with an expert-verified rubric, plus a text-only variant. Across 14 frontier and open models the best score is only 58.7%, and results hold up regardless of which judge model grades.
Coverage is where answers actually fail users, and the rubric-compilation recipe transfers to any task with hierarchical requirements.
Paper: https://t.co/hPjrehEP2T
Learn to build effective AI agents in our academy: https://t.co/LRnpZN7L4c
Confira o vídeo de rarunnycantora! #TikTok https://t.co/frpp9Rj6rQ Esta publicação foi partilhada através do TikTok Lite. Descarrega o TikTok Lite para desfrutares de mais publicações: https://t.co/BJAeJJmEWv