Just wanted to announce my paper on Theoria (91.4% accuracy @ 57% coverage on Humanity's Last Exam (HLE)-Verified Gold text only), something I just released! The goal is to bring deductive reasoning into the heart of how LLMs operate: the model has to *prove* its answer step by step, and skeptical judges audit every step.
Along the way it flagged 33 problems as potential issues in HLE itself, including 10 fully certified proofs that disagree with the published answer key. No human expert has reviewed them yet. Help adjudicate → https://t.co/wq6A8VVd83 (user: HLE / pass: thankyou)
Paper: https://t.co/jx4cUULURq
Code: https://t.co/uiy0MtsELy
Subjects we need adjudicated:
math, physics, chemistry, biology/medicine, CS, engineering, and the humanities
Just uploaded a rough draft of something I've been working on, I'd love some feedback. https://t.co/XqlSaMZv3x
Current statistics on HLE-Verified Gold text only: 99%* accuracy@57% coverage, n=184.
*counts cases where Theoria's answer disagreed with the key but appears to have caught key errors as "correct." Strict floor is 88%. Feel free to try yourself. Requires both claude cli and codex cli.
I created a thing that analyzes scientific reasoning produced by llms. Current statistics on HLE-Verified Gold text only: 99%* accuracy @57% coverage, n=184.
https://t.co/OdJPqGI1cb
*counts cases where Theoria's answer disagreed with the key but appears to have caught key errors as "correct." Strict floor is 88%. Feel free to try yourself. Requires both claude cli and codex cli.