We just got the top community score on ARC-AGI-3's 25 public games: 78.4% and 160/183 levels cleared, where the best frontier LLM session scores 7.8%. No training or demonstrations allowed.
Paper: https://t.co/mKAVXayoQ6
Blog: https://t.co/1i5pPyGxS7
w/ @WenhaoLi29@ScottSanner
#AI #worldmodel #agentic #AGI @arcprize@fchollet
@GregKamradt Following our community leaderboard submission, we would love to present our work on OPINE-World, it's exploration methods and ablations, and current failure modes that we have encountered in our time experimenting with various systems built for ARC-3. Application submitted!
@TheGoldenAnvil Yes! ARC-AGI-3 gives no object representation, action semantic, or domain-specific goal. This makes it quite a lot more difficult than many other domains, meaning our system can likely be adapted to other tasks with a minimal compatibility layer, and even some relaxations.
We just got the top community score on ARC-AGI-3's 25 public games: 78.4% and 160/183 levels cleared, where the best frontier LLM session scores 7.8%. No training or demonstrations allowed.
Paper: https://t.co/mKAVXayoQ6
Blog: https://t.co/1i5pPyGxS7
w/ @WenhaoLi29@ScottSanner
#AI #worldmodel #agentic #AGI @arcprize@fchollet
Learn more about our work at our blog, where we’ve deconstructed some empirical results about our system:
https://t.co/KgRc0FSuhx
And read our full paper here:
https://t.co/rI7fbxcCqx
w/ @WenhaoLi29@ScottSanner
Specifically:
1) An LLM's internal world model cannot be inspected or trusted. We address this by having the agent write the game's dynamics as an object-oriented Python program, accepted when it perfectly models the world or addresses faults as a testable hypothesis.
2) LLMs explore blindly and fall into “hypothesis lock-in”. We address this with ontology error, a Bayesian measure of how much of the game the program still can't explain. The agent spends its moves where that score is highest.
We investigated and found that pure LLMs and existing approaches fail this benchmark for two reasons: their internal model cannot predict the world in an accurate but fault-tolerant manner, and they struggle with hypothesis generation and testing in unknown environments. We propose enhancements to address these shortcomings in our new paper “OPINE-World: Programmatic World Modeling with Ontology-error-Prioritized Interactive Exploration” (https://t.co/mKAVXayWFE)