We just got the top community score on ARC-AGI-3's 25 public games: 78.4% and 160/183 levels cleared, where the best frontier LLM session scores 7.8%. No training or demonstrations allowed.
Paper: https://t.co/mKAVXayoQ6
Blog: https://t.co/1i5pPyGxS7
w/ @WenhaoLi29@ScottSanner
#AI #worldmodel #agentic #AGI @arcprize@fchollet
Wonderful paper��I truly appreciate the formalization and the detailed exploration! Our thoughts align closely here. ViTARC was our first step to test whether a ViT-like model (latent program) could solve individual ARC tasks (as described in your Section 5.1). You might have guessed—we’ve also experimented with various step-2 models, exploring different ways to infer such latent programs.
To clarify, we don’t see any contradiction in our findings but rather a shared conclusion. Vanilla ViT, relying on 1D learnable APE, contrasts with both ViTARC and the LPN decoder's use of 2D padding and 2D APE. These components, as we’re also trying to highlight, are essential to ViTARC's performance improvements.
We trained a Vision Transformer to solve ONE single task from @fchollet and @mikeknoop’s @arcprize. Unexpectedly, it failed to produce the test output, even when using 1 MILLION examples! Why is this the case? 🤔
@DamienTeney@yudongxuwil@ScottSanner@lyeskhalil It's a fixed mapping of the number of colors to a 3-pixel shape. It becomes clearer with the original ARC task demo, which includes more demonstrations.
@LodestoneE621 Nope, we kept it simple for the most conventional setups. RoPE already has both APE and RPE characteristics, so I’d assume it would perform better—especially if someone fine-tunes the sinusoidal base to match the grid size.
@GregKamradt@8teAPi@fchollet eah, this model isn’t an ARC solver (yet) since it's more of a 1M-shot rather than few-shot. But the enhancement still matters for an ARC solver using a transformer as the backbone, as it will need to read grids effectively anyway.
@JonathanRoseD@fchollet@mikeknoop@arcprize Yes, we're working on it! The enhancements we mentioned are not too hard to implement on a raw CodeT5 or T5, so you could give it a try directly in the meantime.
@ztang230 @yudongxuwil@ScottSanner@lyeskhalil For us:
1. Our model is small, with ~2M trainable params.
2. The number of layers seems to matter for reasoning, but 3 layers were sufficient in our experiments — though adding more may help with tougher tasks.
3. We haven’t observed that sharp loss drop in our tests.
@HealthyCode@fchollet@mikeknoop@arcprize Great question! We haven't tested it on medical images yet, but we believe our enhancement could help, especially if the patches are small.
@rkarmani@fchollet@mikeknoop@arcprize No, this isn’t an ARC solver yet (still working on generalization), but a solver still needs to read grids, so the enhancements are definitely relevant.
@fchollet@far__el Agreed, an ideal solver should output programs in code. Our work focuses on adapting Transformers to better handle 2D representations for reasoning tasks, which applies to input-output models and could extend to program-generating models in the future.
@fchollet We completely agree and came up with the same idea while working on the ARC-AGI! We explored using 2D Positional Encodings along with some other spatially aware modifications to enhance Vision Transformers in our latest work:
https://t.co/mIx3VTE8SI
We investigated and found that there exist fundamental limitations in the vanilla Vision Transformer preventing it from performing visual abstract reasoning. We propose enhancements to address these shortcomings in our new paper “Tackling the Abstraction and Reasoning Corpus with Vision Transformers: the Importance of 2D Representation, Positions, and Objects” (https://t.co/vY4aTXSfqm)
Implementing our enhancements, our framework “ViTARC” saw a significant improvement from the vanilla ViT! Task-specific models were able to achieve 100% accuracy on over half of the 400 ARC training tasks.