How does a model understand and explore a world it has never seen?
We introduce 🔭VISTA🔭, a visual harness that gives a VLM long-horizon vision for reasoning in an interactive world. With Claude Opus 5.0, it reaches 100% RHAE on @arcprize's ARC-AGI-3, perfectly solving all 25 public games.
Blog post: https://t.co/bYUWwi0KoQ 🧵
VISTA is also not limited to 2D games. The same approach works when the same world is rendered as a 3D scene, or serialized into a 1D text grid, opening up more opportunities to generalize VISTA to other complex interactive tasks!
Language is discrete. Language models don’t have to be.
🧚Introducing ELF🧚♀️: Embedded Language Flows—a class of diffusion models in continuous embedding space based on continuous-time Flow Matching 🧵