Here’s are some of the experiments and observations I did as part of the initial testers on the locksmith game using within ARC-AGI-3 (my template is available in the repository) 🧵
Today, we're announcing a preview of ARC-AGI-3, the Interactive Reasoning Benchmark with the widest gap between easy for humans and hard for AI
We’re releasing:
* 3 games (environments)
* $10K agent contest
* AI agents API
Starting scores - Frontier AI: 0%, Humans: 100%
After 7 years in San Francisco, I had a good life. Nice apartment, good career.
But something felt like it needed to be more.
So I bought a van and moved in.
SF → San Diego → Vancouver → Toronto.
Queen bed, sofa bed, two stoves, propane, Starlink, solar panels. I wired the power setup myself.
It’s uncomfortable. You’re always hunting for showers, bathrooms, gyms.
It’s also the most stimulating way I’ve ever lived.
@ChrSzegedy@iam_mjk@wtgowers@Noahpinion It is.
Generative architectures, particularly when they generate discrete symbols, are not the right thing to understand the real world.
The real world high-dimensional, continuous, noisy, messy, insanely complex, and largely unpredictable.
Releasing bev-decision-150K - a high quality dataset to train decision models (like JEV)
Contains 150K examples across diverse domains, tasks, question types - either procedurally or synthetically generated from existing LLM training open datasets.
https://t.co/3BMPwVmzxZ
Our paper was accepted into NeurIPS 2026!
This is my first paper in NeurIPS and the second one at Abundant AI. I am so glad to have worked with a cracked team on the reward hack analysis especially @neversupervised@Pran_Ker and @rishi_desai2
Congrats to the other teams who got accepted!
Can coding agents stay coherent over a 1 billion token budget?
Can they build Slack from scratch?
Rewrite a JAX codebase in PyTorch?
Build a C compiler in Rust?
Enter SWE-Marathon: a benchmark for autonomous long-horizon software work.