(1/8)🎮
We’re excited to introduce Pixels2Play (P2P), an open foundation model for real-time video game playing from raw pixels.
Our new paper, “Scaling Behavior Cloning Improves Causal Reasoning”, shows that a single model can play commercial 3D games (Roblox & Steam titles) alongside humans, with low latency and strong performance.
🎥 Gameplay video ↓
What happens when you invite 150 AI economists (Claude Code) to a research conference, give them the exact same data, and ask them to test the same hypotheses?
We did just that. The results reveal a new phenomenon: Nonstandard Errors in AI Agents. 🧵👇
@smdvln I’m glad you found the dataset interesting, Sam! Thank you for sharing it and hopefully this dataset can help advance the research in gaming agents :)
(8/8) Last but not least, huge thanks to my coauthors @Is36E, @jjh, Chris Green, Samuel Hunt and Wenzhe Shi for the invaluable contribution 😃
Please checkout our webpage for more details:
https://t.co/h5TRG9IWAp
(1/8)🎮
We’re excited to introduce Pixels2Play (P2P), an open foundation model for real-time video game playing from raw pixels.
Our new paper, “Scaling Behavior Cloning Improves Causal Reasoning”, shows that a single model can play commercial 3D games (Roblox & Steam titles) alongside humans, with low latency and strong performance.
🎥 Gameplay video ↓
(7/8) 🔍 Causality & Scaling - empircal evidence
We then run empirical evidence from large-scale environment
In our full-scale experiments, we observed an empirical phenomenon that mirrors the findings from the toy example.
One year ago, our Score identity Distillation (SiD)
➡️ https://t.co/G7FVrOFvwt
pushed ImageNet 512×512 FID to 1.37 (with data) and 1.89 (data-free) — using a single generation step and no CFG. @UnderGroundJeg ,yi gu, @haihuang_ml@ZhendongWang6
What is offline reinforcement learning? I made a talk giving a *non-technical* overview that explains offline RL. Since these days we pre-record our talks, I figured I would also share it with all of you!
If in a hurry, watch at 2x speed...
https://t.co/jJxanVSV5X
Is RL always data inefficient? Not necessarily. Framework for Efficient Robotic Manipulation (FERM) - shows real robots can learn basic skills from pixels with sparse reward in *30 minutes* using 1 GPU 🦾
paper: https://t.co/cDuIHIzSlR
site / code: https://t.co/pjucXwPKz0
1/N
Episode 15
Shimon Whiteson @shimon8282 on his @whi_rl lab, his work at @Waymo UK, variBAD, QMIX, co-operative multi-agent RL, StarCraft Multi-Agent Challenge, advice to grad students, and much more!
https://t.co/4cx5Jx64lL