Super cool World-Model results from Genie team! 🧞♂️
Genie 2 generates pixels of an interactive 3D-world from a prompt image.
Amazing progress in only 9 months since the release of Genie-1! 🤯
A cool cameo from our SIMA agent, which can act in this world model out-of-the-box!
Introducing 🧞Genie 2 🧞 - our most capable large-scale foundation world model, which can generate a diverse array of consistent worlds, playable for up to a minute. We believe Genie 2 could unlock the next wave of capabilities for embodied agents 🧠.
@_aidan_clark_ Hard disagree.
This is like saying noone should learn to read or write English because AI will write all text in the future.
Being able to modify AI-generated code will remain a useful skill.
Basic programming literacy is becoming more relevant and accessible, not less!
@NandoDF@scott_e_reed Indeed! Gato was very inspiring to us in SIMA, I remember when it started being blown away by the audacity & boldness to try to fit so many domains/modalities with one model.
Video models are a core part of SIMA's perception, thanks to the Phenaki & SPARC teams for their models!
@j_mcgraph@yaroslavvb I did a bunch of hogwild training of various neural nets & deep RL agents in 2014-2016 (on a big CPU cluster).
Hogwild: poor utilisation of GPU compared to SGD, so slower overall; also hogwild gave worse model results and more variance across runs - making science harder.
So proud of what we and the team have achieved with SIMA!
In this early work, our AI agent plays a number of amazing 3D open world games - but it doesn't play to win, you can tell it what you want it to do...🧵
Introducing SIMA: the first generalist AI agent to follow natural-language instructions in a broad range of 3D virtual environments and video games. 🕹️
It can complete tasks similar to a human, and outperforms an agent trained in just one setting. 🧵 https://t.co/qz3IxzUpto
SIMA is early work towards a General AI Agent, but it is not yet at human-level. We're excited to see where we can take it next - more games, better performance & longer, more complex tasks.
We have work to do, but it's an exciting time to be working on Agents!
Excited to share one of the main things I’ve been working on: scaling towards grounded language agents that can follow natural language instructions across many video game environments—using the same pixels to keyboard/mouse controls as humans do. 1/8
I continue to think the bigger threat of deepfakes is not in convincing people that fake things are real but in offering plausible suspicion that real things are fake