Introducing EdgeBench, a benchmark designed to study how agents learn from environments over at least 12~72-hour runs. We find that performance follows a log-sigmoid function of environment interaction time with high precision.
EdgeBench is built with three ingredients:
- 🌍 Real & Diverse: 134 real-world tasks across 6 task categories, spanning scientific problems, professional knowledge work, software engineering, optimization, formal math, and games.
- ⏳ Ultra-Long-Horizon: Each task supports 12–72 hours of agent work. Recorded human effort averages 57.2 hours.
- 🔁 Informative Feedback: Agents receive real-world feedback for continuous improvement.
After 38,000 hours of agent runs on EdgeBench, a scaling law for learning from environments emerges:
- 📈 As agents interact with task environments over time, their aggregate performance is precisely fit by a log-sigmoid function.
- 🧠 This phenomenon can be explained by an elegant theory of graph exploration.
We are releasing an initial 51 of the 134 tasks, together with the full evaluation framework, to help advance long-horizon agent research. Check our blog & paper for more findings!
Blog https://t.co/nMOzFsOhbT
Paper https://t.co/rZb3eWuvik
GitHub https://t.co/oemXd4UrFw
Dataset https://t.co/P4SQMrM47o
Details below 👇🧵
#NVIDIAIsaac GR00T N1.5 is now accessible to #robotics developers working with a wide range of robot form factors, and available to download from @huggingface. 🎉
Dive into our step-by-step tutorial to learn how to easily post-train and adapt it to the LeRobot SO-101 arm, and put your knowledge to the test by joining the worldwide LeRobot Hackathon happening this weekend.
Learn more. https://t.co/0vLF8Xvpli
Thrilled to announce GR00T N1, our open foundation model for generalist humanoid robots!
GR00T N1 adopts a dual-system design, leverages the entire data pyramid for model training, and supports various robot embodiments.
GR00T N1 embodies years of fundamental research, spanning compositional autonomy stack, synthetic data generation, and scalable training algorithms.
We have made our whitepaper, pre-trained models, training datasets, and codebase publicly available. We can't wait to see how our efforts accelerate the future development of humanoid robotics!
🌐 Codebase: https://t.co/RUIjGYBIJR
🧩 Tech blog: https://t.co/zq2rwzxTij
📃 Whitepaper: https://t.co/pbx9hYOkyl
🚀So excited to share our recent work! GR00T N1 is a 2B model for humanoid robots, which is validated on a series of sim and real robot benchmarks🙌! Try it out!
Excited to announce GR00T N1, the world’s first open foundation model for humanoid robots! We are on a mission to democratize Physical AI. The power of general robot brain, in the palm of your hand - with only 2B parameters, N1 learns from the most diverse physical action dataset ever compiled and punches above its weight:
- Real humanoid teleoperation data.
- Large-scale simulation data: we are open-sourcing 300K+ trajectories!
- Neural trajectories: we apply SOTA video generation models to “hallucinate” new synthetic data that features accurate physics in pixels. Using Jensen’s words, “systematically infinite data”!
- Latent actions: we develop novel algorithms to extract action tokens from in-the-wild human videos and neural generated videos.
GR00T N1 is a single end-to-end neural net, from photons to actions:
- Vision-Language Model (System 2) that interprets the physical world through vision and language instructions, enabling robots to reason about their environment and instructions, and plan the right actions.
- Diffusion Transformer (System 1) that “renders” smooth and precise motor actions at 120 Hz, executing the latent plan made by System 2.
We deploy N1 on GR1 robot, 1X Neo robot, and a large collection of simulation benchmarks. N1 achieves up to +30% boost in diverse manipulation tasks for household and industrial settings.
While humanoid robots are the main focus of N1, our model also supports cross-embodiment. We finetune it to work on the $110 HuggingFace LeRobot SO100 robot arm! Open robot brain runs on open hardware. Sounds just right.
Let’s solve robotics, together, one token at a time.
Links to our Whitepaper, Github repo, HuggingFace model, and open dataset page in the thread: 🧵