@aicheye How do you determine the location at which the racket has to intercept, do you do a trajectory prediction based on a few frames and then run IK for the racket center to reach there?
@AgilexRobotics Cool! How do you guys deal with the force, do you have a sensor input going into the VLA or is there some other way the robot is dealing with that during opening and closing ๐
Seeing @GeneralistAI a few days ago and now @DynaRobotics i kinda understand why Egocentric Human Demonstrations are so important and how companies collecting this data are getting popular!
Super excited to share Dyna-2, the first robot foundation model trained on over 1 million hours of human data. At this scale, we saw the emergence of a cross-embodiment transfer scaling law: training on increasing amount of human data not only improves model prediction on held-out human data but also on robot data the model has never seen before ๐คฏ
We validate that this transfer scaling law does translate to on-robot performance across 3 different robot platforms, and uncovers that both training objective and data matter greatly for this emergence. Crucially, our results position video as a new scaling axis for physical AI.
Beyond these scaling-law oriented results, we also dissect Dyna-2 greatly, going in-depth on its many capabilities, including language following via world modeling, enhanced robustness and precision, zero-shot customer site deployment, and some cool video generation results!
Read our technical blog post for more details! I do think this is a very important result that shows a very different path for robot foundation models from the ones we are marching on. Really excited for the road ahead. We have more releases coming, stay tuned!
@andyzengineer Amazing how naturally the robot performs. Curious to get your thoughts on how task level reasoning would eventually come in the GEN Models. Would a separate model eventually sit on top of GEN-1 for longer horizon tasks, the way Gemini Robotics paired ER 2 with a VLA?
5/ And now, I am actually curious on how would they eventually bring on a higher level reasoning into the model. For eg : If the task was to cook a meal, then apart from just manipulation some sense of reasoning and planning would have to take place
1/Spent a few hours reading @GeneralistAIโs technical blogs today, the GEN-1 model is definitely a game changer. The thing that i found deeply impressive was the way of thinking the company has gone through to build such high capabilities!
We've improved how GEN-1 learns to adapt to new actuators and new robots at the lowest level, with up to 10-20x gains on internal benchmarks. This significantly boosts performance on high-precision tasks like disassembling parts from a NIST board.
Read more about GEN-1 in our blog posts in the comments below.
4/ Having read all the interesting concepts and their terminology on Mastery and how their models have mastered some tasks, it feels like they are solving how do we push โmanipulation of objectsโ to be efficient.
3/ I liked how they talk about System 2 and System 1 thinking and that the System 1 kind of thinking to Robotics can only ever be got via human demonstrations and not by teleoperation. Quite intriguing to see how ideas from psychology are shaping how we build the models!
2/ Their blog on โGoing Beyond VLAs and WAMโ (https://t.co/WVLxBFxZR7) clearly shows the ideology of being a goal driven company rather than an idea driven company and how that led them towards the path of pretraining from scratch instead of using a VLM with an action head.