Preliminary trials of the cable-driven Discover2Walk (D2W) robotic rehabilitation platform for small children with #cerebral#palsy competed! 🤖
Grateful to Hospital Niño Jesús team and above all, the families.
@BioroboticsCSIC@CARobotica_
I'm excited to share that the project I started my PhD with nearly five years ago continues to grow. @Googleorg has supported our research at @BioroboticsCSIC to apply generative AI in robotic therapies for children with cerebral palsy.
@CSIC@CARobotica_
It's Olympics week 🏃♂️ 🏊♀️ 🏓 🎾 , and what better time to share some exciting news! 🎉
After countless hours over ~6 years and unwavering dedication from our incredible team, we're thrilled to announce a major breakthrough in AI and robotics. 🏓🤖 #Robotics#AI#SportsNews
OpenAI is expected to demo a real-time voice assistant tomorrow. What does it take to deliver an immersive, or even magical experience?
Almost all voice AI go through 3 stages:
1. Speech recognition or "ASR": audio -> text1, think Whisper;
2. LLM that plans what to say next: text1 -> text2;
3. Speech synthesis or "TTS": text2 -> audio, think ElevenLabs or VALL-E.
Last year, I made the figure below to show how to make Siri/Alexa 10x better. However, naively going through 3 stages results in huge latency. User experience falls off the cliff if we have to wait 5 seconds for *each* reply. It breaks the immersion and feels lifeless even if the synthesized audio itself sounds real.
Natural dialogues fundamentally don't work like this. We humans
> think about what to say next at the same time as we listen & speak;
> inject "yes, hmm, huh" at appropriate moments;
> predict when the other person finishes and immediately take over;
> decide to talk over the other person organically, without being offensive;
> handle interruptions gracefully. Currently, AI assistants either cannot be interrupted (super frustrating) or simply stop when they detect an audio event and lose train of thought;
> engage in group chat. We are so good at multi-agent conversations.
It's not as simple as making each of the 3 neural nets faster, sequentially. Solving real-time dialogue requires us to rethink the whole stack, overlap each component as much as possible, and learn how to make interventions in real time.
Or perhaps even better - just have 1 NN mapping audio to audio. End-to-end always wins.
I'll sketch out how to design such a model and its training pipeline. Meanwhile, let's wait and see how far OpenAI pushes it!
Excited to share our new Nature paper! In this work, we propose a new display design that pairs inverse-designed metasurface waveguides with AI-driven holographic displays to enable full-color 3D augmented reality from a compact eyeglasses-like form factor.
1/8
AlphaFold-3 is out, the latest iteration of the greatest breakthrough in AI for biology. What's new is that AlphaFold-3 uses diffusion to "render" the molecular structure. It starts from a fuzzy cloud of atoms and then materializes the molecule gradually through denoising.
We live on a timeline where learnings from Llama and Sora can inform and accelerate life sciences. I find this level of generality absolutely mind-boggling. The same transformer+diffusion backbone that generates fancy pixels can also imagine proteins, as long as you convert the data to sequences of floats accordingly.
We are not there yet at a single AGI model, but we have successfully built a menu of general-purpose AI recipes that transfer training, data, and neural architectures across domains. This should not work, but thank god it does!
Muy felices de poder comunicar que el @CARobotica_ ha conseguido la acreditación de Excelencia ASPIRA-MaX Josefa Barba del @CSIC
Ahora a trabajar para la segunda fase e intentar conseguir el Sello de Excelencia ASPIRA-MaX Sagrario Martínez-Carrera. #MaX_CSIC
We trained a robot dog to balance and walk on top of a yoga ball purely in simulation, and then transfer zero-shot to the real world. No fine-tuning. Just works.
I’m excited to announce DrEureka, an LLM agent that writes code to train robot skills in simulation, and writes more code to bridge the difficult simulation-reality gap. It fully automates the pipeline from new skill learning to real-world deployment.
The Yoga ball task is particularly hard because it is not possible to accurately simulate the bouncy ball surface. Yet DrEureka has no trouble searching over a vast space of sim-to-real configurations, and enables the dog to steer the ball on various terrains, even walking sideways!
Traditionally, the sim-to-real transfer is achieved by domain randomization, a tedious process that requires expert human roboticists to stare at every parameter and adjust by hand. Frontier LLMs like GPT-4 have tons of built-in physical intuition for friction, damping, stiffness, gravity, etc. We are (mildly) surprised to find that DrEureka can tune these parameters competently and explain its reasoning well.
DrEureka builds on our prior work Eureka, the algorithm that teaches a 5-finger robot hand to do pen spinning. It takes one step further on our quest to automate the entire robot learning pipeline by an AI agent system. One model that outputs strings will supervise another model that outputs torque control.
We open-source everything! Welcome you all to check out the paper, more videos, and try the codebase today: https://t.co/RwiBT3z78H
Code: https://t.co/ERp4Gl0N36
Yesterday Mohammad Awad from Khalifa University of Science and Technology (UAE) visited the BioRobotics group and the Field Robotics group of the @CARobotica_
Thanks for this visit! .
After two years in stealth mode, we're thrilled to unveil Mentee Robotics and our humanoid robot, @MenteeBot! With AI integration at every layer, from Sim2Real machine learning to NeRF-based algorithms and LLMs, we've achieved a complete end-to-end cycle for tasks.
This is just the beginning and we can’t wait to share more
Congratulations to our PhD student @Romerozabal who has presented this week the scope of his thesis at the scientific days of the @CARobotica_ research center with his talk "Soft robotics to promote walking in young children with cerebral palsy".
📚Last week we had the pleasure to receive the visit of the students of Industrial Electronics and Automation Engineering of the @urjc_uni! Thank you for visiting us and we hope you enjoyed testing our devices. 🤖
#research@CARobotica_@CSICdivulga
Ayer tuvo lugar la primera Jornada Científica en el Centro donde algunos de nuestros doctorandos hicieron unas buenísimas presentaciones. Gracias a Juan Medina, David Pont, Pablo Romero y Alberto Villalonga por vuestras exposiciones. ¡Bravo!
Our paper is online! It introduces the design and validation of the REFLEX prototype, a unilateral active knee–ankle–foot orthosis designed and developed to naturally assist the paretic limbs of hemiparetic patients during gait. https://t.co/jRrRzriQ4Y
Plasticity in the stepping pattern is the feature that makes stepping an ideal target for clinical intervention. The degree of responsiveness in the pattern to variations in treadmill parameters opens exciting new avenues for using the treadmill to train early infant locomotion.