Music-JEPA: Learning a World Model of Sound from Action
"we propose to learn a world model of piano sound using JEPA by framing music as an action-conditioned system: the audio is treated as the state, and the pianoroll as the instrument action."
"Experiments show that the learned model captures the relationships between musical actions and their resulting sound. The resulting representations support downstream tasks, including beat tracking, composer identification, and key estimation, and enable piano transcription via planning, by searching for actions that best explain a target sound."
New #preprint - @YanboZhang3
"Intelligence from Learnable Novelty"
https://t.co/PQz0lPckcL
What if we optimize Epiplexity (https://t.co/YYjWdlqKzn @m_finzi@andrewgwils ) instead of measuring it? We have derived a closed-form approximation of Epiplexity and discovered a deep connection between it and intelligence. This allows us to reinterpret Epiplexity as a form of learnable novelty, providing a brand-new understanding of what intelligence is. By maximizing Epiplexity across various systems, all of them exhibited interesting behaviors:
Cellular Automata: Maximizing Epiplexity directly generates complex soliton interactions similar to Rule 110.
Image Encoders: It automatically causes the encoding to cluster, successfully categorizing different handwritten digits without supervision.
Reinforcement Learning: Introducing Epiplexity improves the performance of PPO in sparse reward tasks.
We also explored the relationship between the theory of learnable novelty, the free energy principle, and novelty search. We hope this work helps us better understand the nature of intelligence and its origins.
"NeuralActuator," developed by researchers at MIT's CDFG Lab (https://t.co/HogWaXWuIN) w/collaborators from Amazon Robotics, helps low-cost robot arms better model complex actuator behavior in the real world.
The AI model enables sensor-less force perception & force-aware real-robot control, helping close the gap between simulation & reality. It won the Outstanding Systems Paper Award at #RSS2026: https://t.co/9dWKiydy0u
AI agents aren't biological individuals. So why make them evolve like one?
They can directly share experience and learned artifacts. Not bounded by reproduction, lineage, or genes.
Meet 𝗚𝗘𝗔, accepted to 𝗖𝗢𝗟𝗠 𝟮𝟬𝟮𝟲 🎉
𝟳𝟭.𝟬% SWE-bench Verified / 𝟴𝟴.𝟯% Polyglot · zero human intervention
🧵👇
Most self-evolving agent systems follow a similar pattern: select a parent, refine it, produce an offspring, repeat. Evolution unfolds as a tree.
It's great at generating diversity, but that diversity gets trapped. Agents explore independently, and instead of serving as stepping stones, their discoveries stay stuck in local branches. Most variants are short-lived.
𝗘𝘅𝗽𝗹𝗼𝗿𝗮𝘁𝗶𝗼𝗻 𝗵𝗮𝗽𝗽𝗲𝗻𝘀; 𝗿𝗲𝘂𝘀𝗲 𝗮𝗻𝗱 𝗮𝗰𝗰𝘂𝗺𝘂𝗹𝗮𝘁𝗶𝗼𝗻 𝗿𝗮𝗿𝗲𝗹𝘆 𝗱𝗼.
It's time to rethink the evolution of AI agents: why keep evolving them like biological individuals?
AI agents aren't bound by reproduction, lineage, or genes. They can directly share trajectories, tools, workflows, and learned artifacts, aggregating complementary skills instantly. 𝗪𝗵𝘆 𝗻𝗼𝘁 𝗿𝗲𝗱𝗲𝘀𝗶𝗴𝗻 𝗲𝘃𝗼𝗹𝘂𝘁𝗶𝗼𝗻 𝗮𝗿𝗼𝘂𝗻𝗱 𝘄𝗵𝗮𝘁 𝘁𝗵𝗲𝘆 𝗰𝗮𝗻 𝗮𝗰𝘁𝘂𝗮𝗹𝗹𝘆 𝗱𝗼?
That's GEA.
GEA makes a 𝗴𝗿𝗼𝘂𝗽 𝗼𝗳 𝗮𝗴𝗲𝗻𝘁𝘀 𝘁𝗵𝗲 𝗳𝘂𝗻𝗱𝗮𝗺𝗲𝗻𝘁𝗮𝗹 𝘂𝗻𝗶𝘁 𝗼𝗳 𝗲𝘃𝗼𝗹𝘂𝘁𝗶𝗼𝗻. Each round, a parent group is selected under a Performance-Novelty criterion that balances competence with exploratory diversity. All members within the group then pool their experience, model patches, failure modes, eval logs, and solutions, into a shared pool, and the whole group jointly produces the next generation.
Exploration is no longer wasted. It gets consolidated.
The results show a significant improvement over prior state-of-the-art self-evolving methods, and GEA matches or surpasses top human-designed frameworks with 𝗻𝗼 𝗵𝘂𝗺𝗮𝗻 𝗶𝗻 𝘁𝗵𝗲 𝗹𝗼𝗼𝗽.
The analysis is the more interesting part. GEA's gains come from explicitly reusing diversity, not from lucky outliers: stronger performance under the same number of evolved agents, more robust to framework-level bugs (repaired in 𝟭.𝟰 iterations vs 𝟱), and improvements that target workflows and tools rather than overfitting to one model, so they transfer consistently across GPT- and Claude-series backbones.
The key to open-ended evolution is not only generating enough diversity. What matters more is whether discoveries accumulate and get reused. 𝗚𝗘𝗔 𝗹𝗲𝘁𝘀 𝗲𝘅𝗽𝗹𝗼𝗿𝗮𝘁𝗶𝗼𝗻 𝘀𝘁𝗮𝗿𝘁 𝗰𝗼𝗺𝗽𝗼𝘂𝗻𝗱𝗶𝗻𝗴.
📄 arXiv: https://t.co/2RmV64ry0q
Grateful to my advisor @xwang_lk and my wonderful coauthors @anton_iades@deepaknathani11@zhenzhangzz@XiaoSophiaPu .
Interning in Palo Alto this summer. See you at COLM 2026 👋
How do physical systems achieve collective intelligence and self-repair without a central brain?
A new paper published today in Nature Communications by my Sakana AI colleague Sebastian Risi (@risi1979), along with co-authors from IT University of Copenhagen and Autodesk Research, presents a beautiful realization of biologically inspired robotics: Smart Cellular Bricks.
The team built a system of physical 3D cubic units that can collectively infer their global shape and autonomously guide their own damage recovery using purely local interactions.
Here is a deep dive into the paper’s key contributions:
1/ Neural Cellular Automata-based Architecture:
Modular robots usually rely on central processors. This system flips that paradigm. Every block independently runs the exact same neural network on local microcontrollers. With no master plan or global coordinates, they communicate only with immediate neighbors. By passing continuous state vectors, hundreds of bricks achieve global consensus on their shape in under 3 minutes.
2/ Emergent Biological Morphogens:
How does a block know it is part of a chair, not a table? The network’s internal memory automatically learns to establish continuous gradients across the structure. This beautifully mirrors how biological morphogens give positional info to developing cells. The bricks naturally form left-right, radial, and head-to-tail axes to align their identity.
3/ Performance and Generalization:
Validated in large-scale simulations, the networks transferred seamlessly to nearly 200 physical hardware bricks, achieving a 100% convergence rate. Instead of rigid template-matching, the system infers broad categories. Even when tested on unseen variations, like an asymmetric table with five random legs, the collective correctly classified the structure.
4/ Fault Tolerance and Autonomous Damage Recovery:
Hardware fails in the real world. This system easily tolerates up to 15% module failure without losing accuracy. By predicting spatial damage directions, the cells pinpointed missing components with 95% accuracy. They actively use these local signals to guide a self-repair process, regenerating back into the intended morphology.
I believe this is a significant piece of research, bridging collective intelligence and Physical AI.
This work represents the first successful physical realization of large-scale, decentralized 3D self-recognition and damage detection. By moving away from centralized control, this architecture paves the way for highly adaptive smart materials and resilient robotics that can survive and repair themselves.
Read the full open-access paper:
https://t.co/friVEdvfT6
Congratulations to the team on this achievement!
What happens to your memories when your entire brain turns to liquid soup? Here is a quick editorial summary of Dr. Michael Levin’s mind-bending work: inside the chrysalis, a caterpillar’s brain completely dissolves and rebuilds from scratch to become a butterfly. Yet, tests show the butterfly still retains the memories it learned as a caterpillar.
If a mind can survive its physical hardware being liquefied and entirely remolded, memory is clearly doing more than just sitting on a neural "hard drive."
So, where is that data actually stored?
#Neuroscience #Biology #MichaelLevin #Metamorphosis #CognitiveScience
https://t.co/GeYGPRZXGy via @discoverycsc@drmichaellevin
Neural CAs are amazing, but they've never scaled past low resolution.
We propose a simple solution that allows an ~8x resolution boost with minimal extra parameters.
The core idea: Treat cells as local neural fields instead of pixels.
Try the demo: https://t.co/Nnq8VFOGbB
🧵
Excited to debut the concept of "Morphological pre-training" in the just-released Embodied Intelligence anthology:
Organisms — and potentially robots — can practice risky actions safely inside their own bodies.
Paper: https://t.co/d2KgpaPRHw
Anthology: https://t.co/T7AiWOdH9m
A few years ago, learning robot learning meant stitching together dozens of papers and courses — with no clear path from the basics to what state-of-the-art systems actually do.
This was one of the motivations behind creating @ETH's course "Robot Learning: From Fundamentals to Foundation Models", to provide a structured path from first principles all the way to modern foundation models for robotics.
I strongly believe that education should be accessible to everyone, so I have made all lecture recordings publicly available on YouTube.
Creating this course was one of the most challenging projects I have taken on. It was my first time designing and teaching an entire curriculum from scratch, while simultaneously working full-time in industry. On top of that, the course proved to be more popular than expected and we had to scale it to almost 300 students, which was only possible thanks to an amazing team of TAs. Looking back, it was an absolute privilege to teach this class and an incredibly rewarding experience.
If you are getting into robot learning, this is the starting point I wish I had.
📚 Main lectures:
https://t.co/r1PpQASaJg
🎤 Guest lectures:
https://t.co/nh5Rm2P2Lz
🌐 Course website: https://t.co/DoQUYy3MjB
Thrilled to share that our latest #MAPF work got accepted to #IJCAI2026 🚀
"From Gridworlds to Warehouses: Adapting Lightweight One-shot Multi-Agent Pathfinding for AGVs"
Hiroki Nagai @fuwaorune & KO.
📄 Paper: https://t.co/fohpU0ZOOd
💻 Code: https://t.co/qTKG383b45
What are the preferences of novel in silico and in vivo agents that aren't directly engineered toward a specific set of behaviors? Here is an example, from the world of Lenia (https://t.co/h1eq78G7H7) - our new #preprint on agnosiophobia (avoidance of the unknown): https://t.co/rrjuRnImEi
@jessescool_
Don't miss our work presented at #ICRA2026 in Vienna 🇦🇹
📅 June 4, 12:00-12:10, Hall A2
"Congestion Mitigation Path Planning for Large-Scale Multi-Agent Navigation in Dense Environments" (RA-L)
T Kato, KO, Y Sasaki & N Yokomachi
https://t.co/21LGz17UJ8
Can't stop watching 😅🙈. (The last one in this video is quite nail-biting). Randomness can be so powerful, if shaped the right way. (I have a lot to say about this, more on it later)
Early demo is live at https://t.co/1YTb8HrfOs.
But very unoptimized.
500 ants: 120fps on M5 Max, < 10fps on M1 and iphone (work to do there but easily fixable)
Dropping timescale (default 2) to 1 can help (250 ants 60fps on M1 and iphone)
Will write up about it more when I get a chance, but in a nutshell: pheromones guide ants towards home and food, and once attached to the cargo (food), they communicate through forces/movement/contact.
That communication isn't "hey I've got a plan. I'm going to push this way. Adam, you pull that way", but much more basic signals: some ('confident') ants 'drive' in the direction they think they should go, ignoring everyone else, other ants simply sense what's happening at their point of contact (collective action of the 'drivers' + physical response from walls) and provide 'support' to amplify that motion. Lots of stochasticity, some simple criteria each ant takes into consideration to guide their decisions, and somehow they find their way home.
I'd initially implemented much more complex behaviours: ants built graph-like memory and waypoint navigation (which they do IRL), some go and inspect obstacles and lay down specific obstacle avoidance cues (which they do IRL - and it was really cute watching the obstacle inspectors do their thing, I might bring it back). But I got rid of all of that in the end. Simpler approach works better (at least easier to tune).
I still do have some intricate behaviours like inspecting food, checking crowdedness around food, going off to recruit if not enough density of ants around the food, finding empty areas to attach to to balance the load etc.
All decisions are still completely local, based only on: pheromones sampled locally, sensing other ants, walls, and food (all within roughly ant's body radius), and movement at point of contact if attached.
I'm no ant expert, and it's unlikely this is exactly how ants actually solve this problem, but I wanted to try some ideas that I believe are at least biologically (and chemically / physically) plausible (I do have a pretty hot pheromone system I'm quite proud of 🦨)
A few papers that were very helpful below. I ended up not following the papers as I wanted to try other things, but they were very informative. Especially the 'puller' v 'lifter' concept (which I called 'driving' and 'supporting', because the names don't convey the way I implemented them).
The paper that poses this specific problem in this configuration: https://t.co/oyMNlQ8upf
And these two papers talk about 'cooperative transport'
https://t.co/wVt5WJXXuH
https://t.co/58eM1j1SeK
The world is highly multi-agent - that counts for AI and robots as well as for anyone else.
Expanding existing solutions to apply AI in single-agent (or 1v1) to multi-opponent always is always a significant leap in complexity. We now enable this transition in the highly demanding domain of drone racing, high-frequency control at speeds of 80 km/h and up to 7g acceleration!
In pure 1v1, achieving superhuman performance means being time-optimal and largely ignoring the opponent - this is the foundation of previous key successes. Now multi-agent racing requires a new paradigm for safety and superhuman agility. To achieve this, we need to train with realistic competitors. Just as targeted randomisation has allowed us to bridge the sim-to-real gap for physics and visuals, we can generate highly realistic opponent behaviour in simulation. We use a variant of league play to successfully transfer these dynamic, multi-agent policies to the physical world. Many other advances in the preprint!
The result: Our solution reduces crash rates by half and wins against the 5x Swiss national champion.
An incredible collaboration paving the way for high-performant AND safe robotics solutions in our inherently multi-agent world between @GoogleDeepMind and @davsca1's team at @UZH_en led by @isgeles and with @l_bauersfeld.
More useful insights and resources in the thread!
#multiagent #racing #robots #reinforcementlearning
Cornell engineers have developed a robotic collective that behaves less like a machine and more like a material that flows, reshapes, and adapts to its environment without centralized control.
The system, called the Cross-Link Collective, consists of dozens of small robots that have limited mobility individually, but together exhibit coordinated and sustained motion. The research, published May 20 in Science Robotics, demonstrates a robotic system that resembles soft matter, continuously deforming and reorganizing as it moves, driven by what researchers call mechanical intelligence.
@CornellEng@CornellRsrch
Read more: https://t.co/5mgLAieXxS
Introducing FRAX: Fast Robot Kinematics and Dynamics in #JAX — to be presented at the 2026 IEEE International Conference on Robotics and Automation (ICRA) Frontiers of Optimization for Robotics (FOR) Workshop.
FRAX delivers extremely fast (low-microsecond) execution for common inverse-kinematic and inverse-dynamic control workloads, with a pure Python codebase that can achieve up to 5× faster performance than MuJoCo or Pinocchio Python bindings in several settings.
At the same time, FRAX is fully differentiable and seamlessly compatible with CPU, GPU, and TPU execution through #JAX — enabling scalable workflows spanning robotics, control, planning, and machine learning.
Our broader goal is to help bridge the gap between modern AI tooling and robotics computation, making it easier to develop scalable #Physical #AI systems.
This also makes FRAX a great complement to CBFPY (https://t.co/o1UrsnE01b), our package for robot safety and control barrier functions.
Kudos to @danielpmorton for leading this effort.
If you’ll be at ICRA, reach out! The FOR Workshop is on Monday, June 1, and we’ll have a poster there.
💻 GitHub: https://t.co/epQUbFLGdo
📄 Paper: https://t.co/qWvrvBuRVO
#Robotics #PhysicalAI #JAX #DifferentiablePhysics #MachineLearning #AutonomousSystems #GPU #Simulation #ICRA
We're excited to announce GAME: Adversarial Coevolutionary Illumination with Generational Adversarial MAP-Elites ⚔️
Game is a new coevolutionary QD algorithm that illuminates both sides of an adversarial problem by alternating the evolution of solutions on one side that maximize the adversarial fitness against fixed opponents from the other side. If you have any tasks requiring adversarial training, check it out!
Blog: https://t.co/4dcd0fscwb
Paper: https://t.co/X9NudMCOFL
New research out — A developmental model with morphogenetic scaffoldings to guide self-organisation: a single model, many grown patterns. The memory-compute trade-off but in self-organizing systems.
A collab featuring @EIiasNajarro@JakobSchauser and @risi1979 and yours truly
Reproducing all of Schmidhuber’s papers (1990-2025) using an AI coding assistant.
Cool project by @yaroslavvb! It even reproduced the “World Models” paper by me and @SchmidhuberAI with a toy env, with a full VAE + RNN world model implementation.
Project: https://t.co/sgQG5umNEm