Highly performant open weights frontier models such as Kimi are a competitive threat to OpenAI & Anthropic, but probably for everyone else these are a win. Hope more US entities will release top quality open weights models as well. Government regulation of AI models to prevent public harm could be done equally for proprietary or open weight models. Classic arguments in favor of open source software apply to AI models too.
"A parcel with snacks has been delivered for Flexion. Retrieve it using the stairs and come up using the elevator. Then unpack it and place the items into the empty drawer on the shelf in the snack area."
One instruction. No human operator. Everything that follows is autonomous.
Today we're introducing Reflect v1.0, our robotics intelligence platform for long-horizon work. From a single natural-language command, the robot understands the task, navigates a multi-floor building, calls elevators, handles doors, uses tools to unpack a box, and puts the items away. The biggest shift in v1.0 is that we use reinforcement learning across every layer, from low-level control to high-level reasoning.
Long-horizon autonomy is unforgiving. The robot must recover on its own when things don't go to plan because in the real world, they never do. Combining reasoning, perception, physical execution and runtime robustness into a single mission-capable system is the foundation required to solve humanoid autonomy.
Our team is just getting started.
#HumanoidRobots #Flexion
Today, we give robots a /skills library that self-evolves and compounds indefinitely! Introducing ASPIRE: a robot solving its 100th task is no longer as clueless as solving its first. Coding agents observe multimodal sensory traces from simulation and real robots, launch an evolutionary search over control programs, and distill the best know-how into an ever-expanding library.
ASPIRE is a new type of continual learning: "training" is skill refinement instead of gradient descent.
"Trained model" is a repo of sensorimotor skills instead of floating weights.
“Distributed training” is a panel of agents each practicing a different skill instead of sharded minibatches.
Here's the beauty: ASPIRE gives the tired terms "sim2real transfer" and "cross-embodiment transfer" a whole new meaning. Bridging the sim-to-real gap is notoriously brutal. An end-to-end policy has to swallow both the visual shift (sim looks toyish next to a real camera) and the subtle contact physics it never quite gets right. ASPIRE sidesteps the mess, because it doesn't ship pixels or weights across the gap, but ships the know-how. The robot still has to practice in the real world, not zero-shot, but it gets there way faster because it isn't rediscovering the strategy from scratch. Same for going single-arm to bimanual hardware, which usually requires new data and retraining from zero. ASPIRE achieves up to ~10x cut in "transfer learning” tokens (yes, tokens are the new unit of *training* compute ;)
Check out our gallery of 150+ tasks and 90+ skills the robots taught themselves, all on the website! Kind of wild that we can ship the "learned weights" as an HTML page rather than a GGUF. We'll open-source the full stack so your own robot library starts compounding from ours!
Deep dive in thread:
why are no data factories selling post-training data?
take a base policy, pick a task, dagger it to three 9s reliability, and sell the dagger data as a premium bundle
pretraining data is a commodity slop game, post-training is value add
but no one’s doing it yet 🤡
Introducing Ineffable Intelligence. Led by David Silver, we're assembling the best engineers and researchers in the world to make first contact with superintelligence. We’ll be solving the hardest problems in AI on the way. Come join us.
https://t.co/zUuvPJGmcq
If anyone’s creating a benchmark for frontier physical AI models, this task is a great one to add to the roster. Sensorimotor end-to-end policies must exhibit the long-term visual memory to track and reason about where the object might be.
It’s also harder to “cheat” on this task -- it can be difficult to do if you’ve got “gaps” in your model’s memory e.g. low-frame rate memory or coarse representations, like language.
GEN-1 nails it (also on the first try with an unseen object). @BerkayAntmen was really trying hard to fool the model here.
Shout out to an excellent task from @RhodaAI.
Constraints are the catalyst of invention. An infinite search space leads to paralysis. The most creative inventions happen when you are forced to solve a problem within appropriately narrow constraints.
The recipe behind today’s frontier reasoning models is surprisingly similar to AlphaGo:
1) Imitate large amounts of human data
2) Scale inference compute to reason better (back then it was Monte Carlo Tree Search, today it's Chain of Thought)
3) Use RL to go beyond imitation
More pretraining improves GEN-0 real-robot performance (via blind A/B evals with closed-loop rollouts).
Improvements are significant in the low-data regime, but the best models thrive with both pretraining and ample post-training.
See blog addendum: https://t.co/LVBdzMxn0f
⏳ The countdown is on and we can hardly wait for @formnext_expo to begin.
Come & meet us at Booth 12.1 – G34 from Nov 15-18th – we have great news to share🤩
Last but not least, in joint work with @nmourdou_, @mfederici_, @gpantalos1, and @markvanderwilk, we show that you can learn invariances in BNNs by just encoding different ones in the prior and letting the posterior figure it out by itself: https://t.co/WbARmHll1K
Scientists at ETH Zurich have leveraged big data from #recruitment platforms and machine learning to study hiring #discrimination. @KOFETH_en https://t.co/7QJaERO3ip
CausalWorld is an open-source simulation framework and benchmark for causal structure and transfer learning in a robotic manipulation environment where tasks range from rather simple to extremely hard. Work done at @MPI_IS and @MILAMontreal. (1/9)
https://t.co/86G08g7izy