Robot policies can move but can't think.
LLMs can think but can't move.
So we connected them.
Real robot: 16.7% → 97.3%
Sim (LIBERO-PRO): 12.8% → 53.3%
This is a neat RL trick from Thinking Machines Inkling.
Instead of optimizing only for task success, they optimize:
Reward = Task Reward − λ × (# reasoning tokens)
Then they vary λ across rollouts and pair it with different effort instructions.
The result: model learns that reasoning is a resource to spend, not something to maximize.
Augmented Reality is the natural interface for robotics.
Not only for giving commands, but also to visualise hardware data (like LiDAR) natively in 3D.
@specs AR glasses, @UnitreeRobotics Go2 Lite quadruped, @dimensionalos OS for the robot
Pre-training folks 👀
This is a super interesting observation from the MAI technical report:
Randomly initialized attention naturally behaves like uniform averaging (i.e., the attention matrix is approximately rank-1). They suggested a surprisingly simple training trick.
Feeling confused? Good. Keep reading 🧵
Inject N code bugs and measure an LLM’s ability to debug them. When fully mature, CD Bench can potentially become a super strong baseline and also an RL environment. Super exciting!
Been poking at this question - if you keep adding cascading bugs to a codebase, how fast do models fall apart?
Setup: inject N bugs (that aim to mirror real dev work) into an open source library, run tests, measure how many get fully fixed. Verification is cheap and asymmetric - we injected the bugs so we know the ground truth clean state. The model only sees corrupted code + test failures.
Ran this on scikit-learn with Qwen2.5-Coder family (3B to 32B), N from 1 to 10. Every model degrades. 32B drops from 100% at N=1 to 52% at N=10. Calling it CD(Controllable Debugging) Bench - very preliminary, one library, limited compute, feedback welcome. (1/n)
@sleenyre +1 to the TP argument. But also, if we want to optimizers like Muon to be supported, we have to leave them unfused. NanoGPT speedruns also have evidence that unfused weights have better convergences.
We trained our 17B/2A MoE Nucleus-Image with Muon. Aurora was exactly what we needed to stabilize our expert updates!
We initially went for 64B/3A model with 1 shared expert (2048 x 2048) and 256 experts (2048 x 512). Our expert output norms died down within 10k iterations while the shared expert got most of the gradients. Aurora might’ve just solved it for us.
this is fascinating, they train an encoder/decoder but use LLM matching the target model's shape for each part, so the latent space is just plain language and they can detect reward hacking, unwanted behavior and more
could even see it being used as an eval to quantify how smart a model is, i love this
We finally know why LLMs hallucinate. It's not the model. It's the geometry.
@OpenAI text-embedding-3-large: 91/3072 dimensions do real work.
@GeminiApp gemini-embedding-001: 80/3072 dimensions do real work.
~97% of your vector database is mathematically empty. Your RAG system is retrieving from noise.
@ashwingop and I present "The Geometry of Consolidation" - a proof that RAG compression has a hard floor no algorithm can beat, set by a single spectral number your embedding model cannot escape.
Every hallucination your RAG pipeline produces? This is why.
Paper + results: https://t.co/zut8pdoPbH
Meet Human Operator from MIT Media Lab: a wearable that lets AI temporarily take control of your hand using electrical muscle stimulation.
Watch it crush piano, draw perfectly, and mix cocktails like a pro — all from a simple voice command.
“I gave an AI a body.”
This isn’t sci-fi. This is tomorrow.
#HumanOperator #MITMediaLab
🚨 Nucleus-Image is now live on fal!
🎨 High-end text-to-image quality for campaigns, product shots and key art
📐 Flexible aspect ratios for social, print, and widescreen layouts
✨ Strong prompt following for concept art, storyboards and brand visuals
Nucleus-Image is now one click away.
Huge thanks and shoutout to @multimodalart and the @huggingface team - there's now an official HF Space
demo!
Just type a prompt and try it → https://t.co/mFVIRUHYMS
17B sparse MoE. 2B active. Apache 2.0.
Ostris AI Toolkit now supports training LoRAs on top of @withnucleusai Nucleus-Image. A 17B param, 2B active, MoE image model.
Currently only trains the shared expert and other non-expert layers, which works well for most LoRAs. Will try to add full MoE support soon.