Latent reasoning and temporary memory can bring inference down to $0.0007 per task!
BDH-CQ is a 150M model that shows it in practice.
It combines learning from examples + internal reasoning in one system:
1. It reasons inside its hidden numerical representation, or latent space, instead of writing out every step. It repeatedly processes its internal state before giving an answer.
2. It also has temporary memory. You show BDH-CQ a few examples of a new puzzle, and it learns the rule without retraining.
The researchers tested it on ARC-AGI-1, a benchmark of visual reasoning puzzles.
Here is what is mostly interesting about BDH-CQ's performance:
- BDH-CQ is surprisingly small – 150M parameters
- It’s extremely cheap to run: an inference cost is just $0.0007 per ARC task – less than one-tenth of a cent.
- Despite the low cost, it achieved 29.5% pass@2 on ARC-AGI-1, solving ~3/10 benchmark tasks.
What do we see? Language may not need to be the main workspace for AI reasoning.
BDH-CQ points to a different scaling path that works for small models: combining internal reasoning + learning from demonstrations
In this new whitepaper with @NVIDIARobotics we highlight the impact of @nvidia GPU-accelerated computing on VSLAM stacks for industrial robotics use cases:
https://t.co/M6eGSa9bIz
Totally agree. Robotics won’t scale on hardware alone. Data will. However AMRs and Humanoids are built to navigate through production lines and logistics areas. Given the confidential nature of such industrial data, Digital Twins and simulation are becoming the foundation of the coming robotics revolution.
Edge AI for robotics, doesn’t need bigger models. We need better understanding. Humans don’t learn by brute force, they learn by extracting meaning. Engineer the right loss functions, and even small models can think differently. We’re cooking something to be shared soon🤖
Biological brains inspired deep learning, now their efficiency is pushing AI to pivot toward more sustainable designs. Nature shows intelligence can thrive with minimal data, energy & parameters. After all, the human brain runs on just ~20W.
DGX Spark from @NVIDIA just landed at our offices! We’ve been dabbling with VLM’s for a while now and preparing something special for robotics ecosystems.
8 years ago deploying computer vision models in production was a nice to have. Today AI is becoming part of our infrastructure.
@nvidia DGX Spark just pushed us forward towards that direction.
@NVIDIAAI@NVIDIARobotics
Just released: New AI Climate Simulator that you can play with. Visualize how geoengineering can slow global warming.
There is no longer any path to limiting warming to 1.5 degrees Celsius (Paris Agreement), unless we use geoengineering. Reflecting 1% of sunlight away from earth would lead to an extra ~1 degree of cooling.
Our simulator lets you explore how geoengineering via Stratospheric Aerosol Injection (SAI) gives us new paths to keep warming to 1.5 degrees. I think SAI is a promising technology worth serious exploration. Check out the simulator here: https://t.co/OxtaQMyDuL
Big thanks to collaborators @jeremy_irvin16, Jake Dexheimer, @dakotagruener, Charlotte DeWald, @DanVisioni, @DWatsonParris, @DougMacMartin, Joshua Elliott, Juerg Luterbacher, Kion Yaghoobzadeh
It's sometimes tricky to manage a large project (such as your thesis or book)! We've compiled some tips on how to keep your document organized. https://t.co/cisTj2Dmkl