I am super excited to share my new paper:
Form Follows Function: Recursive Stem Model https://t.co/InslcXQKQW
I show how a tiny RSM (~2.5M params) can beat Sudoku-Extreme and Maze-Hard:
~better accuracy (~5× reduction in error rate!) and faster training (~20x speedup!) with half the parameters compared to SOTA...
This author-edited video explains everything!
@petergostev How do you explain Claude getting better? I don't think it's learning chess, it's impossible to learn that in 200 games. I think it's mote likely it's finding the optimial opponent
@wgilpin0@wgilpin0 amazing work!
(I haven't read the paper yet) it seems to me that fractals could be the result of autoregressive aspect of the models tested. The discrete, limited capacity reasoning step/token can cause overshoot, needing a smaller correction and so on
We discovered a third pretraining axis beyond parameters and data: exploration.
Scaling exploration monotonically improves existing models across images/video/language, and unlocks end-to-end generation.
In the simplest case, it's just a for loop.
Introducing Explorative Modeling.
TLDR:
- Gains from exploration grow with scale: 7%→36% as data scales, 13%→23% as parameters scale, and gains double at 3× the compute
- Adding exploration to ~SOTA baselines improves data efficiency by 6.2×, FLOP efficiency by 4.1×, parameter efficiency by 47%, and hits a near-SOTA 1.43 unguided FID on ImageNet
- Exploration lets you trade training compute for generalization, and scales how end-to-end your generative model is
- End-to-end Explorative Models (XMs) match diffusion performance on control tasks with up to 256× less inference compute
🧵Thread:
My 2 cents: Not a fan of views that fixate on superficial differences.
• Newton connected a falling apple to the moon's orbit.
• Backprop, predictive coding, and Bayesian inference are approximations of each other.
LLMs and the brain are closer than they may appear.
4 personal thoughts about AI
1) AI/DL/LLMs have nothing in common with the brain. Deep Learning is based on backprop, the brain isn't
2) Anything with backprop is a dead end for understanding the brain. Don't let "brain-inspired" works like JEPA / HRM fool you [1/2]
I am super excited to share my new paper:
Form Follows Function: Recursive Stem Model https://t.co/InslcXQKQW
I show how a tiny RSM (~2.5M params) can beat Sudoku-Extreme and Maze-Hard:
~better accuracy (~5× reduction in error rate!) and faster training (~20x speedup!) with half the parameters compared to SOTA...
This author-edited video explains everything!
The BOLD signal may not be what we thought it was.
BOLD signal changes can oppose oxygen metabolism across the human cortex
https://t.co/3lWmJIdjCU
#neuroscience
Descartes suggested the Mind/Body relationship is the same as the Gravity/Mass relationship.
Doesn't this solve the mind/body problem?
#mindbodyproblem#emergence
"In the Sixth Replies, there is an attempt to explain the presence of the soul within the body by reference to the way in which 'gravity' or heaviness is popularly understood to be present in an object:
"Gravity, while remaining coextensive with the heavy body, can exercise all its force in any one part of the body; for if the body were hung from a rope attached to any part of it, it would still pull the rope down with all its force, just as if all the gravity existed in the part actually touching the rope instead of being scattered through the remaining parts. This is exactly the way in which I now understand the mind to be coextensive with the body - the whole mind in the whole body and the whole mind in any one of its parts (AT VII 442; CSM II 298)."
The comparison is a curious one, particularly since in his writings on physics Descartes argues that 'gravity' or heaviness is not a real quality at all (AT IXB 8; CSM I 182f). "
John Cottingham