Recurrent depth is powerful. OpenAI’s Astra scales it. Yet we still don’t fully understand what happens inside recurrent-depth models, its internal working mechanism. Proud of @leo_raphael_ro on this work that studies the hierarchical structure and recurrent depths in Hierarchical Reasoning Models.
Key findings:
🔁 HRM behaves like a constraint-aware iterative refinement system, with its high- and low-level states playing task-dependent roles.
🏗️ In our Sudoku baseline comparison, recurrence matters more than hierarchy alone: single-state recurrent Transformers perform comparably to HRM.
🔬 Decodable does not necessarily mean causal. Probe-directed ablations behave similarly to random ablations.
Check the paper out ⬇️
📄 Paper: https://t.co/HVT4KvFhPJ
💻 Code: https://t.co/FrwIvN6LZl
(7/7) Overall, HRM appears to perform distributed computation across recurrent latent states.
Our results point to the need for interpretability methods designed specifically for recursive models to test what representations causally do, not only what they encode.
(1/7)🚨 #NewPaperAlert
Excited to share our new paper “Dissecting Hierarchical Reasoning Models: A Mechanistic Study.”
📄 https://t.co/Z3zPO46KvL
💻 https://t.co/hD96f9Hinl
A special thanks to Prof. Jian Kang (@jiank_uiuc) for the mentorship, collaboration and guidance.🧵
(6/7) SAE interventions have larger effects than probes but still do not reveal a dominant feature set. Top ranked SAE features are often no more impactful than random features.
This points to distributed computation across recurrent latent states.