🔥 Frontier labs are redesigning the residual stream to beat the curse of depth: Kimi K3 uses AttnRes, DeepSeek V4 uses mHC, ByteDance proposed HC, and there are more: LNS, KEEL, MoDA…
🙋 But each method was tested with its own training budget, model design and codebase, so the results can't be compared directly. Which one actually works best? More importantly, which one turns architectural depth into effective computational depth?
👇 We built DepthBench to find out.
📄 https://t.co/DpVhu2b44F
[10/11]
So the bigger picture is:
🚀 Depth is an underexplored and promising scaling axis, but only with the correct residual connections, and system-level design are badly needs to fulfill it in practice.
📄 Paper: https://t.co/sZASSRjpaN
💻 Code: https://t.co/QiVk8MYkcp
🤗 Models: https://t.co/Dab4GJQbB0