Due to a collaboration meeting, I cannot attend this year's ICML.
But we do have nice work to present!
Check out our paper to demystify how Transformers solve reasoning tasks using SFT and RL with CoT.
My student Bochen will present the poster, 8 July at 2:30pm #4609! 😃
Current Neural Scaling Laws for Transformer of LLMs rely on intensive empirical model training & curve fitting.
🚀 We present the first tractable, analytical Neural Scaling Law of linear self-attention for multitask sparse feature regression. See our #ICLR2025 paper: https://t.co/NO7lShDPYK
Just summarise several🔑 Key Contributions:
• Reformulated in-context learning dynamics as Riccati ODEs
• Analytical scaling laws for time, model size & data
• Insight into the effect of context sequence length