I wrote a blog about our proposed Entropy Scheduling. This blog presents some laws of post-training and RL for LLMs, and it is closely related to my previous work in some ways. We will briefly present this in #ICML2025
https://t.co/QkEIxt7mzD
Stay tuned for our upcoming release of Entropy Scheduling, an RL training technique similar to Learning Rate Annealing that significantly optimizes final performance and reveals many properties of RL training.
[n/n] Welcome any feedback, comment, and discussion at #alphaxiv or #huggingface!
❤️❤️https://t.co/fx0jJ0IZHo
PS: the first author is looking for an industrial research position! Please contact us if you want to hire him!
🧵[1/n]
Our work, Learning Dynamics in Continual Pre-Training for Large Language Models has been accepted by #icml2025 spotlight!
Paper: https://t.co/ppx1omiPkc
See our interesting and helpful findings in continual pretraining👇
[7/n]
For open-source PT models with unknown information, our scaling law still works! We provide three simple methods (proxy dataset, fitting parameters, etc.) to make our scaling law become applicable again. We test our law on LLaMA-3.2-1B and the results are amazing!
Overall, this result is really interesting and paper is insightful. Quite surprised ive came across these concepts now.
TLDR:
you can predict with
L = L0 + A S_1 - C S_2
where S_1 is area under the lr curve, S_2 is (weighted) area above the lr curve.
https://t.co/Sil0AyXi8u
@kellerjordan0 @ZhangRuichong But larger seqlen always induces smaller loss. The seqlen in valid should keep consistent when enlarging seqlen in train. Is it right?
@kellerjordan0@Grad62304977@wen_kaiyue May you tell me about the schedule like what is T here? There should not exist any situation where annealing are is infinitely large. I believe there is a misunderstanding and I can solve your questions.