"LeRoPE: Learnable RoPE Frequencies Improve Language Modeling"
LLMs use hand-picked RoPE frequencies to understand token positions, but this paper shows they work better when the frequencies are learned.
With just 32 extra parameters, LeRoPE beats RoPE across scales and discovers a dominant positional band around 2.2x the training context.
(1/n)
32 extra parameters → 3.4% saved pretraining compute
RoPE’s rotation frequencies matter, but they’re usually fixed at hand-picked values. In a new work with @kpetro_research, we let the network decide: LeRoPE learns the RoPE frequencies during pretraining.
Thread below!