Slow is Fast: Is Muon the end of Shampoo?
https://t.co/9zo5vGKbcH
Orthogonal Slowest Descent — from a convex optimization convergence bound — washes away any Kronecker preconditioner's eigenvectors, leaving only eigenvalues to balance the gradient's spectrum — leads to Muon.
@alessio__serra I see what you mean. Actually, even in fp32, there are some acceleration tricks for computing exact quantiles, but they make the implementation more complex. So my main point is that histogram approximation is a trade-off between being 'good enough' and 'simple.'
@alessio__serra Good job! However, for the histogram approximation, we compared 1000 bins with 10000 bins and found that the latter brings no gains — so the 1000-bin histogram approximation is indeed sufficient, which is a comparison EQB did not perform.
The Rearrangement Inequality and Its Generalizations https://t.co/vkIrEV31kP
Starting from the rearrangement inequality, this article follows two lines of development — "matricization" and "multi-sequence extension" — to survey its principal generalizations.
Revisiting Convergence Results in Convex Optimization (X)
https://t.co/9A85wAhUv1
drop the monotonicity assumption: dynamic learning rates get average- and last-iterate convergence (why LR decay works), and non-monotone preconditioners get minimal fixes with guarantee
Revisiting Convergence Results in Convex Optimization (IX)
https://t.co/bmxSZwxbX4
Revisits AdaGrad, the seminal work on adaptive gradient algorithms, reproducing its full derivation via "convergence analysis → minimizing the upper bound → optimal preconditioning matrix."
Revisiting Convergence Results in Convex Optimization (VIII)
https://t.co/z3p7R4Q2I9
This article re-examines “schedule-free learning rates” from the perspective of multi-stage training, shifting the scheduling objective to “approaching optimality by the end of each stage” .
@tonysilveti I've read this paper too—it finds that optimal batch size scales as data size^(2/3). I haven't fully digested it yet, but I'll study it properly when I get the chance.
@cedric_chee Thanks for your interest. Whether KDA+MLA uses RoPE is purely a matter of following the experimental results, and has nothing to do with whether I'm at Kimi or not. Just as we continue to use MLA, this too is a result of respecting what the experiments tell us.