The Silver Rate Is Tight for GD!
After non-trivial interaction with ChatGPT 6, we prove the lower bound n^{-(p+C\sqrt{\log\log n/\log n})}, p=\log_2(1+\sqrt{2}), for GD with predetermined nonnegative stepsizes in smooth convex optimization. The silver exponent is optimal.
We proved SOTA (until 2026.09.04) lower bounds on GD convergence rates for smooth convex minimization. Now, the gap is O(N^-1.271) vs Ω(N^-1.450) for the non-anytime rate, and O(N^-1.119) vs Ω(N^-1.184) for the non-anytime rate.
Paper: https://t.co/P39ZUz0rvE
@yfff324 this looks promising! previously, another friend told me he could get a rate around 1.73. but as I mentioned, this hard instance cannot get the silver rate even if you push to the limit. I will share the chat history about the impossibility result here soon
We used GPT-5.6 Sol Pro to prove a new lower bound for gradient descent in smooth convex optimization.
For GD with arbitrary predetermined step sizes, we prove \Omega(T^{-1.9319}).
https://t.co/nPuWtf9q5w
A natural question is: perhaps our hard instance can be pushed further and eventually give a matching 1/T^{1.2715} silver-rate lower bound?
The answer is no.
So if the silver-rate conjecture is true, one must find fundamentally different hard instances.
Schedule-Free Learning
https://t.co/D6OUDuJ05C
We have now open sourced the algorithm behind my series of mysterious plots. Each plot was either Schedule-free SGD or Adam, no other tricks!
We don't expect Bayesian methods to do so well at large scale, but we can now get decent improvements with variational learning to GPT-2. I wrote a blog about this (first one in a long time). Check it out!
https://t.co/c7ftgBol2x
Paper: https://t.co/GUFi1br9av
A thread below.
We show that GD with small initialization can converge for unregularized matrix completion with nearly optimal computational and statistical complexities (https://t.co/LkVBFHr64p).
We believe that our tools can have a broader impact in 1. Provide global guarantees for other tasks like phase retrieval; 2. Can be used to improve the current stability-based generalization bounds.