==== My recommendations today ====
Large Stepsize Gradient Descent for Logistic Loss: Non-Monotonicity of the Loss Improves Optimization Efficiency https://t.co/CEgtjmLRaf
(1/2)
@sarielhp@sushnt I think I understand now. X must contain an
element of D, otherwise you could take E=D. It also must contain at least six elements outside D, otherwise some rotation of the outer six "petals" would contain an empty petal.
Introducing the AlgoPerf: Training Algorithms Benchmark! Compete for a share of the $50,000 prize pool by submitting more effective and efficient neural network training algorithms. Learn more https://t.co/csfXkUCKbN
#Algorithms#MachineLearning#Competition
The code from the original version of "Sharpness-Aware Minimization and the Edge of Stability" had a bug. Patch: https://t.co/opwrhbbNZ8. New version of the paper: https://t.co/sRzzNZUhlt.
Sharpness Minimization Algorithms Do Not Only Minimize Sharpness To Achieve Better Generalization. (arXiv:2307.11007v1 [cs.LG]) https://t.co/gOMaDp0bNK
One of the best known open problems in combinatorics is the union-closed conjecture, which states that if you have a finite collection X of sets such that if A and B belong to X then so does the union of A and B, then at least one element of X belongs to at least half of them. 1/
S. Kale, J. Lee, C. De Sa, A. Sekhari, and K. Sridharan. From Gradient Flow on Population Loss to Learning with Stochastic Gradient Descent. https://t.co/Fxsui7JlMz.
New paper with Peter Bartlett and @obousquet called "The Dynamics of Sharpness-Aware Minimization: Bouncing Across Ravines and Drifting Towards Wide Minima": https://t.co/yM49FWZCNI.
@rvcraiu@niladrichat Here's an upstream paper for context: https://t.co/0eUWcgw7CB. Here a paper that is even further upstream: https://t.co/w29szKpDs1.