HEAVEN OR HELL!
Mercoledì, 8 Luglio, si torna indietro nel tempo per festeggiare il Carnevale di Mezzanotte... ma alle 21.00 👀
L'inferno attende i giocatori italiani di Guilty Gear XX Λ Core Plus R per una serata di Exhibition Matches sul nostro canale Twitch! 🔥
LET'S ROCK!
Happy to finally release Special Demonstrations #21: Ikaruga to everyone! https://t.co/zWY45VJ0fj
@MaZ67086804 finally hit his target of 36mil ALL on Arcade/Normal earlier this year, and we've spent the past few months putting together a replay to do it justice. Enjoy!
2025 has begun, let's celebrate it together with the first Session of the year, Session 25:1
Same place, same times and always ready to have fun and roll with the beautiful game!
!!!Registrations already open from now!!!
Rebirth date: 25/01/'25 & 26/01/'25
#3rdstrike#sf33rd
About the YouTube channel
a-cho will close and the staff responsible for managing the channel will no longer have any authority over it.
In addition, the rights to manage video material that belongs to the manufacturer cannot be transferred to a third party.Please understand
A student reached out asking for advice on research directions in optimization, so I wrote a long response with pointers to interesting papers. I thought it'd be worth sharing it here too:
1. Adaptive optimization.
There has been a lot going on in the last year, below are some papers I personally found interesting.
First of all, this paper by Li and Lan on Nesterov's acceleration of adaptive gradient descent:
https://t.co/D6hykeK2tw
Check Corollary 1 for a simple description of their method. There is one thing I don't like about it: the amount by which we can increase the stepsize at each iteration decreases as t grows. That being said, I don't know if this restriction can be lifted, and perhaps it's the best thing we can get.
Yura Malitsky and I also did some work on adaptive gradient descent, making the stepsizes a bit larger, roughly sqrt(2) improvement over our previous result:
https://t.co/exhFgbjChk
We still don't know if that's the best we can do or if a tighter analysis can give us better methods.
I should also mention that there is more push in the literature on Polyak stepsize, see for instance these two papers:
https://t.co/8tKRReEpx2 (a stepsize very similar to Polyak)
https://t.co/kZhWGqI1sE (Polyak stepsize with momentum)
2. Adagrad-like methods still can be studied, I believe it's an underexplored direction. I wish there was more papers on studying the importance of coordinate-wise stepsizes. One paper on the topic I really liked is this study of when Adam is more useful than SGD:
https://t.co/sF5Abi08h5
There is also some research on new practical methods, for instance, acceleration of DoG is interesting:
https://t.co/VMOdfbL95Z
And I also enjoyed reading this paper by Rodomanov et al. on line-search-inspired stochastic methods:
https://t.co/85uLHGZErQ
3. I also like the direction of getting better assumptions for optimization theory and studying the implications. A good example is the gradient clipping literature:
https://t.co/NMTXzFJScs ((L₀, L₁)-smoothness)
https://t.co/dG8xIoTFPN (same revisited)
https://t.co/goKclD80WG (on heavy-tailed noise)
We need to bridge optimization assumptions with what we know about neural networks, so read about properties of neural networks themselves like this:
https://t.co/6M8l1avBOJ (on scales of layers and how their type affects Lipschitz constants)
4. These days, people are using deep networks of all scales for their tasks, and they have discovered a lot of tricks that haven't been studied thoroughly in optimization literature: quantization, Straight-Through Estimator, (https://t.co/7UK2gsojhm), low-rank techniques such as LoRA, learning-rate warm-up, etc. You should expose yourself to those tricks to get a better understanding of what the current theory is lacking.
If you're considering choosing optimization as the topic for your PhD, here are some extra thoughts. Right now there is less activity than about 5 years ago, most low-hanging fruits seem to have been taken, and the remaining questions seem quite challenging. So if you're looking for a field where it is easy to get publications, it might not be perfect. However, it's still a good field to produce meaningful theory. It's also important who you would work with, i.e. if you can find a good advisor, that often affects one's satisfaction to a larger degree than the topic itself, so make your decision carefully.
As my last word of advice, I definitely encourage testing new methods on neural networks (and preferably not on CIFAR10/CIFAR100, because they give misleading results), at least something like nanoGPT (https://t.co/NTk9KAAqd4). When I was a PhD student, I did a lot of theoretical research testing my methods on logistic regression and that was useful to understand the theory, but I also had the wrong impression about what works and what doesn't because of that. If you can, do both, understand the theory as much as you can, but also learn its limits and failure modes.