Excited to share our new work lead by @rjerryma, where we attempt to demystify RAdam (Liu et al. 2019) and its automatic learning rate warmup schedule.
Paper: https://t.co/hDsfzzfmiv
[1/4]
Congrats @apaszke & #S4TF on releasing "Tensors Fitting Perfectly" - static analyzers are an exciting middle path b/w dependently typed and untyped tensor dimensions.
https://t.co/c01DWZi66e
@srush_nlp - this assert approach may be interesting given prior conversations :)