Spending billions to train the "best" base model? You might be optimizing the wrong thing! 🎯
We show that controlling sharpness during mid-training leads to over 35% less forgetting after fine-tuning / quantization... even when the base model itself gets worse.
🧵 Takeaways for pretraining:
- Use SAM (Sharpness-Aware-Minimization) in the final steps (~10%)
- Try much higher learning rates (yes, even ~10× larger)
1/9
Interested in semantic parsing or generating naturally diverse datasets via LMs?
At #EMNLP2022, we present ReFill, a retrieve-and-edit framework for generating diverse training examples in a new domain by repurposing training examples from existing domains for Text-to-SQL parsing
yacv: Yet Another Compiler Visualizer - is a tool that creates parsing visualizations using manim (by @3blue1brown)
Check demo : https://t.co/kG6IjH6WR1
Docs : https://t.co/KKh66J7QPG
Go ahead and try with your own grammar and string !
A tool to build entire neural networks inside a Minecraft world by @ashutoshbsathe
Here's an MNIST classification network in action. Can also visualize block-by-block each intermediate activation calculation done by the network by moving inside the world
https://t.co/ksDtMhcMGf
Visualizing binarized neural networks by Matthieu Courbariaux et al. in @Minecraft using scarpet programming language by @gnembon_mc https://t.co/kuZ52os2fs