That's all, folks. Will leave ya'll with some nice feedback I got today: "Thanks for streaming! Would definitely recommend this. This is the real deal and it will give a good impression of how difficult ML really is, even for simple problems being solved by an expert!"
It's been a while awhile, going to do some Advent of Code 2023 with machine learning / neural networks in jax. Come hang out, there's a 75% chance this doesn't work https://t.co/oM5brllb0M π #adventofcode#jax#jaxline#haiku#weightsandbiases#deeplearning
New paper π₯³: Transformer inductive biases!
Transformers generalize differently from information stored in:
β£ weights - mostly "rule-based"
β£ context - mostly "exemplar-based"
This effect depends on (a) the training data (b) the size of the transformer
π§΅β¬οΈ
@TheIdOfAlan I do this all the time. When I finally realize what the issue is, I like to move my cursor all the way to one side and then take a running start to try to cross the gap. Hasn't worked yet, but one day...
ArXiv https://t.co/V3HhzAZgmA: Transformers up to 1,000 layers using a new normalization technique. Outperforms baselines on machine translation IWSLT-14 (De-En) and WMT-17 (En-De). New SOTA on multilingual benchmark with 7,482 translation directions with a 200-layer model.
@zacharylipton I had a related beef regarding this very example just yesterday: someone tweeted a paper that used "ERM" abbrev. in the title and abstract without ever saying what ERM stands for in the abstract! π
@DThompsonDev@mvortiz "A foolish consistency is the hobgoblin of little minds." - Ralph Waldo Emerson
I used to be all about this quote when I was younger.
@RodOrtJose Just finished doing this on stream literally 10 minutes ago for the Neural Arithmetic Logic Units paper...If ya'll are interested:
https://t.co/izrO712uWv