The ChatGPT API was released yesterday and it costs 90% less than expected. Here’s five methods (and resources to learn about them) that are **probably** being used to enable this price reduction… 🧵[1/6]
New paper: Equivariant MuZero
https://t.co/yR8mQdtKzH
A version of MuZero that incorporates environment symmetries by design.
Great internship work by @andreeadeac22 👏👏👏
We've made so much progress in the last 7-8 years on complex domains like text and images and basically none on simple sequences of numbers (i.e. time-series forecasting).
In fact, I think we've regressed since good information is so hard to find.
GitHub just wrote an article about how they had to write their own search engine (in Rust, for performance reasons) and a new probabilistic data structure to reduce indexing time because ES and Lucene were blockers for them but sure big data is over 😅
We've just released the first version of our Deep Learning Tuning Playbook! This is our attempt to distill our process for actually getting good results with deep learning. We emphasize hyperparameter tuning since it has been a large pain point. https://t.co/PjeJVWeOzS
How to train very deep NNs without shortcuts, but still achieve competitive results on ImageNet?
Our ICLR paper gives a simple solution derived from kernel approx theory. We hope this could enable further research into deep models.
https://t.co/8CsC54hJok
@BlancheMinerva Here's a few I like for different reasons (not exhaustive):
Careful perturbations:
https://t.co/QssBiBRe4y
https://t.co/mGhqvPBiAE
Cool ML modeling:
https://t.co/SNHHPiwI4Q
https://t.co/Na6Pmh7bhp
Data collection:
https://t.co/SFUN9vQYod
https://t.co/rjHQ8Bnkz2
GradMax: Growing Neural Networks using Gradient Information
abs: https://t.co/IyAuQffXcS
a new approach for growing neural networks that aims to maximize the gradient norm when growing