Summary of Gemini's 60-page technical report.
1. Written in Jax and trained using TPUs. The architecture, while not explained in details, seems similar to Flamigo's.
2. Gemini Pro's performance is similar to GPT-3.5 and Gemini Ultra is reported to be better than GPT-4. Nano-1 (1.8B params) and Nano-2 (3.25B params) are designed to run on-device.
3. 32K context length.
4. Very good at understanding vision and speech.
5. Coding ability: the big jump in HumanEval compared to GPT-4 (74.4% vs. 67%), if true, is awesome. However, the Natural2Code benchmark (no leakage on the Internet) shows a much smaller gap (74.9% vs. 73.9%).
6. On MMLU: using COT@32 (32 samples) to show that Gemini is better than GPT-4 seems forced. In 5-shot setting, GPT-4 is better (86.4% vs. 83.7%).
7. No information at all on the training data, other than they ensured "all data enrichment workers are paid at least a local living wage."
Full report: https://t.co/ZvgTUFPdH5
@united we’ve been waiting over 6 hours. First it was the incoming flight delay, and then crew time out with no re-assignment. Very dissatisfied and disappointed! #unitedairlinedelays#dissapointedcustomers
We're launching Keras Core, a new library that brings the Keras API to JAX and PyTorch in addition to TensorFlow.
It enables you to write cross-framework deep learning components and to benefit from the best that each framework has to offer.
Read more: https://t.co/xmmxBfSZgh
Today we are excited to announce a new partnership with @awscloud! 🔥
Together, we will accelerate the availability of open-source machine learning 🤝
Read the post 👉 https://t.co/fb77d1J2qX
!!!! Ok I recorded a (new!) 2h25m lecture on "The spelled-out intro to neural networks and backpropagation: building micrograd" https://t.co/KQ23lQW1BT .
This is the culmination of about 8 years of obsessing about the best way to explain neural nets and backprop.
After 2 years, Practical Deep Learning for Coders v5 is finally ready! 🎊
This is a from-scratch rewrite of our most popular course. It has a focus on interactive explorations, & covers @PyTorch, @huggingface, DeBERTa, ConvNeXt, @Gradio & other goodies 🧵
https://t.co/nzv7pek0iq
Would you like to make your life easier?
I certainly do! 😄
My brilliant colleagues have just published a blog post on how Merlin Models have been designed to do just that!
✅ DRY across the entire pipeline 🔥
✅ industry best practices
... and more!
https://t.co/FD7yx26diA
We are thrilled to announce Imagen, a text-to-image model with unprecedented photorealism and deep language understanding. Explore https://t.co/mSplg4FlsM and Imagen!
A large rusted ship stuck in a frozen lake. Snowy mountains and beautiful sunset in the background. #imagen
🎉 Introducing ML Course Notes 🎉
A new repo to publish and share course notes related to machine learning, NLP, and AI.
I'll be releasing all my notes there. Contributions appreciated!
https://t.co/f3G6ARdH11
New blog post!⬆️ Deep Neural Nets: 33 years ago and 33 years from now https://t.co/pbZvYh3Mck we reproduce what I think may be the earliest real-world application of a neural net trained end-to-end with backprop (LeCun et al. 1989), try improve it with time travel, and reflect.
An absolute masterclass by World's Top Data Scientists 🙏
The awesome Kaggle Grandmaster Team at NVIDIA shares their winning tips and tricks in this series:
https://t.co/o8bVaZB7Qq
1/4 Black, your friendly #Python code formatter, is no longer beta! The maintainers just released the first ever version we're confident enough to call stable. I'm not crying. YOU'RE crying! 😭
The list of changes is pretty extensive:
https://t.co/k3MwKqKiDH
🚀 In 2022, if you want to learn
🐍 Python
💻 Machine Learning / Deep Learning
📈 Data Science
Then follow your passion, stop slacking and start immediately. There are unlimited free resources available online. Choose carefully and happy learning 🙂