@dwarkesh_sp FWIW, these 'things', can already create novel stories, pictures, videos etc, I think we all can agree that solving mathematics olympiad problem is more that regurgitating facts. Best case, we may be one question away from a fundamental discovery.
Making LLMs run efficiently can feel scary, but scaling isn’t magic, it’s math! We wanted to demystify the “systems view” of LLMs and wrote a little textbook called “How To Scale Your Model” which we’re releasing today. 1/n
For a long time, we’ve been working towards a universal AI agent that can be truly helpful in everyday life. Today at #GoogleIO we showed off our latest progress towards this: Project Astra. Here’s a video of our prototype, captured in real time.
Do you want to train massive deep learning models with ease? Our 10 new tutorial notebooks of our popular UvA DL course show you how, implementing data, pipeline and tensor parallelism (and more) from scratch in JAX+Flax! 🚀🚀
Check them out here: https://t.co/7AtaB7j8l9
🧵 1/11
Cloud TPUs from @googlecloud are one of the most cost-effective ways to train and serve LLMs.
In 2.7, Ray finally will support TPUs natively -- Ray enables a more intuitive TPU developer experience, allowing you to train and serve on massive TPU pods with ease.
Learn more at Ray Summit https://t.co/nsFP1sV4Hl
What problems would you like to solve if you could access limitless computing... We are not limitless yet.. but Cloud TPU v4 pods cluster with 9 EFLOps in aggregate gives you a lot of headspace!
Incredibly excited that Sundar launched the public preview of Cloud TPU v4 Pods at I/O today, with a flythrough video of a datacenter filled with them: https://t.co/dGdUNBjzPk! This is really three separate announcements: https://t.co/LEZTscA7pJ
Incredibly excited that Sundar launched the public preview of Cloud TPU v4 Pods at I/O today, with a flythrough video of a datacenter filled with them: https://t.co/dGdUNBjzPk! This is really three separate announcements: https://t.co/LEZTscA7pJ
I wonder why this video did not receive more attention:
https://t.co/9gqum3DPGm
David Patterson (yes, THE D.P. from Patterson & Hennessy), who also happens to be one of the HW architect of the TPU, just lays bare all the details. Floorplan, arch. details & everything.
Here are the 3 paths to serve recommendations on @googlecloud
🌟 BigQuery ML
🌟 Recommendations AI
🌟 Two Tower Encoders
Read the details by @REWolfe1 @u1trons & Jordan Totten https://t.co/TExTB0QV6f
We just finished comparing Adam, Adafactor & Distributed Shampoo (thanks to @_arohan_) for dalle-mini training 🥳
TLDR: Distributed Shampoo is 🔥 and will become the new default for dalle-mini 🥑
A new architecture for computer vision based on a sparse mixture of experts unlocks the training of large models by routing different inputs to different paths, and achieves top accuracy with about half the compute. Learn more and grab the code → https://t.co/0ciUcukgtp
Dataset distillation enables #ML models to be trained using less data and compute. Today we introduce two novel dataset distillation algorithms and release their distilled datasets, which yield state-of-the-art results for image classification. https://t.co/y17QqSB6U4
Preparing text for processing by an #ML model often involves tokenization, in which text is split into smaller units (e.g., words or word segments). Today we present a new approach that speeds up the process by up to 8x compared to standard methods. https://t.co/aC6bqpnDCu