Llama 3 has been my focus since joining the Llama team last summer. Together, we've been tackling challenges across pre-training and human data, pre-training scaling, long context, post-training, and evaluations. It's been a rigorous yet thrilling journey:
🔹Our largest models exceed 400B parameters and are still training.
🔹Scaling is the recipe, demanding more than better scaling laws and infrastructure; e.g., managing high effective training time across 16K GPUs requires innovative strategies.
🔹Opting for an 8B over a 7B model? Among many others, an upgraded tokenizer expanded our vocabulary from 32K to 128K, making language processing more efficient and allowing our models to handle more text with fewer tokens.
🔹With over 15T tokens processed, our improved tokenization significantly enhances pre-training token efficiency. We're committed to high-quality data curation, including advanced filtering, semantic deduplication, and quality prediction.
🔹Enhanced human data significantly improves our post-training stack that combines supervised fine-tuning (SFT), rejection sampling, proximal policy optimization (PPO), and direct policy optimization (DPO).
🔹We've set the pre-training context window to 8K tokens. A comprehensive approach to data, modeling, parallelism, inference, and evaluations would be interesting. More updates on longer contexts later.
🔹While automated benchmarks are useful, they don't fully reflect a model's grasp of nuanced contexts. Our human evaluations are carefully designed and performed.
What does it take to build LLMs? Beyond data, compute, infrastructure, model, inference, safety, and evaluations--ultimately, it boils down to the synergy of dedicated talents and thoughtful investment.
Exciting updates ahead: we are planning to launch video podcasts with our developers to dive deeper into the tech behind Llama 3. We'll share the research paper. Stay tuned!
"The Construction of a Neon City" 🌆
Supercharging a static Midjourney picture with Photoshop's Generative Fill, WORKFLOW
Time spent: 3 hours 10 minutes. Complexity 6/10
Sound On 🔊
Alright Internet, need some help? During the Starship/Super Heavy launch I was able to capture this photo below, I need help to find this photo to the wall of that Rockstar of Spaceflight. This kid had such enthusiasm and passion the minute that thing broke from the dust! Unfortunately, I was about 300 feet from the group taking this photo and couldn't get to them in time to ask for contact information before then disappeared into the huge crowd. Maybe someone knows someone who was at the launch who knows someone who recognizes someone, would be a cool. Never know...
Photo: me for @SuperclusterHQ - purchase prints or downloads from the launch here: https://t.co/sCsOgQj6PW
Never thought about it before, but you know that famous picture of a bunch of construction workers sitting on a girder way up in the sky and having lunch? Well, here's the photographer who took that picture: Charles C. Ebbets.
We’re here with another industry first. Live Activities for iOS 16.1, now in the Cowboy app. See your ride in real time on your Lock Screen and Dynamic Island for a glance at your riding data in any mode you’re in.
#LiveActivities#iOS#CowboyApp#cowboybikes
Over the years I've had hundreds of pitches from companies on how to increase the revenue of my show. That's not my philosophy. I'm more interested in how to add value to my listeners, not squeeze value out of them. Very rare to see a pitch on how to give listeners more value.