Hooked on @bot.
Have a fitness bot that is on Hevy + Whoop. Pretty amazing to tell me what I should I focus on next.
Household manager is the one though. Amazon, DoorDash, Instacart. I just talk to it and stuff shows up..
This is going to help optimize so many things at home
SkyRL now implements the Tinker API.
Now, training scripts written for Tinker can run on your own GPUs with zero code changes using SkyRL's FSDP2, Megatron, and vLLM backends.
Blog: https://t.co/GAtW81jM38
🧵
One of the greatest cheat codes in life is to never get offended. Train yourself to have a thick skin. Don't take things personally. Let others disagree with you. Being easily offended means you're easily manipulated. Want more peace? Avoid getting offended.
The AI compute software stack consists of 3 specialized layers:
🔧🔧🔧 Layer 1: Training & Inference Framework (PyTorch + vLLM)
• Runs models efficiently on GPUs
• Handles model optimization and model parallelism strategies
• Manages accelerator memory and automatic differentiation
⚡⚡⚡ Layer 2: Distributed Compute Engine (Ray)
• Schedules tasks within jobs and coordinates processes
• Ingests and moves data
• Provides workload-aware failure handling and autoscaling
🏗️🏗️🏗️ Layer 3: Container Orchestrator (Kubernetes)
• Provisions compute resources
• Schedules entire jobs
• Manages user and workload multitenancy
The same recipe: Kubernetes + Ray + PyTorch + vLLM.
I cover a few case studies in this blog, and the results speak for themselves:
• Pinterest cut dataset iteration time from 90 hours to 15 hours (6x improvement) and reduced batch inference cost by 30x.
• Uber increased LLM training throughput by 2-3x.
• Roblox sped up LLM batch inference 9x.
Each layer handles what it does best. The separation of concerns makes this stack so powerful.
https://t.co/tYQ4VD0IRL
Paper from @BytedanceTalk on training graph neural networks on very large graphs. Leverages @raydistributed's fault tolerance and elasticity.
https://t.co/jP794P0d30
We're excited to share the third installment of our series on Ray at Pinterest: Ray Batch Inference. Discover how we're harnessing the power of Ray to efficiently scale and optimize our machine learning models for batch processing. 🚀 Learn more: https://t.co/pZrsJLOh2H
“Resilience matters in success.
And greatness does not come from intelligence. Greatness comes from character, and character isn’t formed out of smart people: it's formed out of people who have suffered.”
— Jensen Huang
Why do 16k GPU jobs fail?
The Llama3 paper has many cool details -- but notably, has a huge infrastructure section that covers how we parallelize, keep things reliable, etc.
We hit an overall 90% effective-training-time.
https://t.co/5gngOZJHBO
Part 2 is here! 🥳 Learn about ‘Ray Infrastructure at Pinterest’ written by Chia-Wei Chen, Raymond Lee, Alex Wang, Saurabh Vishwas Joshi, Karthik Anantha Padmanabhan and Se Won Jang. 📌 https://t.co/2w183KtH57