It’s my last week at Meta this week after an amazing 6 years! I’m grateful for having had the opportunity to contribute to PyTorch both internally and for the community, as well as work on some interesting scaling related problems as part of Llama. Excited for what’s next!
Helping make the AI ecosystem more open has been the lion’s share of my career, from PyTorch to Llama @AIatMeta to the open models we’re now building @reflection_ai. Excited for an open weight future 🚀
For my first post, I’m sharing a letter @NVIDIA signed on why open models matter.
AI will transform every industry, power every company, and be built by every country.
Open models strengthen safety and cybersecurity, accelerate innovation and diffusion, and enable sovereignty.
The world needs both frontier closed models and frontier open models.
https://t.co/AUKzoQ5Ikb
2023: LLMs struggle with 4th grade word problems
2024: LLMs can do high school math
2025: LLMs get a gold medal at the IMO
Now, GPT-5.6 solves famous frontier math/stat questions. The IMO is today and 5.6 one-shotting a perfect score isn't even news.
Where will we be next year?
New blackboard lecture w @reinerpope
How do chips actually work – starting with basic logic gates, and working up to why GPUs, TPUs, FPGAs, and the human brain each look the way they do.
0:00:00 – Building a multiply-accumulate from logic gates
0:16:20 – Muxes and the cost of data movement
0:25:59 – How systolic arrays work
0:39:00 – Clock cycles and pipeline registers
0:51:40 – FPGAs vs ASICs
1:03:14 – Cache vs scratchpad
1:07:16 – Why CPU cores are much bigger than GPU cores
1:11:49 – Brains vs chips
1:15:22 – A GPU is just a bunch of tiny TPUs
Look up Dwarkesh Podcast on YouTube/Spotify/etc to watch. Enjoy!
Or go to the Presidio, jump in the ocean, get a coffee at The Mill, watch sunset at Twin Peaks, ride a bike anywhere, see live music, eat a burrito, take a grass nap in GG Park, have beer at The Page, watch the Bay Bridge lights, wander Chinatown, wander Ferry building, run across GG Bridge, walk Fort Funston, eat the best meal of your life with friends…drive any direction for 2hrs. And be deeply grateful for the heavenscape you live in.
I have gone from the most AI pilled person in the world to a straight up boomer in like a week. The slop keeps slopping and I feel people should have to take a test before they are allowed to use claude
For 15+ years, Jump Trading has partnered with @nvidia to advance accelerated computing in financial research.
Today, we’re deploying NVIDIA’s Vera Rubin NVL72 to support large-scale AI infrastructure. We build for research velocity.
Learn more: https://t.co/NohuMwqTj8
This is just ridiculously wrong.
PyTorch is software, so it only needs linear thinking and is not seminal work?
Let me tell you, my PhD was in AI systems, and I would be so thrilled if I had created PyTorch. It was published at NeurIPS (a top AI conference), has 64K citations, and has a transformative impact on the entire AI field.
Most of computer science is engineering and building software. Saying we cannot call that research or seminal work, is the real linear thinking.
If you feel like giving up, you must read this never-before-shared story of the creator of PyTorch and ex-VP at Meta, Soumith Chintala.
> from hyderabad public school, but bad at math
> goes to a "tier 2" college in India, VIT in Vellore
> rejected from all 12 universities for US masters despite 1420 on the GRE
> fuckit.jpg
> goes to the US anyway on a J-1 visa to CMU with no plan
> applies for masters (again) to 15 universities
> rejected from all except USC and with late admissions, NYU in 2010
> finds this guy called Yann LeCun (before he was famous)
> starts getting into open source
> rejected from all jobs including DeepMind
> only job is Amazon as test engineer
> his PhD mentor helps him get a job at a small startup (MuseAmi)
> rejected from DeepMind
> couldn't get H-1B because of J-1 home return issue; gets waiver through months of approval with USCIS and US State Dept
> very low on confidence
> In 2011/12 builds one of the fastest AI inference engines on phones
> rejected from DeepMind
> emailed Yann again and joins FAIR because of Torch7 open-source work
> scrapes through bootcamp at Facebook, struggling on an HBase task
> L8/L9 engineers at Facebook struggle to get ImageNet working
> figures out numerics / hyperparam issue as an L4
> first big win!
> FAIR goes well, runs 3 person torch7 team and co-creates PyTorch
> because of politics, management wants to shut down PyTorch
> cries-at-bar.jpg, literally
> eventually some people save PyTorch and it launches in 2017
> gets a EB-1 green card!
> the rest is history...
Think about that. He went to a tier 2 college. Was rejected from all Masters programs 2x. Rejected from every single job except Amazon test engineering. Rejected from DeepMind 3x. Nearly had his baby project shut down. Struggled with visa issues. After 12 years of failures (2005-17), he eventually rose to became a VP at Meta one of the most influential people in AI!
Soumith's story is one of resilience and he's living proof that no matter how down in the dumps you are, there's always hope.