This Fall at CMU we're teaching a new course on AI Agents!
The goal is that you learn how to create a scaffold, build evals, and train an agentic LLM using RL.
We'll try to balance theory and practice, and introduce modern frameworks and best practices.
physics, philosophy, biology and economics
master these at a fundamental level and you will be more formidable than 99% of machine learning researchers with computer science PhDs
frontier research is about abt modeling complex systems more than anything else - not degrees
New work with @AlecRad and @DavidDuvenaud:
Have you ever dreamed of talking to someone from the past? Introducing talkie, a 13B model trained only on pre-1931 text.
Vintage models should help us to understand how LMs generalize (e.g., can we teach talkie to code?). Thread:
One of the most substantive classes with @ChaseLochmiller at Stanford. We went deep on economics of the datacenter:
- Where is the ~$650B of AI infra capex actually going this year?
- Who's capturing the margin, who's getting squeezed?
- How the bottleneck has moved from GPUs to power, and where it goes next
- The economics of neoclouds
Meet the new Stitch, your vibe design partner.
Here are 5 major upgrades to help you create, iterate and collaborate:
🎨 AI-Native Canvas
🧠 Smarter Design Agent
🎙️ Voice
⚡️ Instant Prototypes
📐 Design Systems and DESIGN.md
Rolling out now. Details and product walkthrough video in 🧵
This is the DeepSeek moment for Voice AI.
Today we’re releasing Chatterbox Turbo — our state-of-the-art MIT licensed voice model that beats ElevenLabs Turbo and Cartesia Sonic 3!
We’re finally removing the trade-offs that have held voice AI back.
Fast models sound robotic. Great models are slow. And none are built for trust. We fixed all three. Chatterbox Turbo is transparent, auditable, and built for a world that needs proof.
Excited to release new repo: nanochat!
(it's among the most unhinged I've written).
Unlike my earlier similar repo nanoGPT which only covered pretraining, nanochat is a minimal, from scratch, full-stack training/inference pipeline of a simple ChatGPT clone in a single, dependency-minimal codebase. You boot up a cloud GPU box, run a single script and in as little as 4 hours later you can talk to your own LLM in a ChatGPT-like web UI.
It weighs ~8,000 lines of imo quite clean code to:
- Train the tokenizer using a new Rust implementation
- Pretrain a Transformer LLM on FineWeb, evaluate CORE score across a number of metrics
- Midtrain on user-assistant conversations from SmolTalk, multiple choice questions, tool use.
- SFT, evaluate the chat model on world knowledge multiple choice (ARC-E/C, MMLU), math (GSM8K), code (HumanEval)
- RL the model optionally on GSM8K with "GRPO"
- Efficient inference the model in an Engine with KV cache, simple prefill/decode, tool use (Python interpreter in a lightweight sandbox), talk to it over CLI or ChatGPT-like WebUI.
- Write a single markdown report card, summarizing and gamifying the whole thing.
Even for as low as ~$100 in cost (~4 hours on an 8XH100 node), you can train a little ChatGPT clone that you can kind of talk to, and which can write stories/poems, answer simple questions. About ~12 hours surpasses GPT-2 CORE metric. As you further scale up towards ~$1000 (~41.6 hours of training), it quickly becomes a lot more coherent and can solve simple math/code problems and take multiple choice tests. E.g. a depth 30 model trained for 24 hours (this is about equal to FLOPs of GPT-3 Small 125M and 1/1000th of GPT-3) gets into 40s on MMLU and 70s on ARC-Easy, 20s on GSM8K, etc.
My goal is to get the full "strong baseline" stack into one cohesive, minimal, readable, hackable, maximally forkable repo. nanochat will be the capstone project of LLM101n (which is still being developed). I think it also has potential to grow into a research harness, or a benchmark, similar to nanoGPT before it. It is by no means finished, tuned or optimized (actually I think there's likely quite a bit of low-hanging fruit), but I think it's at a place where the overall skeleton is ok enough that it can go up on GitHub where all the parts of it can be improved.
Link to repo and a detailed walkthrough of the nanochat speedrun is in the reply.
Here's my 6 hour conversation with @dhh, a legendary programmer, creator of Ruby on Rails, author, and race car driver. This was a fun and inspiring conversation on everything from the future of programming & AI to the nature of happiness & productivity to the value of family, getting married and having kids.
X limits video length to 6 hours. So this full convo doesn't fit (by a few minutes). So, the first 6 hours are here on X. The full version is up everywhere else (see comment).
Timestamps:
0:00 - Episode highlight
1:21 - Introduction
2:32 - Programming - early days
19:57 - JavaScript
30:16 - Google Chrome and DOJ
38:03 - Ruby programming language
45:14 - Beautiful code
1:03:15 - Metaprogramming
1:06:36 - Dynamic typing
1:13:55 - Scaling
1:26:47 - Future of programming
1:44:18 - Future of AI
1:50:13 - Vibe coding
1:58:45 - Rails manifesto: Principles of a great programming language
2:23:11 - Why managers are useless
2:32:32 - Small teams
2:38:39 - Jeff Bezos
2:53:57 - Why meetings are toxic
3:01:43 - Case against retirement
3:09:00 - Hard work
3:14:38 - Why we left the cloud
3:17:48 - AWS
3:27:07 - Owning your own servers
3:33:19 - Elon Musk
3:43:01 - Apple
3:54:48 - Tim Sweeney
4:06:22 - Fatherhood
4:32:04 - Racing
4:59:08 - Cars
5:04:26 - Programming setup
5:19:35 - Programming language for beginners
5:32:53 - Open source
5:41:46 - WordPress drama
5:53:03 - Money and happiness
6:01:56 - Hope
Nice - my AI startup school talk is now up! Chapters:
0:00 Imo fair to say that software is changing quite fundamentally again. LLMs are a new kind of computer, and you program them *in English*. Hence I think they are well deserving of a major version upgrade in terms of software.
6:06 LLMs have properties of utilities, of fabs, and of operating systems => New LLM OS, fabbed by labs, and distributed like utilities (for now). Many historical analogies apply - imo we are computing circa ~1960s.
14:39 LLM psychology: LLMs = "people spirits", stochastic simulations of people, where the simulator is an autoregressive Transformer. Since they are trained on human data, they have a kind of emergent psychology, and are simultaneously superhuman in some ways, but also fallible in many others. Given this, how do we productively work with them hand in hand?
Switching gears to opportunities...
18:16 LLMs are "people spirits" => can build partially autonomous products.
29:05 LLMs are programmed in English => make software highly accessible! (yes, vibe coding)
33:36 LLMs are new primary consumer/manipulator of digital information (adding to GUIs/humans and APIs/programs) => Build for agents!
Thank you again for the invite @ycombinator and congrats again on an awesome events! I'll post some links/references in the reply.
wrote a new post, the gentle singularity.
realized it may be the last one like this i write with no AI help at all.
(proud to have written "From a relativistic perspective, the singularity happens bit by bit, and the merge happens slowly" the old-fashioned way)
@veggie_eric Something like ima.copilot, where we can create notes out of the conversation, and have a section in the app/website where we can review our notes and even interact with it.
What would truly open-source AI look like? Not just open weights, open code/data, but *open development*, where the entire research and development process is public *and* anyone can contribute. We built Marin, an open lab, to fulfill this vision:
The AI Data Center market is a $250+BN market expected to grow to >$500BN in 3 yrs & > $1T by the end of the decade.
The caveat is that not all business models will thrive—some will win big, others will falter.
Here is a framework for understanding who captures the most value: