CMU is putting its new AI Agents course on YouTube for free.
It is 11-768, taught this Fall by Graham Neubig and Daniel Fried, and this is one of the first agent courses I’ve seen that goes all the way from building the harness to training the model behind it.
The early part covers the things you would expect: tool use, context management, skills, memory and planning.
Then it gets much more interesting.
There are separate sections on coding agents, GUI agents and deep-research agents, followed by supervised fine-tuning, reinforcement learning, RL systems, sandboxing, credential management, OpenHands, LangGraph, observability, multi-agent interaction, human-agent interaction and search.
The assignments are probably the best description of the course.
First, build an agent harness.
Then build evaluations for it.
Then train an agent with reinforcement learning.
The first assignment is already public on GitHub, including the starter code. The lectures are being uploaded to YouTube as the semester runs, and Graham Neubig has said the full set of videos will go into the public playlist.
This is considerably more useful than learning one agent framework and assuming you understand agents.
You get to see the whole problem: the model, the harness around it, the tools and memory it uses, the evals that tell you whether it works, and the training that can change the model itself.
Course:
https://t.co/dOk3OintCS
YouTube:
https://t.co/sfFYWv4S2E
The safety data for Waymo gets better and better.
Comparison w/ human drivers on rate of crashes with serious injury:
Latest data: 270M miles, 20X better
Mar 2026: 170M miles, 13X better
Before that (forget exactly when): 10X better
There are very few people I’ve learned more from than Mike Moritz. But Ausländer reveals depths of him that even those lucky enough to know him may not have seen.
It’s an extraordinarily moving exploration of family, exile, identity, and what it means to be an outsider. While my history could not be more different from Mike’s, I’ve often felt like an outsider myself, and that part of his story resonated with me a lot.
It’s deeply personal, beautifully written (his writing is legendary for a reason!), and a reminder that understanding where we come from can shed light on who we become.
I can’t recommend it highly enough.
We finished the Training Agents series. Six live sessions over six months, from evaluating agents to training them inside real environments. All of it is on the Hugging Face YouTube channel and all of the code is open.
Here's what we did and who made it happen:
1. Agentic Evaluations WorkshopWhere agent evals actually stand, and why benchmark scores don't match what people see in use. With Avijit Ghosh and Nathan Habib (Hugging Face), Arvind Narayanan (Princeton), Pierre Andrews (Meta), J.J. Allaire (UK AI Security Institute) and Mahesh Sathiamoorthy (Bespoke Labs).
2. RL for Agents Workshop Environments, rollouts, reward design and the inference bottlenecks that appear when you move from RL for LLMs to RL for agents. With Lewis Tunstall (Hugging Face), Will Brown (Prime Intellect), Ofir Press (Princeton) and Alex Zhang (MIT CSAIL).
3. Training Agents 1: SFT on agent traces Public coding-agent traces turned into prompt/completion data, a TRL + LoRA fine-tune on Hugging Face Jobs, metrics in Trackio, and an honest look at what the first eval numbers can and cannot tell you. Joined by Sergio Paniego and Quentin Gallouédec.
4. Training Agents 2: Distillation Off-policy, on-policy and self-distillation for moving capability from a teacher into a smaller coding agent.
5. Training Agents 3: Reinforcement learning GRPO after SFT: group sampling, verifiable reward functions, reading the reward/KL/length curves, and three experiments, one of them with a deliberately gameable reward so we could watch the hacking happen.
6. Training Agents 4: From reward functions to environments The reward stops being a function and becomes a place the agent acts in. We walked the reset()/step() contract from Gym to LLM agents, built an OpenEnv environment and pushed it to the Hub, plugged it into TRL's GRPOTrainer, then trained a real coding agent (OpenCode) through Harbor with AsyncGRPOTrainer on Hugging Face sandboxes.
The series has passed 300k views. Thank you to every speaker, to the TRL team, and to everyone who showed up live with questions.
Playlist: https://t.co/DYE1pvnywV
The team at @stripe is setting the standard for internal AI platforms: minion coding agents, a custom prototyping rig, and now their company brain, Kai.
On today's episode of How I AI, Sharadh shows us how 1.5 engineers and 2 weeks got them a company brain, including:
- projects as governance
- skill routing + telemetry
- a skills platform that works for 10k teammates
Plus, he and I debate the merits of gentle parenting your AI (esp when your company is running evals.)
Full episode on YT: https://t.co/2MiumBjBNH
That’s what Recursive Self Improvement (RSI) could look like for hardware.
Assuming every humanoid is better than the one previously built precision wise, and the software keeps getting better and better as well.
And scaling laws of a different type (manufacturing scale) apply for hardware for the first time. Materials become the only shortage. Labor is no longer a shortage that constrains production. Nutzo! This one video tells you so much about what could happen.
There are very few companies where the scale of their vision can be seen in a 10 second video. This might be one example.
Now think of what @elonmusk has with @SpaceX / @X / @Tesla, not even counting @neuralink:
- He has a frontier class model w/ @grok
- He has a leading app for coding w/ @cursor_ai
- He has user productivity w/ @bot
- He has the transport to space w/ @SpaceX
- He has the satellite communication w/ @Starlink
- He has the data for autonomous driving w/ @robotaxi
- He has the data for social media w/ @X
- He has the fab, and therefore custom chips w/ @TerafabCorp
- He has the terrestrial datacenter/hyperscaler
- He will have space datacenter/hyperscaler
-He has robotics with @TeslaAIBot
- He owns RSI in software and hardware
Each of these keeps getting better because of an inherent data compounding advantage.
He is stitching together one of the best vertically integrated stacks with at scale distribution of both hardware and software.
He is unapologetic to end-of-life’ing products when they reach 80% of their market potential so he can be super efficient in capital allocation.
It’s a masterclass on non-linear thinking that is still so obvious to the market from a strategy communications perspective. You don’t even have to cogently articulate the full strategy and people could put it together.
Yes, the critics would say execution is a risk. But then again, he does have a track record for that.
Say what you may about him, but this level of clarity in strategy with continued compounding value is just hard to find easily.
Sharing lecture videos for the **How to AI (Almost) Anything/Multimodal AI** course I taught at MIT in spring 2026.
This course became quite a hit the last time I shared it in spring 2025. Spring 2026's updated version contains updated topics on multimodal agents, reasoning, self-evolving AI, and new modalities like touch and smell. Also includes slides (not videos) of guest lectures on multimodal AI for health, design, manufacturing, cities, & transportation.
Youtube playlist: https://t.co/awhFUqHGBt
Course website and materials: https://t.co/m7LZaCQ4fw
Today's AI can be applied to almost anything - from language to vision, audio, sensors, medical data, music, art, smell, and taste. This course covers the principles of AI (focusing on deep learning and foundation models), how we can apply AI to novel real-world data modalities, and multimodal AI that can process many modalities at once, such as connecting language and multimedia, music and art, sensing and actuation, and more.
MIT autumn 2011.
At 84, Marvin Minsky explained a radical idea: there is no single “self” inside your head.
He argued that the mind is a collection of countless small, imperfect processes working together. Intelligence isn’t one thing it’s what emerges when these “stupid” mechanisms interact.
In his 13-lecture course, The Society of Mind, he explores consciousness, emotion, pain, creativity, and why we can literally be of two minds.
Minsky co-founded the MIT AI Lab, won the Turing Award, and advised 2001: A Space Odyssey.
The entire course is free: lectures, readings, assignments, and the book.
13 lectures that might completely change how you think about your own mind.
Recommend watching @vkhosla’s latest talk on company building. Some very useful constructs
• Hardest founder job is deciding whose judgment to trust on which topic.
• There’s a difference between a $0M company and a $0B company. It’s reflected in the mindset, first 10 hires, how you talk about yourself, whether you actually think big.
• The way to think about burn is like driving a race car. When you see a curve, you want to brake as late as possible and as quick as you can. Ability to reduce burn quickly (e.g., marketing spend, contractors vs FTE) gives flexibility in experimenting
• Obstinate on vision, flexible on tactics. Zigzag to Everest. Don’t take the fast-revenue side trail if it isn’t the mountain.
• Waste some money on purpose. Test assumptions you don’t know are assumptions. Square Cash was “waste” and became huge.
* Organized chaos > perfect management early.
• Hire for the risks in the business, not just open roles. Experts often know the previous version of the world; large innovations rarely come from people who “knew the industry.” First 10 hires hire the next 500.
• Survive long enough to give luck a chance. 3 things you control, 3 your competitors control, 4 are luck.
wrote down some of the design thinking behind Grok Bot.
persistent roles, clear state, scoped context, coordinated teams — an interface designed to move you from operating AI to delegating work.
https://t.co/37On6hzsyl
How does one RL post-train a 397B model for long-horizon knowledge work? 👩💼
We share every step we took to bring Qwen 3.5 397B from 16.1% Pass@1 to 27.3% on APEX-Agents using DPPO, including final models weights and the full training script🚀 This is the first of many works from Mercor Research on open model training research.
Full blog: https://t.co/GWgR2DB7iy
Source code: https://t.co/96h8ADfvBp