Karpathy just described what hiring looks like in 2026:
"Build a large project with Claude Code — like a Twitter clone. Make it secure. Have real agents using the platform doing stuff. The interviewer uses parallel agents trying to break in to verify security."
One person. Multiple agents. Shipping and defending production code simultaneously.
This is not a future job description.
This is happening right now.
The founders who get there first are not the smartest ones in the room. They are the ones who stopped doing everything themselves and built agents to do it for them.
Here is the complete playbook — 13 agents, exact prompts, 90-day build plan ↓
Read this before your competition does.
More AI agent observations below (I keep adding to the list):
1. Hermes agents write to their own memory after every task. Which means starting today versus starting in 6 months is an unfair advantage for you.
2. We're maybe 12 months from an agent that can watch you work for a week and then do your job without any instructions. The screen recording plus agent memory plus local model combination makes this possible right now
3. The real reason local models matter for founders: you can ship a product where the AI runs entirely on the customer's device and you never touch their data. Zero privacy concerns. Zero server costs. Zero compliance headaches.
That changes which industries you can sell to overnight. Healthcare, legal, finance, all the regulated verticals that won't send data to the cloud just opened up.
4. Every company needs to be rebuilt as a "second brain" before agents can be useful. That means every process, every decision, every piece of institutional knowledge has to exist in a format an agent can read. Most companies have none of this.
5. Agent costs are the new headcount. Won't be crazy for companies to spend 50%+ of their total headcount cost on tokens.
6. Agents are accidentally creating internal competition at companies. The marketing agent and the sales agent are optimizing for different metrics and working against each other without anyone realizing it. It took humans decades to develop cross-functional alignment. Nobody thought about it for agents.
7. The YAML config file is becoming the new org chart. Who reports to who, what permissions they have, what tools they access, all defined in a config file. The company's structure is literally a file you can version control, fork, and deploy. That's new.
8. The first agents that can smell a scam are going to be worth billions. Right now agents will happily wire money to a fake invoice because it matched the format. The trust layer is completely missing.
9. We're about to find out that most "expertise" was actually just memory. Knowing the tax code. Knowing the case law. Knowing which supplier charges what. When an agent holds all of that in context, the expert's value shifts from "I know things" to "I know which things matter." Much smaller group of people.
10. We're all running the same models. The differentiation is in what you feed them. Two founders with the same agent, same model, same tools will get wildly different results based purely on the quality of their knowledge base. Garbage context in, garbage output out. Forever.
11. The most underbuilt category in AI right now: agents for old people. 70 million boomers who need help with medical forms, insurance claims, and appointment scheduling.
12. Agent latency is the new page load speed. If your agent takes 45 seconds to respond, your customer already switched to one that takes
13. Skills files are the new apps. A SKILL.md that tells an agent how to do one thing well is more valuable than a SaaS subscription that does the same thing behind a login screen.
14. AI hardware... how do you create devices that are good businesses that people want? It'll be a $30 dongle you plug into existing dumb devices to give them an agent brain. Smart toaster doesn't need to be built from scratch. It needs a $30 brain attached to a $15 toaster.
15. Your agent can read faster than you can think. The bottleneck in every agent workflow is now the human approval step. We're the slow part. That's a strange thing to sit with.
16. Agents made the 80/20 rule violent. The 20% of work that matters is now the only work humans do. The 80% just disappeared. Entire job descriptions were hiding inside that 80%.
17. The thing I keep coming back to: the best businesses right now are being built by people who are just slightly ahead of their customers. Not 10 years ahead. 6 months ahead. That's the sweet spot. Far enough to lead. Close enough to be understood.
Anthropic's Claude team just showed how to build an AI agent with real memory in under 30 minutes.
24-minutes. free. by the people who built Claude.
one person + 10 agents with memory = a team that runs 24/7, remembers every customer improves itself.
worth than $500 vibe-coding course.
Ask Claude to map your entire app's architecture into a single HTML page and JSON file.
The HTML is for you. The JSON is for the next agent working on a new feature.
Your codebase now explains itself.
🚨 Anthropic just showed a 24-minute workshop on how to actually do prompts for Claude.
Taught by the people who built it.
Free. No registration. No paywall.
I've seen $300 courses that don't cover what they teach in the first 8 minutes.
Watch it and bookmark it now.
Google Cloud AI engineer just showed how they go from idea to deployed app at Google in 30-minutes using Claude.
26-minutes. free. by Google AI team.
one person + Claude + Google Cloud = a full engineering org running on a laptop.
worth more than any $500 vibe-coding course.
There will be no AI jobpocalypse.
The story that AI will lead to massive unemployment is stoking unnecessary fear. AI — like any other technology — does affect jobs, but telling overblown stories of large-scale unemployment is irresponsible and damaging. Let’s put a stop to it.
I’ve expressed skepticism about the jobpocalypse in previous posts. I’m glad to see that the popular press is now pushing back on this narrative. The image below features some recent headlines.
Software engineering is the sector most affected by AI tools, as coding agents race ahead. Yet hiring of software engineers remains strong! So while there are examples of AI taking away jobs, the trends strongly suggest the net job creation is vastly greater than the job destruction — just like earlier waves of technology. Further, despite all the exciting progress in AI, the U.S. unemployment rate remains a healthy 4.3%.
Why is the AI jobpocalypse narrative so popular? For one thing, frontier AI labs have a strong incentive to tell stories that make AI technology sound more powerful. At their most extreme, they promote science-fiction scenarios of AI “taking over” and causing human extinction. If a technology can replace many employees, surely that technology must be very valuable!
Also, a lot of SaaS software companies charge around $100-$1000 per user/year. But if an AI company can replace an employee who makes $100,000 — or make them 50% more productive — then charging even $10,000 starts to look reasonable. By anchoring not to typical SaaS prices but to salaries of employees, AI companies can charge a lot more.
Additionally, businesses have a strong incentive to talk about layoffs as if they were caused by AI. After all, talking about how they’re using AI to be far more productive with fewer staff makes them look smart. This is a better message than admitting they overhired during the pandemic when capital was abundant due to low interest rates and a massive government financial stimulus.
To be clear, I recognize that AI is causing a lot of people’s work to change. This is hard. This is stressful. (And to some, it can be fun.) I empathize with everyone affected. At the same time, this is very different from predicting a collapse of the job market.
Societies are capable of telling themselves stories for years that have little basis in reality and lead to poor society-wide decision making. For example, fears over nuclear plant safety led to under-investment in nuclear power. Fears of the “population bomb” in the 1960s led countries to implement harsh policies to reduce their populations. And worries about dietary fat led governments to promote unhealthy high-sugar diets for decades.
Now that mainstream media is openly skeptical about the jobpocalypse, I hope these stories will start to lose their teeth (much like fears of AI-driven human extinction have).
Contrary to the predictions of an AI jobpocalypse, I predict the opposite: There will be an AI jobapalooza! AI will lead to a lot more good AI engineering jobs, and I’m also optimistic about the future of the overall job market. What AI engineers do will be different from traditional software engineering, and many of these jobs will be in businesses other than traditional large employers of developers. In non-AI roles, too, the skills needed will change because of AI. That makes this a good time to encourage more people to become proficient in AI, and make sure they’re ready for the different but plentiful jobs of the future!
[Original text in The Batch newsletter.]
Andrej Karpathy: "90% of your AI coding bill is paying for context you didn't need to send"
Here are 10 things senior AI engineers stopped wasting tokens on:
1. Auto-context loading 50 files for a 30-line fix: $1.20/turn for tokens you'll never read. 80% input waste, every session
2. Running Opus on lint, format, and rename tasks: $0.60 for what Haiku nails at $0.02. 30x overpay on the cleanup tier
3. Tool call loops that re-send the full repo on every retry: 5x context cost per agentic flow. fixing these alone cuts 30-50% of bills
4. Sonnet as the default model: Kimi 2.6 matches its quality on most coding tasks at 1/6 the cost. defaulting to Sonnet in 2026 is leaving 60-70% on the table
5. Streaming responses on stable-prefix workflows: kills your prompt cache. you pay 10x for tokens that should have cost cents
6. "Just in case" file includes: 80,000-token prompts that should be 3,000. context bloat is the silent budget killer
7. Per-session knowledge rebuilding: 10 min writing a SKILL.md once vs paying agents to re-figure out your environment every run. $4 vs $0.30 per execution
8. Single-model setups: premium tier on every task is the most expensive mistake in AI coding right now
9. Asking 10 small questions one at a time: 10 separate input prefix charges vs one batched call. 70-90% savings on routine workflows
10. Buying Claude Pro + ChatGPT Plus + Cursor Pro: you seriously use one. the other two are habit, not utility
what actually compounds instead:
- context discipline (grep before fetching, always)
- prompt caching on every stable prefix
- multi-model routing (Kimi 2.6 default, Opus for the 10%)
- graduated skills via SKILL.md files
- profiling tool calls before optimizing prompts
- the routing mindset (right model for right task)
in 12 months, the gap between developers shipping on $200/month and $4,000/month budgets won't be skill
it'll be how well they route
study this.
This works really well btw, at the end of your query ask your LLM to "structure your response as HTML", then view the generated file in your browser. I've also had some success asking the LLM to present its output as slideshows, etc.
More generally, imo audio is the human-preferred input to AIs but vision (images/animations/video) is the preferred output from them. Around a ~third of our brains are a massively parallel processor dedicated to vision, it is the 10-lane superhighway of information into brain. As AI improves, I think we'll see a progression that takes advantage:
1) raw text (hard/effortful to read)
2) markdown (bold, italic, headings, tables, a bit easier on the eyes) <-- current default
3) HTML (still procedural with underlying code, but a lot more flexibility on the graphics, layout, even interactivity) <-- early but forming new good default
...4,5,6,...
n) interactive neural videos/simulations
Imo the extrapolation (though the technology doesn't exist just yet) ends in some kind of interactive videos generated directly by a diffusion neural net. Many open questions as to how exact/procedural "Software 1.0" artifacts (e.g. interactive simulations) may be woven together with neural artifacts (diffusion grids), but generally something in the direction of the recently viral https://t.co/z21CP5iQfu
There are also improvements necessary and pending at the input. Audio nor text nor video alone are not enough, e.g. I feel a need to point/gesture to things on the screen, similar to all the things you would do with a person physically next to you and your computer screen.
TLDR The input/output mind meld between humans and AIs is ongoing and there is a lot of work to do and significant progress to be made, way before jumping all the way into neuralink-esque BCIs and all that. For what's worth exploring at the current stage, hot tip try ask for HTML.
THE ENTIRE AI INDUSTRY JUST GOT HUMILIATED
a tiny model trained in just a few hours on a single graphics card is planning 48x faster than billion-dollar supercomputers.
It actually understands physics instead of just memorizing patterns.
yann lecun was right the whole time
for three years every major lab told you the same story. scale is all you need. just throw more GPUs at it. just train on more tokens. eventually the model will "wake up" and understand the world.
it was a lie. or at minimum, a very expensive bet that just lost.
LeCun kept saying generative AI is a dead end. predicting the next pixel or the next token is fundamentally wasteful, the model burns trillions of parameters memorizing surface details instead of learning how reality actually works.
he proposed JEPA instead. predict abstract concepts in a compressed thought space. don't paint the world pixel by pixel, understand it.
the problem was JEPA kept collapsing. left to its own devices the model would cheat, mapping a dog, a car, and a human to the same point in latent space. technically minimizes the loss. learns absolutely nothing.
every fix was ugly. seven loss terms. frozen encoders. EMA tricks. stop-gradients. the kind of duct-tape engineering that should have been a red flag.
then LeCun's team dropped LeWorldModel.
they replaced all the hacks with one regularizer that forces the latent space into a gaussian distribution. the model can no longer cheat. to make accurate predictions it has to actually encode physics.
15 million parameters. single GPU. trains in hours.
plans 48x faster than foundation world models.
detects physically impossible events on its own.
meanwhile OpenAI is raising another $40B to train GPT-6 on a data center the size of manhattan.
the entire scaling thesis just got embarrassed by a model that fits on a gaming PC.
Vibe Coding with Codex - Complete Guide
Build a Web App, Desktop App & iOS App
with Codex + GPT‑5.5
(No Coding Needed, Beginner Friendly)
In this video you will learn:
> Vibe Coding Basics + Vocab
> How to build a web app using Codex
> Add db, auth + storage with @Firebase
> Github Basics
> Add AI Features (API's)
> Deploy to internet (@vercel)
> Convert Web App into Desktop app & iOS App
Chapters
00:00 Intro
01:12 Setting up Codex
01:57 The basics of Vibe Coding and Codex
02:20 Projects, Files, App
03:45 Example App - Microsoft Paint
04:25 Running app locally
06:39 Save My Code - Use Github
10:37 Quick Review before building app
12:24 Building a web app - The Prompt
15:32 Creating Web App Project
16:34 Explaining Firebase (Database, Storage, Auth)
18:42 Setting up Firebase Project
22:54 Prompting Codex to build our app
24:49 Inspect Element - Console
26:24 Verify Data being Stored in Database
27:51 Making Changes to App
31:29 Fixing Storage Permissions with Codex
32:20 GPT API - Adding AI to our app
36:20 Making more changes using screenshots
39:02 Queuing vs Steering
40:23 Deploying our app to Vercel (App on Internet)
43:09 Convert web app to desktop app and iOS app
47:15 Web App and Desktop App work Now
48:36 Now let's run the iOS app
50:27 All three apps work!
51:11 Making Changes to iOS app
52:51 Testing agent skill feature of our app
53:55 Summary of what we did
Fireside chat at Sequoia Ascent 2026 from a ~week ago. Some highlights:
The first theme I tried to push on is that LLMs are about a lot more than just speeding up what existed before (e.g. coding). Three examples of new horizons:
1. menugen: an app that can be fully engulfed by LLMs, with no classical code needed: input an image, output an image and an LLM can natively do the thing.
2. install .md skills instead of install .sh scripts. Why create a complex Software 1.0 bash script for e.g. installing a piece of software if you can write the installation out in words and say "just show this to your LLM". The LLM is an advanced interpreter of English and can intelligently target installation to your setup, debug everything inline, etc.
3. LLM knowledge bases as an example of something that was *impossible* with classical code because it's computation over unstructured data (knowledge) from arbitrary sources and in arbitrary formats, including simply text articles etc.
I pushed on these because in every new paradigm change, the obvious things are always in the realm of speeding up or somehow improving what existed, but here we have examples of functionality that either suddenly perhaps shouldn't even exist (1,2), or was fundamentally not possible before (3).
The second (ongoing) theme is trying to explain the pattern of jaggedness in LLMs. How it can be true that a single artifact will simultaneously 1) coherently refactor a 100,000-line code base *and* 2) tell you to walk to the car wash to wash your car. I previously wrote about the source of this as having to do with verifiability of a domain, here I expand on this as having to also do with economics because revenue/TAM dictates what the frontier labs choose to package into training data distributions during RL. You're either in the data distribution (on the rails of the RL circuits) and flying or you're off-roading in the jungle with a machete, in relative terms. Still not 100% satisfied with this, but it's an ongoing struggle to build an accurate model of LLM capabilities if you wish to practically take advantage of their power while avoiding their pitfalls, which brings me to...
Last theme is the agent-native economy. The decomposition of products and services into sensors, actuators and logic (split up across all of 1.0/2.0/3.0 computing paradigms), how we can make information maximally legible to LLMs, some words on the quickly emerging agentic engineering and its skill set, related hiring practices, etc., possibly even hints/dreams of fully neural computing handling the vast majority of computation with some help from (classical) CPU coprocessors.
The Codex Super-App (Full Beginners Guide)
The All Purpose Interface for AI Agents
Part 1: Codex Basics
Install Codex, Projects, Chats, Documents, Plugins, Custom Skills, Automations
Part 2: Multitasking with Codex
- iOS App Designs
- Build an iOS App
- Landing Page
- Launch Video
- Investor Deck
- Social Media Automation
TIMESTAMPS:
Part 1: Codex Basics
00:00 Intro
02:54 Downloading Codex
03:20 Overview of Codex interface
03:56 Chats, Prompting, & Built in Search
04:53 Creating Projects
07:37 Creating Spreadsheet
09:43 How Files are stored and mentioned within projects
10:42 Quick Codex Overview
12:47 Search (CMD G) and Folder Organization
14:29 Skills and Plugins
16:29 Using Calendar Plugin
18:07 Creating Automations on Codex
19:18 Learn about Plugins (Figma)
21:37 Built in Image Gen
22:37 MCP Example (Paper for Design)
24:17 Opening Chats in mini-window
25:26 Steering vs Queueing the Agent
27:35 Creating Own Skill with API's
31:34 Using YouTube Researcher Skill (That we created)
33:24 Creating Automation with your custom skill
Part 2: Multitasking (More chaotic and fun)
35:27 Part 2 Multitasking: Building iOS App, Web App, Investor Deck, Launch Video, Mobile Designs, and Automated X Posts
37:54 Creating Project
38:31 Planning my 6 Projects
40:25 Mobile Design Skill
41:47 Setting up iOS App
45:08 Implementing Desings into Mobile app
46:13 Creating a landing page that collects user info
46:45 Tally for form submission (Great for lead magnets)
49:43 Organizing and Renaming Chats for multitasking
52:12 Database for Mobile App (Supabase)
53:19 Generating app icons
54:08 Launch Video (Remotion)
59:32 Remotion Video Timeline and Seeing the Video Editor
01:05:37 Editing instructions for Remotion (Gridlines)
01:07:11 Editing Web App
01:09:46 Using CLAUDE CODE Inside Codex for Design (Terminal)
01:17:20 Forking a Chat to create investor deck
01:19:09 Using Claude 4.7 Opus for Designing Deck
01:20:22 Testing Canva Export (It's good)
01:22:33 Running Mobile App on Actual Phone (Not Simulator)
01:28:58 Finishing up All Projects (Mobile App, Landing Page Launch Video)
01:31:56 Exporting Deck and making changes in Canva
01:33:13 Deploy to Vercel using the Vercel Plugin
01:33:44 Adding Song to Remotion Video
01:35:26 Setting up x Post automations (Typefully)
01:37:57 Our App is on Testflight!
01:39:58 Final Remotion Video
01:41:04 Final Thoughts, Reflections, Summary
🚨 Holy shit...A developer on GitHub just built a full development methodology for AI coding agents and it has 40.9K stars on GitHub.
It's called Superpowers, and it completely changes how your AI agent writes code.
Right now, most people fire up Claude Code or Codex and just… let it go. The agent guesses what you want, writes code before understanding the problem, skips tests, and produces spaghetti you have to babysit.
Superpowers fixes all of that.
Here's what happens when you install it:
→ Before writing a single line, the agent stops and brainstorms with you. It asks what you're actually trying to build, refines the spec through questions, and shows it to you in chunks short enough to read.
→ Once you approve the design, it creates an implementation plan so detailed that "an enthusiastic junior engineer with poor taste and no judgement" could follow it.
→ Then it launches subagent-driven development. Fresh subagents per task. Two-stage code review after each one (spec compliance, then code quality). The agent can run autonomously for hours without deviating from your plan.
→ It enforces true test-driven development. Write failing test → watch it fail → write minimal code → watch it pass → commit. It literally deletes code written before tests.
→ When tasks are done, it verifies everything, presents options (merge, PR, keep, discard), and cleans up.
The philosophy is brutal: systematic over ad-hoc. Evidence over claims. Complexity reduction. Verify before declaring success.
Works with Claude Code (plugin install), Codex, and OpenCode.
This isn't a prompt template. It's an entire operating system for how AI agents should build software.
100% Opensource. MIT License.
🚨 Anthropic's own team just showed how to actually use Claude Code properly.
30 minutes. free. the person who created Claude Code.
watch the workshop. bookmark it.
worth more than every $500 course you almost bought.
you've been using Claude without knowing 40 of its commands.
Then read the guide below.
Karpathy didn't make a course.
He made THE course.
3 hours. Free.
Tokenization. Attention. Hallucinations. Tool use. RLHF. DeepSeek. AlphaGo.
Every behavior you've ever wondered about in an LLM - where it comes from, why it exists, how it was engineered.
The gap between engineers who understand this and engineers who don't isn't technical depth.
It's the ability to conceive of entirely different things.
i built a free claude skill on the most critical content principle in 2026:
The Minto Pyramid (from a 1970s McKinsey book on clear communication)
this is the hidden formula behind almost every piece of viral content on the internet.
i use it religiously. it's how i average 10m+ impressions a month and pull about $5k mrr just from X revenue share.
the tweets you save, tiktoks you can't scroll past, essays people quote for years...
they're all built on this one McKinsey principle
the whole idea is really simple:
1. open with your conclusion in one sentence
2. then explain why you think that, with proof/examples
it works because it forces you to:
> get straight to your core point (which captures attention in 2 seconds)
> give a real opinion (which drives engagement)
> back it up with evidence, so people actually trust it
every post i write runs through this Minto claude skill (linked below) before i publish
import it and type "minto this" on any draft.
1. the skill will analyze what you’re trying to say
2. find the 1-sentence core opinion your whole piece should be built on
3. and tell you how to restructure everything else around it engineered for max virality
plus you get the actual visual pyramid structure of your post
notice something?
this post is itself a minto pyramid.
line 1 = the answer.
everything after = the proof.
you got this far without scrolling, which is pretty much the whole point :)
once you go minto, you never go back