AgentEng..
@karpathy "agentic engineering..... I do think it's kind of an engineering discipline....
How do you coordinate them to go faster without sacrificing your quality bar?
And doing that well and correctly is the realm of agentic engineering."
@karpathy and I are back! At @sequoia AI Ascent 2026. And a lot has changed. Last year, he coined “vibe coding”. This year, he’s never felt more behind as a programmer.
The big shift: vibe coding raised the floor. Agentic engineering raises the ceiling.
We talk about what it means to build seriously in the agent era. Not just moving faster. Building new things, with new tools, while preserving the parts that still require human taste, judgment, and understanding.
"The biggest prize is in figuring out how you can keep ascending the layers of abstraction....
The leverage achievable via top tier "agentic engineering" feels very high right now.”
It is hard to communicate how much programming has changed due to AI in the last 2 months: not gradually and over time in the "progress as usual" way, but specifically this last December. There are a number of asterisks but imo coding agents basically didn’t work before December and basically work since - the models have significantly higher quality, long-term coherence and tenacity and they can power through large and long tasks, well past enough that it is extremely disruptive to the default programming workflow.
Just to give an example, over the weekend I was building a local video analysis dashboard for the cameras of my home so I wrote: “Here is the local IP and username/password of my DGX Spark. Log in, set up ssh keys, set up vLLM, download and bench Qwen3-VL, set up a server endpoint to inference videos, a basic web ui dashboard, test everything, set it up with systemd, record memory notes for yourself and write up a markdown report for me”. The agent went off for ~30 minutes, ran into multiple issues, researched solutions online, resolved them one by one, wrote the code, tested it, debugged it, set up the services, and came back with the report and it was just done. I didn’t touch anything. All of this could easily have been a weekend project just 3 months ago but today it’s something you kick off and forget about for 30 minutes.
As a result, programming is becoming unrecognizable. You’re not typing computer code into an editor like the way things were since computers were invented, that era is over. You're spinning up AI agents, giving them tasks *in English* and managing and reviewing their work in parallel. The biggest prize is in figuring out how you can keep ascending the layers of abstraction to set up long-running orchestrator Claws with all of the right tools, memory and instructions that productively manage multiple parallel Code instances for you. The leverage achievable via top tier "agentic engineering" feels very high right now.
It’s not perfect, it needs high-level direction, judgement, taste, oversight, iteration and hints and ideas. It works a lot better in some scenarios than others (e.g. especially for tasks that are well-specified and where you can verify/test functionality). The key is to build intuition to decompose the task just right to hand off the parts that work and help out around the edges. But imo, this is nowhere near "business as usual" time in software.
@karpathy "The biggest prize is in figuring out how you can keep ascending the layers of abstraction....
The leverage achievable via top tier "agentic engineering" feels very high right now.”
@claudeai 43/ But because intelligence is now a commodity, the nature of human work must adapt:
➡️ Directors: navigate Knightian uncertainty, transform vague intent into constraints, orchestrate agent swarms.
If a computer is generating most code, how do we make it easy for a software engineer to verify it does what they intend? What is the role of programming language design (e.g., what Rust did for memory safety)? What is the role of formal verification? What is the role of tests, CI/CD, and development workflows?
2025 will be the year of agentic systems
The pieces are falling into place: computer use, MCP, improved tool use. It's time to start thinking about building these systems.
At Anthropic, we're seeing a few best practices emerge - we wrote a blog post with our findings:
Emergent agent cooperation…especially Claude.
‘There are several drivers of the "cooperative bootstrap" that we see in Claude agents. Initial cooperation is higher, punishment of free-riders seems to be more effective, and mutation of strategies has a more cooperative bias.’
Worried that there aren't enough Multi-Agent LLM evals? Fear not!
Today in a new paper, @aronvallinder and I take a step in the right direction by studying the Cultural Evolution of Cooperation among LLM Agents.🧵
https://t.co/gWN7pXbBOC
New @AnthropicAI Computer Use feels surreal.
But don't take my word for it. We made a template on Replit for you to try.
Watch me fork the template, ask the agent to go to YouTube, find a video, and even skip the ads -- all in a few minutes.
Path to Ascencion
Dario's Vision on the AGI future
TLDR: AGI 2026 earliest, how to get there, and what to do with it. Overall, the first honest public discussion of the internal company view of what may be a good outcome
> Appreciate him putting this out there, as he will clearly be labelled crazy by a bunch of people
> For the first time, there's a plausible narrative of how we get from here to there.
> I’ve heard this multiple times in private conversations, but its great to have an official view from someone actually building it
What is AGI?
> “Strong AI” (aka super intelligence/AGI) could arrive as early as 2026
-> Given we are at the end of 2024 and an AI research to product cycle is roughly 18 months, this implies that several current research directions might actually bear fruit
Definition of Strong AI
> Country of geniuses in a datacenter:
* Smarter than a Nobel prize winner across many fields
* Has complete access to digital interfaces ie audio, video, search for both input and output. It can communicate and instruct humans
* It can autonomously plan and carry out tasks over long periods of time
* Does not have a physical embodiment but can control any robots its connected to
* The training resources can redeployed to run millions of its instances, and the model can absorb and generate information at 10x-100x human speed
* Each of these millions of instances can act independently, or can collaborate with other instances as necessary
> It’s a society of minds, working together.
> Note the accesses and capabilities the AGI has. The counterpoint to this is that an AGI that is constrained to not have interface access, limited to a box with no outputs, is not an AGI (one reason I have been going on and on about what a wonder ScarJo ChatGPT Voice is)
> Demis/Eric Schmidt at Google have been talking about agents in production next year
> Noam Brown at OpenAI just started hiring for multi-agent co-ordination research. OpenAI just released their Swarm example framework for what a group of agents working closely would look like,
> So all three labs are getting close to having autonomous single agents working, perhaps not at Nobel IQ levels, or able to co-ordinate well, but those things can be improved in parallel.
The middle ground between safety and capability
> He then discards the fast takeoff (ie within an hour) Singularity scenario due to physical bottlenecks such as hardware and need to run experiments. It will take time to get stuff done.
> Equally he discards the no-movement scenario, ie the intelligence is hamstrung by regulation and nothing happens.
> Instead he picks a middle route: an intelligence at first limited by all kinds of walls, which it works to scale and overcome
> he then speculates how the AGI will affect 5 areas of human endeavor
A. Biology
> He acknowledges the limits ie experiments that take time to run in the physical world, regulations and intrinsic complexity. Notably Yudkowsky, the prophet of AI doom, has always refused to take this view of the real constraints that any takeoff scenario will face
> Dario emphasizes that he wants to use AI not as a data analysis tool, but a principal investigator that improves very aspect of what a biologist does (planning, purchasing, management etc)
> Points out that most of progress in biology comes from big breakthroughs like CRISPR, and that we usually average 1 per year
> He expects AI to push this up 10x and maybe up to 1000x
> Does not see clinical trials as a hurdle. Trials take long because our drugs suck, and don’t usually provide clear indication of improvement. This will change if AI only produces the most efficacious drugs with improved measurement techniques of more accurate endpoints, where the FDA can stop the trial when the super efficacy of a drug means it it would be unethical not to provide it to the placebo group
> Expects advances in imaging, and AI models better able to predict the body’s reaction to drugs
> Expects 100 years of breakthroughs in 5-10 years such as
+ reliable prevention and treatment of nearly all infectious diseases
+ eliminating most cancer
+ Prevention and/or cure genetic disease, Alzheimer’s
+ biological freedom (choice of body type, gender, hair color, etc)
+ Doubling human lifespan
> Points out that this is on no policymakers roadmap, as it would increase the working age population and therefore make the whole Socials Security solvent for example.
> Most of our social infrastructure is completely unprepared to deal with this at all
B. Neuroscience
> Expects the same 100 years of breakthroughs in 5-10 years. Primarily coming from much better measurement techniques like optogenetics and neural probes, better digital models, and a similar ability to intervene both physically and behaviorally
> Expects
+ most mental illnesses to be cured
+ psychopathy and other structural conditionals maybe amenable to “brain restructuring”
+ performance management of the brain, akin to coaching
+ neurological freedom, the ability to easily experience the peaks of human experience
> mind uploading is not inside the 5-10 year window, so its a bit sci-fi for now (!)
C. Economy and Inequality
> admits he’s not an confident on this as he is on scientific advances
> does not think AI will solve economic planning and will then lead to socialism (ie the perfect central planner)
> Expects
+ health advances to be widely distributed fairly rapidly in the developing world, and this alone to lead to rapid increases in GDP 10-20% per annum
+ advances in food security due to a second green revolution basically solving for abundance
+ climate change mitigations including replacing of factory farming with energy efficient lab meat
+ greater state organizational capacity to do things with AIs for execution
D. War and Peace and Justice
> has no certainty on how AI will affect this
> wants a Western block of democracies to secure the critical bottlenecks for AI and then dole out the benefits to other nations willing to play ball
> AI could be used to reduce bias in systems by having consistency in previously imprecise criteria like “cruel and unusual punishment”
> crypto smart contracts were not smart enough to self adjudicate but AI contracts could be
> AI interpretability techniques could be used to examine and critique internal thought process of judicial decision making in order to improve it
> AI could be used to aggregate opinions and preferences amongst citizens to improve the performance of democratic governments
> improve provision of government services such as the DMV
E. Work and Meaning
> envisions a divorce of economic payments from use of labor. Want to spend a few years making a movie.. do so. Never mind that the AI can do it better, if you enjoy it, do it
> does not know what humans will get paid for if the AI can do everything
> in the short term, if AI gets good at 90% of human tasks, the remaining 10% become the human economy and expand. This is in the first 5-10 years.
> in the long term the AI will be able to do everything. His long term is a 10+ year timeline.
> envisions a societal transition but not sure what comes next, could imagine humanity living in the surface runoff of the deluge of the AI supereconomy.
In My Opinion
Probabilities
> He sounds like he’s fairly confident of an earlier than 2029 timeline. In fact I would say he has a 30% probability of a 2026 emergence.
> One should assume Dario is comfortable stating this in public because the cumulative five year probability is now in 80%> range, with 10 year probability in the 99% range (which is why he has a 5-10 year range that he provides as an afterthought)
> This is high! And soon!
> And unlike Kurzweil, who was super hand wavy and didn’t have a path to get there, Dario is actually building it.
Growth
> First time I’ve heard a company head state a GDP number, 20% per annum in developing countries, raising the rest of the world to US levels within 5-10 years.
> Paul Christiano, formerly on OpenAI's model alignment team, had a 40% Dyson Sphere probability by 2040, with an implied GDP growth rate of 640% for the next 2 decades.
> I can’t even imagine what Paul’s potential future looks like, so at least I can think about what Dario’s numbers are.
> This would definitely feel like Accelerando to us I think.
> The levels of growth in the US would be unprecedented. This country would MOVE. The further along you are on the tech curve, the more able you are to distribute new innovations. Just as 1 in 8 Americans take Ozempic or other GLP-1 drugs, just as MRNA COVID vaccines were manufactured first by American firms, every single new innovation would be fast deployed here first.
Society
> Lots of midlife and existential crisis ahead. What happens when a 70 year old feels like 35 mentally and physically? What happens when one can choose racial attributes or height or gender.
> What happens when some of us embrace the future faster than the rest?
Overall
> This make me much more optimistic about the sci-fi future in the next 5 years. We are going to make it.
> Ending off with a modified waitbutwhy. Just to let you know exactly where we are
A new essay exploring what we can expect from AGI-like systems in the 5-10 year period after they are developed. Possibly the most comprehensive breakdown of the potential benefits that I've ever read on the subject.
Written by our CEO, Dario Amodei.
https://t.co/G4x7HFOqNn
Just one example of the agents we are building at FutureHouse to automate scientific research. If you're excited about accelerating science, check out our opportunities: https://t.co/SZlw8lQagN
Super excited to announce what we have been working on in the last six months - Agent Q is out now! This is a framework for self-supervised agent reasoning and search that can self-correct and autonomously improve by self-play and RL on real tasks on the real internet! 👇
Introducing The AI Scientist: The world’s first AI system for automating scientific research and open-ended discovery!
https://t.co/jC7g5GPVsE
From ideation, writing code, running experiments and summarizing results, to writing entire papers and conducting peer-review, The AI Scientist opens a new era of AI-driven scientific research and accelerated discovery.
Here are 4 example Machine Learning research papers generated by The AI Scientist.
We published our report, The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery, and open-sourced our project!
Paper: https://t.co/lTQ8UenFHk
GitHub: https://t.co/Im53whVeAq
Our system leverages LLMs to propose and implement new research directions. Here, we first apply The AI Scientist to conduct Machine Learning research. Crucially, our system is capable of executing the entire ML research lifecycle: from inventing research ideas and experiments, writing code, to executing experiments on GPUs and gathering results. It can also write an entire scientific paper, explaining, visualizing and contextualizing the results.
Furthermore, while an LLM author writes entire research papers, another LLM reviewer critiques resulting manuscripts to provide feedback to improve the work, and also to select the most promising ideas to further develop in the next iteration cycle, leading to continual, open-ended discoveries, thus emulating the human scientific community. As a proof of concept, our system produced papers with novel contributions in ML research domains such language modeling, Diffusion and Grokking.
We (@_chris_lu_, @RobertTLange, @hardmaru) proudly collaborated with the @UniOfOxford (@j_foerst, @FLAIR_Ox) and @UBC (@cong_ml, @jeffclune) on this exciting project.
9/ For one, companies that market to developers will soon start “marketing” to coding agents as well. After all, your agent might decide what cloud you use and which database you choose. Agent-friendly UI/UX (often: a good CLI) will be prioritized.
5/ In this new world, every engineer becomes an engineering manager. You will delegate basic tasks to coding agents, and spend more time on the higher level parts of coding: understanding the requirements, architecting systems, and deciding what to build.
You should check out OpenDevin if you are learning about or building AI agents today.
The team has published now a technical report on it.
OpenDevin is a platform to develop generalist agents that interact with the world through software.
Features include:
- an interaction mechanism for interaction between agents, interfaces, and environments
- environment: sandboxed operating system + web browser available to the agents
- interface to create and execute code
- multi-agent support
- evaluation framework
--
10 implemented agents
MIT License
28K GitHub stars
160 contributors
1.3K contributions
AI Agents That Matter
abs: https://t.co/SpImx2EY8J
Performs a careful analysis of existing benchmarks, analyzing across additional axes like cost, proposes new baselines
1. AI agent evaluations must be cost-controlled
2. Jointly optimizing accuracy and cost can yield better agent design
3. Model developers and downstream developers have distinct benchmarking needs
4. Agent benchmarks enable shortcuts
5. Agent evaluations lack standardization and reproducibility