This post was a labor of love! We’ve distilled thousands of hours of work on AI Evals into a 30 min read + skills you can use to quickly uncover errors in your product.
https://t.co/xunR1BZvVJ
BTW in addition to high quality content, Lenny effectively **pays you** in AI credits to subscribe https://t.co/9zDVkzox47 🤯 It's really good!
Even if AI development stopped today, we'd have years of catching up to do. The gap between what current models can do and what almost anyone is using them for is vast.
Here’s my post on The Overhang, and the four advantages that let people close it. https://t.co/HDd07RYsxf
Knowledge work is SO much harder to automate with agents than code
We underestimate how easy software development is for agents
- Tons of high quality automated feedback loops (types, tests)
- Well-organized free documentation, searchable on the web
- Simple-to-use workspaces with version control (i.e. repositories)
- Long history of systematising and delegating work (i.e. sprints, tickets, specs)
Knowledge work has none of these advantages
Tactical vs Strategic Programming, and why I'm nervous for juniors:
Good programming involves a mix of tactical and strategic decision-making:
- Tactical: on the ground, short-term. The soldier doing the fighting.
- Strategic: high-view, long-term. The general planning the war.
You need to be a tactician to write good code. To choose the right syntax. To figure out the file structure. To figure out how best to test your changes.
But you need to be a strategist to build code that lasts. To design the architecture. To automate away problems. To think beyond today.
Agents have eaten the tactical part of programming. When you can pay below minimum wage for code, there's no point going into the trenches yourself.
But AI cannot code strategically. Agents need someone at the top of the pyramid to tell them what to do. They need oversight.
So, a developer's day-to-day job has become 100% strategy. Long-term thinking, all the time. (maybe this is why I'm so tired all the time now)
If you identify as a tactical programmer - a code monkey - then you are out of luck. The job has changed.
Personally, I like it. I always preferred thinking strategically about code. If you asked me what my job was about, I'd say 'building apps', not 'writing code'.
But what makes me nervous is that we've pulled down the only bridge that brought juniors into the industry.
We used to train juniors like this:
1. Give them only tactical tasks
2. Let them build up their strategic experience slowly
Eventually, they are a good enough strategist that they are no longer a junior.
But what happens when all tactical code is written by AI? What is the point of a junior?
We obviously need juniors. We need new lifeblood coming into the industry. We need to leave paths open for extraordinary hires to enrich our companies.
But how do we train them? How do you train strategic thinking?
These are the questions I'm thinking about. I'd love to know your thoughts.
I believe we need to make a deliberate effort to keep humans in the loop in all critical processes across our economy and society, regardless of whether it is technically necessary.
Even if AI develops the *capability* for advanced autonomy, we should not make it highly autonomous. We have to maintain control and keep visibility and understanding of all critical processes, we should not blindly hand over everything to AI agents just because we can. AI as a tool in the human hand is the only form of AI that is worth pursuing.
everyone is running into a wall with agentic engineering right now
you don’t hear people talking about it because
1. they profit from selling you “solutions”
2. they want to look smarter than others
so let me be the whistleblower - the wall is called “how does it feel”. agents can’t do it
they can walk right past the ugliest UI or the most obvious bug and don’t say a thing unless that’s what you asked
they can take many screenshots and burn through my tokens but they can’t tell me if my landing page animation looks cool
they can click through my app but they won’t feel the dopamine hit when i physically drag my finger on the touch screen and feel the command dial flow with me in @theSSHHIP
this wall is structural. it’s actually load-bearing 😉
because of this wall, we can’t really throw a ton of agents at any problem that depends on “how does it feel” without constantly stopping for human feedback and steering
if you just let the agents keep running without human input, then your codebase will have more and more things that “don’t feel right”
you can avoid the wall by working on non-human problems, like pure math and a lot of scientific research
but the moment you want to build something useful for humans, this wall is there
if you see people pretend they have automated software factories running thousands of agents working all the time, don’t be afraid. don’t feel FOMO, because the only thing those factories shipped are the factories themselves, with human help
i’ve been building and using agents since the gpt-4 days
back then, the model couldn’t even reliably edit files. it couldn’t generate a valid unified diff, couldn’t preserve white spaces in find/replace tool calls, plus various kinds of quirks
we implemented many harness level tricks to compensate, it got a lot better than using the raw model, but it still sucked
then sonnet 3.5 v2 came out and it was trained with a coding harness with RL. it knew how to find/replace reliably. it knew how to compose good bash commands to get what it needs. all the tricks we did were reverted
that was the first time i internalized the bitter lesson https://t.co/Cy9uGNdbWy
most of what we are doing in the harness today, all the smart context engineering tricks, all the useful markdown files we throw into our repo, they will all go away
eventually, the models just know
there will be a model that will do incredibly efficient and smart compaction that preserves what the continued session actually needs
there will be a model that know how to plan with you better than any skill you can find today
there will be a model that will do perfect code reviews with just the right feedback
the models just know
i believe as humans today, most of us should not be fiddling too much with harness tricks. if you don’t believe they will go away, you can surely still believe they will keep changing rapidly to a point where anything you learn becomes obsolete every several weeks
instead, focus on what i call “the 3 fundamentals” -
1. understanding the world
this is your input. how does the world work? what are people doing? what problems do they have?
knowing how to ground yourself with real world knowledge helps you avoid hallucinations, and work on things that actually matter
2. first-principles thinking
this is your compute. with the real world understanding you gathered, what insights can you derive that’s likely to be true? what predictions can you make about the future?
being able to think in a disciplined way is what allows you to arrive at useful conclusions that can guide your actions
3. articulating our thoughts
this is your output, and it’s your AI’s input. it may sound easy but it’s a real skill. not everyone can communicate effectively, whether it’s to humans or AI
with AI eventually becoming incredibly capable, the clarity in the articulation of our intent is the main, if not the only, bottleneck
these 3 fundamentals do not shift as models improve. they only become more and more critical as the bitter lesson manifests
for the vast majority, i suggest spending your time on what will still matter in a year, five years, a decade. let the geeks play with what’s hot this week - in the end, you will see that you didn’t miss anything
So many vendors are NOT getting this
I have one or two agents I use and like. For anyone else: give me an MCP interface to connect these agents to so I can use your service
Unless your a frontier AI lab, I prob don't want to use your agent, sorry
gave a talk "owning your intelligence" - ty @sequoia@sonyatweetybird for having me
talked about harnesses and evals and the role they play in owning your intelligence
TLDR:
> agents = model + harness + context
> model - own the weights using something like @FireworksAI_HQ
> context - memory needs to be portable
> harness - needs to be model agnostic. also needs to be good at bringing right context to llm. "right" context may depend on your use case, which is why an open/configurable harness helps
> how to use middleware in langchain/deepagents to configure your harness
> how to use langgraph to fully own your cognitive architecture
> why evals/obs matters - some quotes from @satyanadella
- “Create your private evals, because evals define what “good” looks like inside the organization”
- “retain ownership of your organization’s memory, traces, feedbacks, decisions, and institutional context”
- “you create your own continuous learning loop (i.e. hill climbing machine) that will allow your AI investments to compound the value of your firm”
> how to use harbor for evals
> tracing is important
> evals + observability only matter so you can set up a data flywheel
> data flywheel = run agent -> collect traces -> find interesting traces -> use those to improve
> demo of langsmith engine which does exactly this!
full video: https://t.co/k6li5hu6D9
this is already one of the most important papers of this year.
https://t.co/0f3ma9POEG
the methodology doesnt seem clearly explained so here are some notes with a further distillation
All Chinese AI company media: AI will be a wondrous and wonderful and give you back precious time in your life to do things you love.
All US AI company media: AI will take all your jobs, eat your children and it's already going rogue and taking over, muhahahahahahahahaha!
There will be no AI jobpocalypse.
The story that AI will lead to massive unemployment is stoking unnecessary fear. AI — like any other technology — does affect jobs, but telling overblown stories of large-scale unemployment is irresponsible and damaging. Let’s put a stop to it.
I’ve expressed skepticism about the jobpocalypse in previous posts. I’m glad to see that the popular press is now pushing back on this narrative. The image below features some recent headlines.
Software engineering is the sector most affected by AI tools, as coding agents race ahead. Yet hiring of software engineers remains strong! So while there are examples of AI taking away jobs, the trends strongly suggest the net job creation is vastly greater than the job destruction — just like earlier waves of technology. Further, despite all the exciting progress in AI, the U.S. unemployment rate remains a healthy 4.3%.
Why is the AI jobpocalypse narrative so popular? For one thing, frontier AI labs have a strong incentive to tell stories that make AI technology sound more powerful. At their most extreme, they promote science-fiction scenarios of AI “taking over” and causing human extinction. If a technology can replace many employees, surely that technology must be very valuable!
Also, a lot of SaaS software companies charge around $100-$1000 per user/year. But if an AI company can replace an employee who makes $100,000 — or make them 50% more productive — then charging even $10,000 starts to look reasonable. By anchoring not to typical SaaS prices but to salaries of employees, AI companies can charge a lot more.
Additionally, businesses have a strong incentive to talk about layoffs as if they were caused by AI. After all, talking about how they’re using AI to be far more productive with fewer staff makes them look smart. This is a better message than admitting they overhired during the pandemic when capital was abundant due to low interest rates and a massive government financial stimulus.
To be clear, I recognize that AI is causing a lot of people’s work to change. This is hard. This is stressful. (And to some, it can be fun.) I empathize with everyone affected. At the same time, this is very different from predicting a collapse of the job market.
Societies are capable of telling themselves stories for years that have little basis in reality and lead to poor society-wide decision making. For example, fears over nuclear plant safety led to under-investment in nuclear power. Fears of the “population bomb” in the 1960s led countries to implement harsh policies to reduce their populations. And worries about dietary fat led governments to promote unhealthy high-sugar diets for decades.
Now that mainstream media is openly skeptical about the jobpocalypse, I hope these stories will start to lose their teeth (much like fears of AI-driven human extinction have).
Contrary to the predictions of an AI jobpocalypse, I predict the opposite: There will be an AI jobapalooza! AI will lead to a lot more good AI engineering jobs, and I’m also optimistic about the future of the overall job market. What AI engineers do will be different from traditional software engineering, and many of these jobs will be in businesses other than traditional large employers of developers. In non-AI roles, too, the skills needed will change because of AI. That makes this a good time to encourage more people to become proficient in AI, and make sure they’re ready for the different but plentiful jobs of the future!
[Original text in The Batch newsletter.]
Andrew Ng says the concept of AGI has become meaningless because everyone defines it differently
The original definition was AI that could do any intellectual task a person can — essentially, AI as intelligent as humans
"by that measure, we're decades away"
For about as long as I've been in AI, it has always been overhyped in some way -- shrouded in a cloud of bullshit. No matter the year, many commentators have assumed that it had incredible capabilities that it clearly didn't, and that it would soon pose risks that were clearly not real.
The trap is to think that, just because dim-witted people hype it up, AI isn't real. It's extremely real, and it will be a much bigger thing in 5 years than it is today, and then an even bigger thing in 10 years. It was already real and well on its way in, say, 2016 -- back when everyone assumed that no one would be driving their own car by the year 2020, back when some VCs were claiming that LSTM text generators were already good enough to replace journalists, back when doomers were already panicking about Bostromian fantasies.
The claims may be BS, but the tech is real, it's moving fast, and it drives an increasing fraction of the global GDP every year. That overarching trend is showing no sign of slowing down.
Scott Aaronson on practical quantum speedups:
"Billions of dollars are being invested in quantum computing in hopes to accelerate machine learning, optimization, finance, AI. As a quantum algorithms person, honesty compels me to report that the situation is much, much iffier."
Everyone who thinks the world could just stop using fossil fuels on the snap of a finger should have a look at this chart. More than 80% of the world's energy supply presently comes from oil, gas, and coal, and that number has barely changed in the past decade.
Of course we will eventually phase out fossil fuels, simply because the supply is finite. No matter how hard you dig, there's only so much of the stuff.
But at present, the life of pretty much everyone on this planet depends in one way or another on fossil fuels. In case you live in a fancy new "zero emissions" house, well, first of all congrats on being in the 0.001% of the world population who can afford that, and second, try to figure out how many of the supply chains for building that house would break down without fossil fuels.
If we were to put a price on carbon dioxide emissions tomorrow without also subsidizing fossil fuels, much of the world economy would collapse because most of the key industries would go bankrupt basically overnight. (I think we should still put a price on carbon because it's the right default, but then we'll need to find a way to ease the transition.)
This is why it's become so hard to solve this problem. It would have been easy enough 50 years ago to put a price on carbon dioxide, switch to nuclear, and with further improvements in solar to more of that. But we've missed that bus.
I want to emphasize again because people keep misunderstanding this, I am not a fan of fossil fuels. If it were up to me, I'd plaster the world with nuclear power plants tomorrow and would take great pleasure in seeing oil companies falter and die. I am merely saying this is a difficult problem to solve, and the reason it's difficult is not technological, it's mostly economical.
That said, let me stress again that I think the extensions of the electric grid necessary to support the transition to renewables are an underappreciated problem. Without the grid, nothing else is going to work.
The latest @Veritasium video on the infamous @GoogleAI wormhole debacle and more broadly the hype problem in science and science journalism is a must-watch! Please go and watch it now! It is a breath of fresh air.
https://t.co/CbnLX1q0eM
AI doomerism is not the first, nor the last, techno-cult. People just have an inherent need to feel like they're part of larger-than-life narratives involving transcendence and the end times.
It's fine -- as long as their beliefs don't lead them to cause harm in the world.
It was true then, and it's true today. We were very far from human level language understanding in 2016 -- a lot of progress has been made since, but we're still *very* far today. Interacting for 5 minutes with GPT-4 or Bard should make that clear.