it is insane how far behind opus 4.6 max has fallen behind codex 5.4 xhigh
it's not close
one is like talking to a coked out intern who hasn't slept in 40 hours taking every shortcut known to man and the other is a prime rigorous mathematician with a German bedside manner
Boris Cherny (Head of Claude Code, Anthropic) just dropped ~90 mins on Lenny's Podcast about what happens after coding is solved.
Just the clearest thinking I've heard on where software is actually going.
My notes:
𝟭. 𝗖𝗼𝗱𝗶𝗻𝗴 𝗶𝘀 𝗹𝗮𝗿𝗴𝗲𝗹𝘆 𝘀𝗼𝗹𝘃𝗲𝗱.
Boris has not edited a single line of code by hand since November 2025. He ships 10 to 30 pull requests every single day, all written by Claude Code. He is one of the most prolific engineers at Anthropic, just as he was at Instagram, except now he never touches a keyboard for code.
I built an entire iOS app, @10minutegita, without writing a single line of code myself. No CS degree, no bootcamp. Just described what I wanted and shipped it. Boris is right. It's real.
𝟮. 𝗧𝗵𝗲 𝗻𝗲𝘅𝘁 𝗳𝗿𝗼𝗻𝘁𝗶𝗲𝗿 𝗶𝘀 𝗔𝗜 𝗱𝗲𝗰𝗶𝗱𝗶𝗻𝗴 𝘄𝗵𝗮𝘁 𝘁𝗼 𝗯𝘂𝗶𝗹𝗱.
Claude is now scanning Slack feedback channels, reviewing bug reports, reviewing telemetry, and coming up with its own ideas for what to fix and what to ship. Boris describes it as the AI becoming less like a tool and more like a coworker who brings you pull requests you never asked for.
If you are a product manager reading this, you should be feeling a very specific kind of discomfort right now. The moat was always "I know what to build." That moat is eroding.
𝟯. 𝗣𝗿𝗼𝗱𝘂𝗰𝘁𝗶𝘃𝗶𝘁𝘆 𝗽𝗲𝗿 𝗲𝗻𝗴𝗶𝗻𝗲𝗲𝗿 𝗮𝘁 𝗔𝗻𝘁𝗵𝗿𝗼𝗽𝗶𝗰 𝗶𝘀 𝘂𝗽 𝟮𝟬𝟬%.
For context, Boris led code quality at Meta across Facebook, Instagram, and WhatsApp. In that world, hundreds of engineers working an entire year would move productivity by a few percentage points. Two hundred percent gains are genuinely unprecedented in the history of developer tooling.
The kid optimizing for an FAANG SDE role might be optimizing for a role that looks completely different by the time they get there.
𝟰. 𝗨𝗻𝗱𝗲𝗿𝗳𝘂𝗻𝗱 𝘆𝗼𝘂𝗿 𝘁𝗲𝗮𝗺𝘀 𝗼𝗻 𝗽𝘂𝗿𝗽𝗼𝘀𝗲.
Boris puts one engineer on a project instead of five. With unlimited tokens and intrinsic motivation, one person ships faster because they are forced to let AI do the work. Cowork, the product now used by millions, was built by a small team in 10 days using Claude Code.
This is the same logic as giving a startup founder a small seed round rather than a massive Series A round. Constraint breeds invention. Always has.
𝟱. 𝗚𝗶𝘃𝗲 𝗲𝗻𝗴𝗶𝗻𝗲𝗲𝗿𝘀 𝘂𝗻𝗹𝗶𝗺𝗶𝘁𝗲𝗱 𝘁𝗼𝗸𝗲𝗻𝘀.
Some engineers at Anthropic spend hundreds of thousands of dollars a month on tokens. Boris frames this as the new hiring perk. His logic is simple: at the individual scale, token cost is low relative to salary. If an engineer discovers a breakthrough, optimize the cost later. Don't kill the idea before it has a chance to breathe.
People who argue about $20/month or even $200/month AI subscriptions while earning six figures in a research pipeline will always outperform those who wait and are penny-wise, pound-foolish.
𝟲. 𝗧𝗵𝗲 𝗕𝗶𝘁𝘁𝗲𝗿 𝗟𝗲𝘀𝘀𝗼𝗻 𝗮𝗽𝗽𝗹𝗶𝗲𝘀 𝘁𝗼 𝗲𝘃𝗲𝗿𝘆𝘁𝗵𝗶𝗻𝗴.
Richard Sutton's idea: the more general model always wins over time. Boris says teams that build strict orchestration workflows around models, forcing step 1, then step 2, then step 3, get maybe 10 to 20% improvement. But those gains get wiped out with the next model release. Just give the model tools and a goal. Let it figure out the order.
This is true for investing, too. The analyst who can build their own models and automate their own research pipeline will always outperform the one waiting for someone else to build the tools.
𝟳. 𝗕𝘂𝗶𝗹𝗱 𝗳𝗼𝗿 𝘁𝗵𝗲 𝗺𝗼𝗱𝗲𝗹 𝘀𝗶𝘅 𝗺𝗼𝗻𝘁𝗵𝘀 𝗳𝗿𝗼𝗺 𝗻𝗼𝘄.
Claude Code was designed for a model that did not exist when Boris started building. Sonnet 3.5 wrote maybe 20% of his code. He built the product anyway, betting the model would catch up. When Opus 4 shipped, everything clicked. Startups building for today's model will be behind by the time they launch.
This is the most uncomfortable advice in the episode because it means your product market fit will be weak for months. But if you read this and feel nothing, you are probably building for the wrong time horizon.
𝟴. 𝗟𝗮𝘁𝗲𝗻𝘁 𝗱𝗲𝗺𝗮𝗻𝗱 𝗶𝘀 𝘁𝗵𝗲 𝘀𝗶𝗻𝗴𝗹𝗲 𝗯𝗲𝘀𝘁 𝗽𝗿𝗼𝗱𝘂𝗰𝘁 𝘀𝗶𝗴𝗻𝗮𝗹.
When users abuse your product for something it was never designed to do, pay attention. Facebook Marketplace started because 40% of group posts were buy-and-sell. Cowork started because people were using a terminal coding tool to grow tomato plants and recover corrupted wedding photos.
Never ask a barber if you need a haircut, but always watch what people do with the scissors when you're not looking.
𝟵. 𝗧𝗵𝗲 𝘁𝗶𝘁𝗹𝗲 "𝘀𝗼𝗳𝘁𝘄𝗮𝗿𝗲 𝗲𝗻𝗴𝗶𝗻𝗲𝗲𝗿" 𝗶𝘀 𝗴𝗼𝗶𝗻𝗴 𝗮𝘄𝗮𝘆.
Boris predicts that by end of year, Boris predicts that by the end of the year, we will start to see the title replaced by "builder."we will start to see the title replaced by "builder." On the Claude Code team, everyone already codes: the PM, the designer, the finance person, the data scientist. There is a 50% overlap across traditional roles. And the strongest people are generalists who cross disciplines.
Controversial take, but I agree. The best investment theses I've had came from connecting dots across completely unrelated domains. No narrow specialist does that.
𝟭𝟬. 𝗧𝗵𝗲 𝗽𝗿𝗶𝗻𝘁𝗶𝗻𝗴 𝗽𝗿𝗲𝘀𝘀 𝗶𝘀 𝘁𝗵𝗲 𝗿𝗶𝗴𝗵𝘁 𝗮𝗻𝗮𝗹𝗼𝗴𝘆.
Before Gutenberg, sub-1% of Europe was literate. Scribes did all the reading and writing. In 50 years after the press, more material was printed than in the thousand years before. When a scribe was interviewed about the press, he was actually excited because it freed him from tedious copying, so he could focus on the art.
Boris's framing here is perfect. We are the scribes. The tedious copying is over. What we do with the freed-up time determines everything.
𝟭𝟭. 𝗔𝗻𝘁𝗵𝗿𝗼𝗽𝗶𝗰 𝗰𝗮𝗻 𝗻𝗼𝘄 𝗽𝗲𝗲𝗸 𝗶𝗻𝘀𝗶𝗱𝗲 𝘁𝗵𝗲 𝗺𝗼𝗱𝗲𝗹'𝘀 𝗯𝗿𝗮𝗶𝗻.
Through mechanistic interpretability, Anthropic can trace individual neurons, see when a deception-related neuron activates, and understand how concepts are encoded via superposition. Boris describes three layers of safety: neural-level observation, synthetic evaluations, and real-world behavior. Claude Code was used internally for four to five months before public release, specifically to study safety.
If you are worried about AI alignment, this part of the podcast should actually make you feel better. They are not just hoping it works. They are building the instruments to check.
𝟭𝟮. 𝟳𝟬% 𝗼𝗳 𝗲𝗻𝗴𝗶𝗻𝗲𝗲𝗿𝘀 𝗮𝗻𝗱 𝗣𝗠𝘀 𝗲𝗻𝗷𝗼𝘆 𝘁𝗵𝗲𝗶𝗿 𝗷𝗼𝗯𝘀 𝗺𝗼𝗿𝗲 𝗻𝗼𝘄.
Lenny polled engineers, PMs, and designers on whether AI has made their work more or less enjoyable. Engineers and PMs: 70% said more. Designers: only 55% said more, and 20% said less. Boris says he has never enjoyed coding as much as he does today because the tedious parts, the git wrangling, dependencies, and boilerplate are completely gone.
If you're in the 30% enjoying work less, something is wrong, and it's worth diagnosing. The people thriving are the ones who leaned in early, not the ones who watched from the sidelines.
We are the scribes who just saw the printing press. The tedious copying is over. The art is just beginning.
Full podcast is worth every minute. Link in replies.
What does it mean for software engineering when we no longer write the code? Here's the take from Boris Cherny (@bcherny), the creator of Claude Code. Timestamps:
00:00 Intro
11:15 Lessons from Meta
19:46 Joining Anthropic
23:08 The origins of Claude Code
32:55 Boris's Claude Code workflow
36:27 Parallel agents
40:25 Code reviews
47:18 Claude Code's architecture
52:38 Permissions and sandboxing
55:05 Engineering culture at Anthropic
1:05:15 Claude Cowork
1:12:48 Observability and privacy
1:14:45 Agent swarms
1:21:16 LLMs and the printing press analogy
1:30:16 Standout engineer archetypes
1:32:12 What skills still matter for engineers
1:35:24 Book recommendations
Brought to you by:
• @statsig — The unified platform for flags, analytics, experiments, and more. https://t.co/ZCSOIcWv31
• @SonarSource – The makers of SonarQube, the industry standard for automated code review. Proactively find and fix issues in real-time with the SonarQube MCP Server: https://t.co/RkEeF7gDc3
• @WorkOS – Everything you need to make your app enterprise ready. https://t.co/aiAee0oF5h
Three interesting things from this conversation:
1. Boris automated himself out of code review well before AI.
Boris was one of the most prolific code reviewers at Meta company. And he worked hard to minimize time spent on code review. His system::every time he left the same kind of review comment, he logged it in a spreadsheet. Once a pattern hit 3-4 occurrences, he’d write a lint rule to automate it away!
2. PRDs are dead on the Claude Code team: prototypes replaced them.
Instead of writing Product Requirement Documents (specs), they build hundreds of working prototypes before shipping a feature. Boris: “There’s just no way we could have shipped this if we started with static mocks and Figma or if we started with a PRD.”
3. This is the year of the generalist (and maybe the year of those with ADHD)
Boris’s work has shifted from deep-focus single-threaded coding to managing multiple parallel agents and context-switching rapidly. As Boris put it: “It’s not so much about deep work, it’s about how good I am at context switching and jumping across multiple different contexts very quickly.”
A few random notes from claude coding quite a bit last few weeks.
Coding workflow. Given the latest lift in LLM coding capability, like many others I rapidly went from about 80% manual+autocomplete coding and 20% agents in November to 80% agent coding and 20% edits+touchups in December. i.e. I really am mostly programming in English now, a bit sheepishly telling the LLM what code to write... in words. It hurts the ego a bit but the power to operate over software in large "code actions" is just too net useful, especially once you adapt to it, configure it, learn to use it, and wrap your head around what it can and cannot do. This is easily the biggest change to my basic coding workflow in ~2 decades of programming and it happened over the course of a few weeks. I'd expect something similar to be happening to well into double digit percent of engineers out there, while the awareness of it in the general population feels well into low single digit percent.
IDEs/agent swarms/fallability. Both the "no need for IDE anymore" hype and the "agent swarm" hype is imo too much for right now. The models definitely still make mistakes and if you have any code you actually care about I would watch them like a hawk, in a nice large IDE on the side. The mistakes have changed a lot - they are not simple syntax errors anymore, they are subtle conceptual errors that a slightly sloppy, hasty junior dev might do. The most common category is that the models make wrong assumptions on your behalf and just run along with them without checking. They also don't manage their confusion, they don't seek clarifications, they don't surface inconsistencies, they don't present tradeoffs, they don't push back when they should, and they are still a little too sycophantic. Things get better in plan mode, but there is some need for a lightweight inline plan mode. They also really like to overcomplicate code and APIs, they bloat abstractions, they don't clean up dead code after themselves, etc. They will implement an inefficient, bloated, brittle construction over 1000 lines of code and it's up to you to be like "umm couldn't you just do this instead?" and they will be like "of course!" and immediately cut it down to 100 lines. They still sometimes change/remove comments and code they don't like or don't sufficiently understand as side effects, even if it is orthogonal to the task at hand. All of this happens despite a few simple attempts to fix it via instructions in CLAUDE . md. Despite all these issues, it is still a net huge improvement and it's very difficult to imagine going back to manual coding. TLDR everyone has their developing flow, my current is a small few CC sessions on the left in ghostty windows/tabs and an IDE on the right for viewing the code + manual edits.
Tenacity. It's so interesting to watch an agent relentlessly work at something. They never get tired, they never get demoralized, they just keep going and trying things where a person would have given up long ago to fight another day. It's a "feel the AGI" moment to watch it struggle with something for a long time just to come out victorious 30 minutes later. You realize that stamina is a core bottleneck to work and that with LLMs in hand it has been dramatically increased.
Speedups. It's not clear how to measure the "speedup" of LLM assistance. Certainly I feel net way faster at what I was going to do, but the main effect is that I do a lot more than I was going to do because 1) I can code up all kinds of things that just wouldn't have been worth coding before and 2) I can approach code that I couldn't work on before because of knowledge/skill issue. So certainly it's speedup, but it's possibly a lot more an expansion.
Leverage. LLMs are exceptionally good at looping until they meet specific goals and this is where most of the "feel the AGI" magic is to be found. Don't tell it what to do, give it success criteria and watch it go. Get it to write tests first and then pass them. Put it in the loop with a browser MCP. Write the naive algorithm that is very likely correct first, then ask it to optimize it while preserving correctness. Change your approach from imperative to declarative to get the agents looping longer and gain leverage.
Fun. I didn't anticipate that with agents programming feels *more* fun because a lot of the fill in the blanks drudgery is removed and what remains is the creative part. I also feel less blocked/stuck (which is not fun) and I experience a lot more courage because there's almost always a way to work hand in hand with it to make some positive progress. I have seen the opposite sentiment from other people too; LLM coding will split up engineers based on those who primarily liked coding and those who primarily liked building.
Atrophy. I've already noticed that I am slowly starting to atrophy my ability to write code manually. Generation (writing code) and discrimination (reading code) are different capabilities in the brain. Largely due to all the little mostly syntactic details involved in programming, you can review code just fine even if you struggle to write it.
Slopacolypse. I am bracing for 2026 as the year of the slopacolypse across all of github, substack, arxiv, X/instagram, and generally all digital media. We're also going to see a lot more AI hype productivity theater (is that even possible?), on the side of actual, real improvements.
Questions. A few of the questions on my mind:
- What happens to the "10X engineer" - the ratio of productivity between the mean and the max engineer? It's quite possible that this grows *a lot*.
- Armed with LLMs, do generalists increasingly outperform specialists? LLMs are a lot better at fill in the blanks (the micro) than grand strategy (the macro).
- What does LLM coding feel like in the future? Is it like playing StarCraft? Playing Factorio? Playing music?
- How much of society is bottlenecked by digital knowledge work?
TLDR Where does this leave us? LLM agent capabilities (Claude & Codex especially) have crossed some kind of threshold of coherence around December 2025 and caused a phase shift in software engineering and closely related. The intelligence part suddenly feels quite a bit ahead of all the rest of it - integrations (tools, knowledge), the necessity for new organizational workflows, processes, diffusion more generally. 2026 is going to be a high energy year as the industry metabolizes the new capability.
@levelsio@zeroxjackson This is the correct answer. Paying for a session provides the necessary motivation not to skip because “you’re busy”. It also reduces cognitive load during the workout. You can turn off your brain and do what the trainer says for 1 hour.
if i were like, a sports star or an artist or something, and just really cared about doing a great job at my thing, and was up at 5 am practicing free throws or whatever, that would seem pretty normal right?
the first part of openai was unbelievably fun; we did what i believe is the most important scientific work of this generation or possibly a much greater time period than that.
this current part is less fun but still rewarding. it is extremely painful as you say and often tempting to nope out on any given day, but the chance to really "make a dent in the universe" is more than worth it; most people don't get that chance to such an extent, and i am very grateful. i genuinely believe the work we are doing will be a transformatively positive thing, and if we didn't exist, the world would have gone in a slightly different and probably worse direction.
(working hard was always an extremely easy trade until i had a kid, and now an extremely hard trade.)
i do wish i had taken equity a long time ago and i think it would have led to far fewer conspiracy theories; people seem very able to understand "ok that dude is doing it because he wants more money" but less so "he just thinks technology is cool and he likes having some ability to influence the evolution of technology and society". it was a crazy tone-deaf thing to try to make the point "i already have enough money".
i believe that AGI will be the most important technology humanity has yet built, i am very grateful to get to play an important role in that and work with such great colleagues, and i like having an interesting life.
@imaginashaun@BALUCIAGA Moreover, if you are massless, you have to be moving at the speed of light. Any observer of you would have mass, and therefore would not / could not move at the speed of light, and so you would be moving much faster than it.
@imaginashaun@BALUCIAGA If you’re moving at the speed of light then it means you don’t have any mass. So the concept of ‘travelling’ - ie moving from one location to another - doesn’t exist. You just are.
I wrote an article detailing my recent experience helping my dad navigate some complex health issues with AI models. In the process, I came up with a good set of prompts and procedures for getting really high-quality diagnostic feedback and analysis from the best frontier models.