I had a great time sitting down with 동아일보 (Korea's #2 newspaper) to discuss AI, Playco, Silicon Valley, and how I don't use AI when writing :)
https://t.co/Pxu2uI6Yqe
the most underrated hire right now is a great product person.
when i say product person i'm def not talking about a product manager. perhaps i think there has to be somewhat of a new role. i don't have a good name for it yet but maybe something like "product thinker".. someone with an intuitive grasp of the product as it exists, where it's soft, where it sings, & how to iterate it toward something even sharper. in some sense, this person has to cohesively hold in their head where this product should be 2 years from now & work backwards from that.
i say this cuz when building was hard, engineering was the bottleneck & the status hierarchy often reflected that. building is no longer hard. which means the variance in outcomes has shifted almost entirely to judgment on what to build, how to sequence it, & how to talk about it.
& the story matters as much as the thing. internally, it organizes the team around a shared model of why. externally, it shapes the interpretive frame users bring to their first experience. you can't retrofit narrative onto a product & expect it to land, it has to be load bearing from the start.
the rarest version of this person sits at the intersection of culture & deep technology. someone genuinely bilingual. they know what's technically possible & they know which cultural currents are real vs. ephemeral. that combo is what separates products that feel inevitable from products that feel assembled.
before ppl clap back with this person has always been valuable, i know.. i am just saying now they might be the most *important* person in the room. their value compounds like never before.
We're killing the DM at @Zapier. Starting with the executive team.
We've long held Default to Transparency as a value. That value has largely encouraged communication in public channels. But as the company grew, DMs are a hard habit to resist and break.
But every DM is a gap in our Shared Brain. It's context that is lost for humans and AIs. As a result the cost of DMs keeps going up.
So earlier this year I posted about our exec transparency leaderboard. The leaderboard has become quite the competition internally…
I'm 3rd today. My co-founder @BryanHelmig has held the top spot as long as I can remember…
It sets a standard for the rest of the company. In fact, since last year we’ve seen the % of Slack messages in public channels go from 33% to 46%.
What the leaderboard measures
Transparency is a team sport, and a disinfectant. Every month we track what percentage of our execs' Slack messages happen in public channels versus private DMs.
When your CEO debates strategy in a DM, that decision is invisible to every agent and every team that needs to know what was decided and why. The decision happens but the reasoning vanishes.
When that conversation happens in a channel, it stays. New hires can search it, agents can read and verify it, etc. Your Shared Brain knows what's true now: ask it a question and the answer reflects the latest reality.
Taking It to the Next Level
Reducing DMs are one way to increase transparency and open up context for humans and AI, but there are other mechanisms that help too. Three things beyond the leaderboard:
1. Meetings get recorded, transcribed, and become queryable
2. We run a shared skills library. Anyone on the team can encode a workflow they've figured out into a skill and share with the team
3. And we keep score. It's a silly scoreboard, but it subtly drives positive behaviors
Raising Your Ambition
In order to get the most of AI in your company, the AIs need context. So making your context queryable is one of the most practical moves you can make to improve the effectiveness of your AI agents.
P.S. I’m coming for #1, Bryan...
SoftBank’s investor presentation is one of the greatest things ever made. I’ve been thinking about it all day. These are the real slides shown in a speech where Masayoshi Son said he wouldn’t retire for at least another decade. The goose stuff is perfect.
https://t.co/sk9cDhdWIE
More PixiJS 3D testing 👀
Take any live 2D PixiJS content (sprites, text, graphics) and render it onto 3D geometry
Our goal is to make working between 2D and 3D as seamless as possible in our WebGPU/WebGL renderer
Let us know what features you'd like to see!
Sneak peek at our latest test for the upcoming PixiJS 3D engine 👀
WebGPU first with WebGL fallback
Optimized for performance from day one
Really excited to get this into your hands... More to come soon!
Models by @quaternius and @KenneyNL 🙏
cold open: google campus. a conference room named “moonshot serenity 4b.” twelve people are in a meeting titled: pre-sync for sync alignment on ai velocity.
sundar sits calmly at the head of the table.
a pm clicks to slide 1 of 187.
“the agenda today is simple,” she says. “how do we move faster while preserving our culture of not doing that?”
everyone nods.
then the door opens.
noam shazeer walks in.
the room goes silent.
noam: “i’m leaving.”
a vp of gemini reliability, brand, trust, latency, policy, and vibe raises a hand.
“leaving… this meeting?”
noam: “google.”
someone gasps. someone else opens a doc titled retention narrative draft final final noam v7.
sundar blinks once.
“noam, we brought you back.”
“for two point seven billion dollars.”
“technically you licensed some technology and reacquired talent.”
“that sentence is why we need legal in the room.”
legal is already there.
cut to: openai.
sam altman stands beside a whiteboard that just says ship.
an engineer walks by carrying a server rack and what appears to be the future.
sam: “we can offer speed, compute, and one meeting.”
noam: “one meeting per week?”
sam: “no. one meeting. total.”
back at google, the emergency retention committee forms instantly. it has 31 members.
a director says, “what if we give him a new title?”
“he already co-leads gemini.”
“distinguished super co-lead?”
“google fellow?”
“he already left google, founded a company, got brought back for billions, then left again. he’s folklore.”
meanwhile, a gemini launch review begins.
pm: “we’re ready to announce the model.”
policy: “can it answer questions?”
eng: “yes.”
policy: “too risky.”
marketing: “can we call it experimental?”
research: “the model is better than the last one.”
brand: “better is aggressive.”
trust & safety: “what about ‘more contextually adjacent to usefulness’?”
a staff engineer whispers, “openai just shipped a model while we were discussing the adjective.”
cut to noam’s exit interview.
hr: “what could google have done better?”
flashback montage:
a chatbot blocked because it might be too good.
a launch delayed because a button was the wrong shade of responsible blue.
a spreadsheet comparing twelve ai product names.
a meeting where someone says “we need a single coherent ai strategy” and three new strategies are created before lunch.
noam: “nothing comes to mind.”
hr: “great. we’ll mark that as positive attrition.”
later, sundar calls him privately.
“google is still google. best researchers. best infrastructure. billions of users.”
“yes.”
“so why leave?”
noam looks out the window.
“because you have everything except permission.”
silence.
sundar, softly: “we can create a permission working group.”
cut to all-hands.
sundar addresses the company.
“noam is leaving. this is not a loss. it is an opportunity to reflect on our operating model.”
chat explodes:
“is this recorded?”
“which gemini?”
“can we ask gemini why people keep leaving?”
“it said ‘insufficient context.’”
a vp steps up.
“to honor noam’s legacy, we’re launching project attention.”
applause.
“it will study whether attention is, in fact, all we need.”
a researcher raises a hand. “didn’t we answer that in 2017?”
“yes. but now we need enterprise readiness.”
final scene: noam arrives at openai. badge works instantly.
receptionist: “yeah, we just made one.”
no pre-read. no doc. just a whiteboard, five people, and a model running somewhere hot enough to toast bread.
sam: “ready?”
noam smiles.
cut back to google. a calendar invite appears:
meeting: reduce meetings task force kickoff
duration: 90 minutes
required attendees: 214
sundar sighs, opens gemini, and types:
“how do we move faster?”
gemini responds:
“have you considered leaving google?”
smash cut to credits.
The Factory team seems very smart, and I'm eager to see how well their cost-optimizing new model router works. If they can pull it off, that is a huge accomplishment.
For Amp: we may try to build a model router that picks the best model for a task, but we don't intend to build a cost-optimizing model router based on the current state of the models. Here's why.
Every time we've looked into using cheaper models in Amp, we've benchmarked on tasks that reflect how people use agents for coding today. On these real tasks, the expensive frontier model was not only the best (obviously), but also usually the fastest and cheapest, when measuring end-to-end task completion.
Why? Cheaper-per-token models are less capable, which means that on complex real-world tasks they spend more tokens and time fixing mistakes along the way.
You can find plenty of cases where cheaper models are indeed faster and cheaper end-to-end. But such cases were rarer than we expected, and the differences were fairly small.
If you can easily detect such cases, then there is an opportunity here. But even then, on the AI hedonic treadmill, once people get a taste of frontier intelligence, they don't want to go back to using those more primitive prompts where cheaper models suffice. (Which is a good part of human behavior! It's how we decided to stop living in caves!)
If your tasks can be handled just as well by non-frontier models, I would strongly advise you to uplevel how you use agents and what you produce to stay competitive against people who are using frontier models.
In a power-law world, with rapid intelligence advances, try to get to the frontier and stay there.
"post-AGI, no one is going to work and the economy is going to collapse"
"i am switching to polyphasic sleep because GPT-5.5 in codex is so good that i can't afford to be sleeping for such long stretches and miss out on working"
Feel free to slave away at your 9-5 living in South Bay with the copium that a quicker commute is worth the sacrifice
I’ll be spending my prime making the most out of the beautiful city of San Francisco
@sdamico Agreed but quoted tweet is super misleading re: vitamins. Only in "well nourished typical" populations. Exercising 4+ days a week completely changes things, e.g. daily vitamin C reduces colds by 50% in athletes
> return flight to nyc gets canceled by snowstorm
> call united
> immediately connected with customer service (rare)
> voice is uncanny, def AI but they gave it a human-like accent
> takes ~20 min to get rebooked (pretty good imo)
> I ask if it's AI
> "haha no ma'am but I get that a lot"
> I ask it to calculate 228*6647
> it runs the calculation
> ggs
PixiJS v8.16.0 is out 🎉
- Return of Canvas 2D renderer
- Tagged text for inline styling
- Major SplitText improvements
- Cube textures, external textures, mip level rendering
- Broad engine stability updates
Plus a small preview for 3D 👀
https://t.co/i0ej5rL4eQ
A few random notes from claude coding quite a bit last few weeks.
Coding workflow. Given the latest lift in LLM coding capability, like many others I rapidly went from about 80% manual+autocomplete coding and 20% agents in November to 80% agent coding and 20% edits+touchups in December. i.e. I really am mostly programming in English now, a bit sheepishly telling the LLM what code to write... in words. It hurts the ego a bit but the power to operate over software in large "code actions" is just too net useful, especially once you adapt to it, configure it, learn to use it, and wrap your head around what it can and cannot do. This is easily the biggest change to my basic coding workflow in ~2 decades of programming and it happened over the course of a few weeks. I'd expect something similar to be happening to well into double digit percent of engineers out there, while the awareness of it in the general population feels well into low single digit percent.
IDEs/agent swarms/fallability. Both the "no need for IDE anymore" hype and the "agent swarm" hype is imo too much for right now. The models definitely still make mistakes and if you have any code you actually care about I would watch them like a hawk, in a nice large IDE on the side. The mistakes have changed a lot - they are not simple syntax errors anymore, they are subtle conceptual errors that a slightly sloppy, hasty junior dev might do. The most common category is that the models make wrong assumptions on your behalf and just run along with them without checking. They also don't manage their confusion, they don't seek clarifications, they don't surface inconsistencies, they don't present tradeoffs, they don't push back when they should, and they are still a little too sycophantic. Things get better in plan mode, but there is some need for a lightweight inline plan mode. They also really like to overcomplicate code and APIs, they bloat abstractions, they don't clean up dead code after themselves, etc. They will implement an inefficient, bloated, brittle construction over 1000 lines of code and it's up to you to be like "umm couldn't you just do this instead?" and they will be like "of course!" and immediately cut it down to 100 lines. They still sometimes change/remove comments and code they don't like or don't sufficiently understand as side effects, even if it is orthogonal to the task at hand. All of this happens despite a few simple attempts to fix it via instructions in CLAUDE . md. Despite all these issues, it is still a net huge improvement and it's very difficult to imagine going back to manual coding. TLDR everyone has their developing flow, my current is a small few CC sessions on the left in ghostty windows/tabs and an IDE on the right for viewing the code + manual edits.
Tenacity. It's so interesting to watch an agent relentlessly work at something. They never get tired, they never get demoralized, they just keep going and trying things where a person would have given up long ago to fight another day. It's a "feel the AGI" moment to watch it struggle with something for a long time just to come out victorious 30 minutes later. You realize that stamina is a core bottleneck to work and that with LLMs in hand it has been dramatically increased.
Speedups. It's not clear how to measure the "speedup" of LLM assistance. Certainly I feel net way faster at what I was going to do, but the main effect is that I do a lot more than I was going to do because 1) I can code up all kinds of things that just wouldn't have been worth coding before and 2) I can approach code that I couldn't work on before because of knowledge/skill issue. So certainly it's speedup, but it's possibly a lot more an expansion.
Leverage. LLMs are exceptionally good at looping until they meet specific goals and this is where most of the "feel the AGI" magic is to be found. Don't tell it what to do, give it success criteria and watch it go. Get it to write tests first and then pass them. Put it in the loop with a browser MCP. Write the naive algorithm that is very likely correct first, then ask it to optimize it while preserving correctness. Change your approach from imperative to declarative to get the agents looping longer and gain leverage.
Fun. I didn't anticipate that with agents programming feels *more* fun because a lot of the fill in the blanks drudgery is removed and what remains is the creative part. I also feel less blocked/stuck (which is not fun) and I experience a lot more courage because there's almost always a way to work hand in hand with it to make some positive progress. I have seen the opposite sentiment from other people too; LLM coding will split up engineers based on those who primarily liked coding and those who primarily liked building.
Atrophy. I've already noticed that I am slowly starting to atrophy my ability to write code manually. Generation (writing code) and discrimination (reading code) are different capabilities in the brain. Largely due to all the little mostly syntactic details involved in programming, you can review code just fine even if you struggle to write it.
Slopacolypse. I am bracing for 2026 as the year of the slopacolypse across all of github, substack, arxiv, X/instagram, and generally all digital media. We're also going to see a lot more AI hype productivity theater (is that even possible?), on the side of actual, real improvements.
Questions. A few of the questions on my mind:
- What happens to the "10X engineer" - the ratio of productivity between the mean and the max engineer? It's quite possible that this grows *a lot*.
- Armed with LLMs, do generalists increasingly outperform specialists? LLMs are a lot better at fill in the blanks (the micro) than grand strategy (the macro).
- What does LLM coding feel like in the future? Is it like playing StarCraft? Playing Factorio? Playing music?
- How much of society is bottlenecked by digital knowledge work?
TLDR Where does this leave us? LLM agent capabilities (Claude & Codex especially) have crossed some kind of threshold of coherence around December 2025 and caused a phase shift in software engineering and closely related. The intelligence part suddenly feels quite a bit ahead of all the rest of it - integrations (tools, knowledge), the necessity for new organizational workflows, processes, diffusion more generally. 2026 is going to be a high energy year as the industry metabolizes the new capability.
This has been said a thousand times before, but allow me to add my own voice: the era of humans writing code is over. Disturbing for those of us who identify as SWEs, but no less true. That's not to say SWEs don't have work to do, but writing syntax directly is not it.