In a world where everyone has a hammer, everything looks like a nail.
I'm a big believer in taking the time to pause, reflect, and write. I just finished writing a pretty long braindump of what I believe about the state of technology and people, why I believe it, and what has to change as a result of it (specifically in the products I'm building). I wrote it for myself, and now I'll refine to to be more consumable for a secondary audience (the team).
Writing does force such a unique clarity and organization that I refuse to give it up. Reading forces a type of focus and rigor that helps me think harder, for longer. Very anti docs and slides for the sake of it, but that feels like something that happens when the incentives aren't aligned.
Three bullet points have their purpose, but the long-pagers still do the job of fully elucidating the why and the how in a way that we probably shouldn't be offloading to AI just yet - not just because it's functional, but because it's also fun, and feels really nice to debate these things with our fellow humans in a substantive way.
From a friend who joined a fast-growth, later-stage startup:
"I was wondering why no one is doing PRDs here. Then I realized that the Head of Product and most PMs are 25-year-olds who don't have the attention span to *read* even a 1-page doc.
You lose them after 3 bullet points."
Sarah and Mike have been incredible partners to Huxe and myself. Startups aren’t easy - and you want people alongside you that have grit, thoughtfulness, and vision.
I’ve worked closely with @mvernal during the Huxe journey and he is truly one of a kind. I tell people he’s the kind of exec I wish I had at Google - maybe for me that’s shorthand for saying that he embodies the exact balance of pushing you to do your best, see the big picture, while supporting you wholeheartedly and unfailingly along the way. He’s extremely fair, kind, smart - and somehow also very funny. Mike helped me remember why I love building things so much - it’s really about the people along the way.
Grateful to have the Conviction crew, and not surprised one bit to see them crushing it.
In 2018, Sarah Guo became the youngest general partner in Greylock's 60-year history. She was 28. Four years later, she quit to launch Conviction, a firm staked entirely on AI.
Before ChatGPT shipped, she seeded Baseten and Harvey; each is now worth over $11 billion. In Conviction's first year, she wrote early checks into Sierra, Cognition, and Mistral; those three companies are now worth, together, $54 billion.
Andrej Karpathy worked out of Conviction's office until Anthropic hired him in May. Guo has been close to Jensen Huang for over a decade. Her first two calls after starting the firm were to Sam Altman and Nat Friedman.
And yet the investor closest to the AI frontier is betting against its biggest companies.
The two big frontier labs, worth close to a trillion dollars apiece, no longer just want to build the models. They also want to build every product and company on top of them, leaving nothing for anyone else.
The market is paying as though they might succeed. Of the $300 billion in venture capital deployed in the first quarter of 2026, the biggest quarter in the history of the trade, 65% went to only four companies: Anthropic, OpenAI, xAI, and Waymo.
Guo is betting the labs can't build everything, and she spends her days making sure of it. She won Harvey its first client. She flew across the country to take a single Baseten candidate to a four-hour lunch. On one wedding anniversary, she spent the whole weekend on back-to-back calls, keeping two founders on the line so they couldn't speak to rival firms. Twice a year, she flies the world's brightest young founders to San Francisco and inducts them into the fight.
In the months @domcooke spent reporting this piece, @saranormous had her fourth child, walked the Met Gala in 45 pounds of chainmail, and still answered her founders' texts within minutes.
Guo's parents arrived from China in 1987 with $50, built a company, and took it public at $1.2 billion. Then it went bankrupt. Guo grew up inside that startup. She built its first website, did her homework in a cubicle, and slept over for bug bashes. She loved it.
If two labs build everything, no one gets to do that again.
Welcome to Sarah's Wager. Read it below.
We first met @markiewagner axe throwing. She hit 10 bullseyes in a row and I knew we needed to be friends
We went on a 3 hour walk and she told me she'd been dreaming of automating work to machines since the 8th grade. She felt the time had finally come and she was ready to take her big swing. At that exact moment, two ladybugs landed on us simultaneously. Our ever auspicious journey officially began and @geniusventures committed the very first check alongside @traestephens at @foundersfund
Since then, as Markie would say, we've been on a quest together. Building the founding team, signing the first design partners, hosting amazing events, and of course raising this Series A led by the incredible @LM_Braswell at @kleinerperkins
Poetic has all the things: Stanford dropout, Thiel Fellow, FF/OpenAI/KP backing, a ridiculously talent dense team, 0 to 8 figures of ARR in less than a year, etc
But a lot of companies have those things
What I've never seen before is the soul in which Poetic is being built with. The intensity, urgency, meaning, togetherness. Anyone who has spent time in the “Forge” house knows exactly what I'm talking about
We are proud to be among the first believers, proud to write our largest check ever in the Series A, and proud to be on this inevitable quest to build Poetic with a truly generational team and investor base.
So excited for the world to finally meet @PoeticHQ
I’ve left Google DeepMind.
The last two years have been an incredible whirlwind.
A couple years ago, I joined a small startup called Codeium. There, I got to ship Windsurf, train SWE-1 (a frontier agentic coding model), go to DeepMind in the $2.4B acquisition. Now, I decided to leave the acquisition money and DeepMind.
I’m grateful to the mentors, teammates, and friends I worked with along the way.
At Windsurf, thanks to @_mohansolo and Douglas Chen, I got to see what a fast moving startup that ships relentlessly and builds for the future looks like. I learned from @thenickmoy how excellent research leadership can drive outsized innovation.
At DeepMind, I got to push the frontier of agentic coding, be part of the amazing team that shipped Antigravity and contributed to Gemini 3. DeepMind is a rare place: deeply curious people, exceptional research taste, and access to enormous compute and Google-scale infrastructure.
A few things that I learned:
1. Finding the right hill to climb. Now more than ever, there are a multitude of directions to push the frontier in AI research. It’s easy to optimize for the wrong benchmark or capability. You should step back regularly to question if you are climbing the right hill, and adjust course often.
2. The secret to being a fast-moving team. Moving quickly is not just about working hard and long hours. It requires making concrete bets about where the world will be in 6 months, aligning around them, and cutting everything else. This was our journey from the Codeium Extension → Windsurf IDE → SWE-1 → Antigravity → Antigravity CLI
3. Silicon Valley is small. Since the split of Windsurf to DeepMind and Cognition, many of my colleagues have gone to other exciting places - Thinking Machines, OpenAI, xAI, Cursor, fast-moving startups, or started their own companies. I’m grateful to have worked with so many talented, hungry people whose stories are not yet finished.
So what’s next?
We are living in one of the most exciting and powerful times in human history. Just like we transformed software engineering, soon every industry, every unit of work will be radically transformed, democratized, accelerated. With this comes new challenges, and new doors of frontier research to be opened.
More soon.
Woke up this morning to see a ton of new Japanese users! The feedback has been amazing to see in the Discord, thank you for trying Huxe out and thank you for reviewing it @jetdaizu 🇯🇵
新しく来てくれた皆さん、ようこそ!🎉
会えて嬉しいです!よろしくお願いします���
"...Huxe shows how democratization of AI is done, and I have become a fan. It is, hands down, the best way to listen to AI, and my favorite one."
💕 neat!!!
@EverydayAI_@kevinrose I do this too! Each agent does this after they’re done. I feel like my method is definitely not token efficient but I catch a lot fewer dumb mistakes
It's gotten to a point where I feel like I'm living in between some sort of utopian future or a black mirror episode right before everything goes bust.
I set up long running tasks (both for coding and research), review some of the more complicated implementations and debate it with the orchestrator - reach some point of agreement - then I leave it alone to go work.
In the meantime, I go do my "human" things - talk to the team and users, fiddle designs, read the things I actually want to read (instead of summarize), go try new apps, write out my thoughts for new ideas, learn new skills (currently sculpting), and exercise.
Then I get a notification to come back to the terminal and repeat the loop. It's a crazy way to work, but also very... peaceful.
Judging by my tl there is a growing gap in understanding of AI capability.
The first issue I think is around recency and tier of use. I think a lot of people tried the free tier of ChatGPT somewhere last year and allowed it to inform their views on AI a little too much. This is a group of reactions laughing at various quirks of the models, hallucinations, etc. Yes I also saw the viral videos of OpenAI's Advanced Voice mode fumbling simple queries like "should I drive or walk to the carwash". The thing is that these free and old/deprecated models don't reflect the capability in the latest round of state of the art agentic models of this year, especially OpenAI Codex and Claude Code.
But that brings me to the second issue. Even if people paid $200/month to use the state of the art models, a lot of the capabilities are relatively "peaky" in highly technical areas. Typical queries around search, writing, advice, etc. are *not* the domain that has made the most noticeable and dramatic strides in capability. Partly, this is due to the technical details of reinforcement learning and its use of verifiable rewards. But partly, it's also because these use cases are not sufficiently prioritized by the companies in their hillclimbing because they don't lead to as much $$$ value. The goldmines are elsewhere, and the focus comes along.
So that brings me to the second group of people, who *both* 1) pay for and use the state of the art frontier agentic models (OpenAI Codex / Claude Code) and 2) do so professionally in technical domains like programming, math and research. This group of people is subject to the highest amount of "AI Psychosis" because the recent improvements in these domains as of this year have been nothing short of staggering. When you hand a computer terminal to one of these models, you can now watch them melt programming problems that you'd normally expect to take days/weeks of work. It's this second group of people that assigns a much greater gravity to the capabilities, their slope, and various cyber-related repercussions.
TLDR the people in these two groups are speaking past each other. It really is simultaneously the case that OpenAI's free and I think slightly orphaned (?) "Advanced Voice Mode" will fumble the dumbest questions in your Instagram's reels and *at the same time*, OpenAI's highest-tier and paid Codex model will go off for 1 hour to coherently restructure an entire code base, or find and exploit vulnerabilities in computer systems. This part really works and has made dramatic strides because 2 properties: 1) these domains offer explicit reward functions that are verifiable meaning they are easily amenable to reinforcement learning training (e.g. unit tests passed yes or no, in contrast to writing, which is much harder to explicitly judge), but also 2) they are a lot more valuable in b2b settings, meaning that the biggest fraction of the team is focused on improving them. So here we are.
I didn't realize it but Gemini added the video links to the chat (not the generated mini app that I asked for) but the link doesn't take you to a video, it takes you to a google search that's pre-filled to search for a video. Annotation is excellent but I don't love the bad video link, it really bugs me when this happens lol
I've been waiting to try this model (specifically via API) to see what interesting Instagram / FB related things can be built.
Since API isn't available yet, I tried via https://t.co/HChjkoVc7O one of my favorite use cases of "Given a location, plan a walking food tour for me with videos" and did a side by side of Gemini vs. Muse Spark vs. Claude Opus 4.6
Gemini:
- Obviously very advantaged with Maps and commensurate data, creates a very familiar Maps-looking UI with the links to website and easy CTA to get directions
- However does NOT give me links to YouTube videos about the suggested restaurant, which I asked for :(
Claude Opus 4.6:
- Bad map, and links were sadly not actual links but that annoying thing LLMs do when they're cheating and send you to a YouTube link that's a search for the business name lol
Muse Spark:
- Pretty good map and actually links to unique Reels!
- Each recommendation had an actual Reel attached that wasn't a landing page for a search
- I appreciated the little annotations too that seemed sourced from Meta properties
Overall very exciting progress - I'm looking forward to playing with the API and buiding on it, It would be quite a unique unlock for developers~
1/ today we're releasing muse spark, the first model from MSL. nine months ago we rebuilt our ai stack from scratch. new infrastructure, new architecture, new data pipelines. muse spark is the result of that work, and now it powers meta ai. 🧵