One of the more subtle pitfalls of starting a startup when you're too young is that you can't hire well, because you haven't had enough experience to be a good judge of people. You can judge technical ability but not character, so you end up hiring smart jerks.
One of the more subtle pitfalls of starting a startup when you're too young is that you can't hire well, because you haven't had enough experience to be a good judge of people. You can judge technical ability but not character, so you end up hiring smart jerks.
The first time I encountered the title Chief of Staff (CoS) was probably 20 years ago, and I'll admit I was skeptical. The first CoS I met was often used as a surrogate CEO -- showing up to meetings because the CEO couldn't. I didn't think that was a good thing, and it left me with a pretty negative impression of the role.
Over time, though, I watched the CoS role evolve from something that only larger companies embraced, now all the way down to surprisingly early-stage startups. Maybe the CoS role just made a bad first impression on me, and I needed to open my mind to the possibility that it was actually a good idea.
Over time, I've worked with numerous Chiefs of Staff, and my perspective has changed considerably. The purpose of this post is to highlight a strong, three-part series, that my Balderton colleagues recently wrote on the Chief of Staff role.
A few reactions:
I liked the framing. The right question isn't "Should companies our size have a Chief of Staff?" It's "Has the founder become the bottleneck?" A Chief of Staff should solve coordination problems, not compensate for missing executives.
I liked the discussion distinguishing the Executive Assistant and Chief of Staff roles. That's an important distinction that's often poorly understood. I might draw the boundary a bit differently (e.g., competitive, financial model), though. If the work is helping the CEO manage their day, it's EA work. If it's helping the CEO think, decide, and run the business, it's Chief of Staff work.
I liked the qualities they emphasize. Broadly, I think founders should be looking for a flexible, adaptable, low-ego, get-shit-done person who's comfortable operating through influence rather than positional authority.
On authority, I'd make one distinction. I completely agree that a Chief of Staff needs exceptional access to the CEO and visibility into the business. Where I'd differ is on formal authority. To me, the role isn't about giving someone positional authority -- it's about giving them the CEO's confidence. A great Chief of Staff gets things done through influence, judgment, and credibility. They're a leveraging resource for the CEO, not another executive running part of the company.
My addition to the interview process would be to spend more time together. Have dinner. Spend half a day solving real problems. Don't just evaluate the candidate -- experience what it's like to work together. A Chief of Staff is ultimately a thought partner, and there's no substitute for seeing that dynamic firsthand. Heck, do a try-and-buy if the candidate is open to it.
Overall, well worth reading if you're considering hiring a Chief of Staff, or becoming one. Link in the comments.
"AI is like oil."
Anthropic sells that barrel for $56. OpenAI sells it for $26. Elon sells it for $1. Zuck sells it for $1.50. Google sells it for $1. The Chinese sell it for 50 cents.
Same barrel. Wildly different prices.
what do you think will happen?
Highly recommended for founders, especially first time founders. YC is like beauty pageant. It takes a unknown girl from the village and make her a beauty queen over night!
Every company that changes the world starts with someone deciding to build.
If you're making something people want, we'd love to hear from you.
Apply to the YC Fall 2026 batch by July 27: https://t.co/gNl84ElBrq
Every company that changes the world starts with someone deciding to build.
If you're making something people want, we'd love to hear from you.
Apply to the YC Fall 2026 batch by July 27: https://t.co/gNl84ElBrq
The buried lede in the self-driving company data: governance came first. Access policies, token proxies, audit logs, zero trust, all before agents touched anything. Everyone asks what agents can do. The better question is what yours are allowed to touch, and who logged it.
The largest open-weights model ever announced, and by its own benchmark table it still trails the top two closed models. Both facts matter. The open-frontier gap is now measured in weeks, and the verification burden just moved to the buyer.
Introducing Kimi K3: Open Frontier Intelligence
🔹 2.8 Trillion Parameters, 1 Million Context, Native Multimodal
🔹 Kimi Delta Attention enables up to 6.3x faster decoding in million-token contexts
🔹 Attention Residuals deliver ~25% higher training efficiency at <2% additional cost
🔹 Built for long-horizon agentic coding and self-evolving workflows
Kimi K3 is now live on on https://t.co/zrk6zZxZUo, Kimi Work, Kimi Code, and the Kimi API.
Open Weights by July 27, 2026.
🔗 API: https://t.co/XCrgjXAqMw
🔗 Tech blog: https://t.co/YTfiMSNM1f
AI-generated apps just stopped being screenshots. Artifacts that call MCP connectors mean the dashboard you describe can fetch live data and take actions, scoped to each viewer's own permissions. The long tail of internal software that never got built is about to get built.
Claude Code artifacts can now call MCP connectors, letting you build dashboards and apps that can fetch information and take actions for each viewer on demand.
Available on Pro, Max, Team, and Enterprise plans. Not available on publicly-shared artifacts.
A frontier lab shipping its first model with the caveat "not the strongest available" is the tell. The market is splitting: rent the smartest generalist through an API, or take an open base and own what it learns about your domain. Both bets are now fully priced.
Today, we are introducing Inkling.
Inkling reasons efficiently across text, image, and audio modalities. We are making the full weights available.
https://t.co/Ghebq5mG30
Available today for fine-tuning on Tinker. Play with it in the Inkling Playground. 🧵
Free frontier AI for every US teacher, and the model is the least interesting part. Skills library, curricula grounding, standards mapping for all 50 states: the scaffolding is the product. The same will be true in every regulated profession.
We're introducing Claude for Teachers: free access to premium Claude capabilities for verified K-12 educators in the US, with a library of teaching skills and a direct connection to evidence-based curricula, mapped to academic standards in all 50 states.
https://t.co/5hZZijVPCV
The most telling AI news this week: a frontier lab CEO asking for pre-deployment risk assessment. Every other high-stakes industry already works this way. When the builders start requesting the exam, buyers should stop accepting vendors who skip it.
The bottleneck for agent adoption was never model capability. It is whether you can define what good work looks like precisely enough to score it. Companies that cannot eval their knowledge work cannot delegate it, to AI or to anyone else.
One of the many properties that code has that makes it highly amenable to agents is that you can more or less quickly test it. You can either go see if the application works manually, or you can actually run a test on what you built.
Most other areas of work don’t have this benefit. You only get the testing when the final product hits the real world in some capacity - a stock trade is executed, a contract is negotiated, a sales pitch is delivered, and so on.
There’s probably going to be a whole new set of opportunities for how we begin to test the rest of work in this way. Ultimately it will mean more agents being layered into workflows.
It also means we need much better evals on most of our workflows. Most work today in enterprises doesn’t have an associated eval to know if something broke or improved with a model, prompt, or system change.
The enterprises that are able to eval their knowledge work the best also stand to gain the most from AI. Will become a critical aspect of agent adoption over time.
Same model, different language, different behavior. For anyone deploying AI across markets, this research decides whether your agent is one product or several products wearing the same name. Benchmarks won't catch the difference. Your users will.
In previous research, we found that Claude expresses over 3,000 values, like honesty and warmth. In new work, we asked how the values Claude expresses vary between Claude models and across languages.
We analyzed 300K+ anonymized conversations to find out.https://t.co/PgxsMXipt5
Every RL benchmark resets the world between episodes. Real operations never reset: past decisions compound, environments drift, fixed policies decay. Scoring agents on persistence instead of episodes is the eval enterprises have been missing.
Today we present Morpheus, a persistent enterprise simulation platform designed to make Continual Learning a reality. Morpheus is the world’s first real world Reinforcement Learning environment.
Every Reinforcement Learning environment operates in the game world. Benchmarks like Atari, OpenAI Gym, MuJoCo, and Procgen are all small, game-like worlds that reset every few minutes.
But the real world never resets. A business keeps running and evolving everyday.
We tested how frontier LLMs would perform in realistic and dynamic business environments 🧬on Morpheus. The main conclusion was that LLMs are not continual learners.
🧵Here’s how we did it and what we learned:
The viral "Chinese models hit 46%" chart measures one routing platform, not the market. Most enterprise traffic goes direct through vendor APIs and agentic tools that never touch OpenRouter. The trend is real, but this chart cannot tell you how big it is.
Key takeaway: "The seller learns more and more about you as you use what you purchased, while you learn very little about what the seller is learning in return."