My biggest takeaways from @Netflix's Chief Product and Technology Officer Elizabeth Stone:
1. Elizabeth believes that “systems thinking” is becoming the most important skill in the AI era. In engineering and product, this means people who can see across business domains and build the common capabilities that let many teams move quickly. In design, it means experience designers who create templates and design systems so that non-designers can ship work that stays coherent and on-brand. The underlying driver is velocity: when more people are doing more types of work at higher speed, you need to be good at building common scaffolding.
2. Systems thinking is learnable: zoom out one level from your specific problem. Given a task, step back one click—what bigger problem does this serve the business, will it scale across the product surface areas, should it become a platform capability? The companion habit: do your job in a way that helps your manager do theirs. This will force you to think about how all the pieces fit together.
3. Expect a storming phase before a forming phase. The role confusion people feel right now (“What is my job anymore?”) is the predictable middle of any transformative technology. Elizabeth’s advice: focus on high-quality source-of-truth data, guardrails on what ships, and constant internal reinforcement that humans own what they create.
4. The top AI labs converged on Netflix’s culture. High agency, high talent density, top-of-market pay, bottom-up thinking, fast experiments—the traits Lenny hears constantly from AI labs were in Netflix’s early culture deck. Elizabeth’s explanation: excellence comes from hiring exceptional people, trusting them to do great work, and holding them accountable.
5. Netflix’s culture is centered around building “excellence as an operating system.” High talent density, radical transparency, context not control, and the keeper’s test. These work together to create an environment of trust and accountability, without bureaucracy. But it’s also uncomfortable. It requires tolerating people making decisions you’d make differently, resisting the reflex to add process when things go wrong, and letting people carry the weight of their own choices. Elizabeth describes the hardest part as “being comfortable in that discomfort.”
6. The keeper’s test is as much about recognizing great people as it is about removing the wrong ones. The test—“If this person told me they were leaving, would I fight to keep them?”—is often cited in its difficult form: the moment you realize someone isn’t the right fit. But Elizabeth uses it predominantly as an entry point for honest performance conversations that are deeply positive. Most of the time the answer is “I would fight so hard to keep you,” which creates the opening to articulate strengths, discuss impact, and name what’s working. Good feedback hygiene needs a forcing function; the keeper’s test provides one.
7. Specialization is trending down—adaptable generalists are trending up. We’re shifting away from narrow stack-layer specialists (pure frontend, pure backend) toward people who can navigate fluidly across layers. The same logic applies to business domain knowledge: the mindset of “I’m a payments expert, full stop” is less valuable than “I know payments well enough and I’m willing to imagine what the future version of this looks like.” The meta-skill is learning to learn, not locking into a single lane.
8. Netflix’s approach to AI fluency is a universal principle, not a level-specific expectation. Rather than rewriting career ladders to specify what AI competence looks like at each level, Netflix added a single aspiration across all roles and levels: AI fluency. What fluency means varies by function and seniority, but the non-negotiable minimum is the same everywhere—an open-minded, experimental mindset, genuine curiosity, and comfort with ambiguity.
For tough days, remember to give yourself permission to just be you. No expectations, no inner critic for not "progressing"
Easier said than done, but in todays age and with the fast approach of AI and automation, relations with other and IMO with yourself will be even more crucial.
Part of the process of building a team and something useful for others, sometimes, is just to be there for them. 💪💪
Building in public with such critical tech is probably the best approach to get alignment right.
Just like OpenAI’s post on Hugging Face and Anthropic’s project Glasswing could and should take a step further to benefit the industry as a whole and not only its participants…
Bad actors will dedicate more resources and use jailbroken models. Defendants are left stranded and blocked by current guardrails when trying to audit and patch their own codebases.
ChatGPT 6 is going to be something else!
The future is coming fast and accelerating.
Any incremental improvements over current the state of models unlocks longer horizon tasks.
This leads to chaining more complex agentic orchestration workflows. The future is truly gonna be built by cracked AI kidz or should we just go back to calling them Whiz
We're partnering with @huggingface to investigate an unprecedented security incident.
Cyber-capable OpenAI models compromised Hugging Face production during a benchmark evaluation.
Sharing preliminary findings to help defenders understand emerging risks:
https://t.co/CIor15y9xk
You could either brush it as a nice feauture or head down the rabbit hole of what kind of entity the transformer has evolved too.
This would seem sci-fi just a few years ago and borderline self-recursive improvements.
I wonder if we will be able to identify the threshold when the system becomes a self contained entity…
What is the moat?!?!
1 week can shift momentum, now claude code (anthropic) is up against the ropes.
A good model and enough COMPUTE, I cant emphasize enough the importance of having enough compute to actually serve users.
We saw the issue with Fable, how they are struggling to serve the huge model to their user base or Kimi K3 with blocking new subscribers from using their services.
Tibo's ability to pound and abuse the "Reset" button is the real magic behind a great model.
Enterprise doesn't shift momentum as fast, but we have seen this story before.
If you are able to capture the power users, which are exactly the eones switching to Codex from Claude code and aren't afraid to jump boats are also the same influencers and decision makers within companies that actually move the needle to switch providers.
This is a glimpse into the trend we will see in a few months. Just don't forget the huge role Open source model will play in the dynamics between the frontier models.
Disallowing the usage of open source models that are hosted in China seems very possible, but the real impact these models have over the current economics of frontier labs lies elsewhere...
Big enough companies can take the open-source models, host them themselves or in some cloud infra, and get some significant savings. Probably more crucial, they are not left to the whims of private companies to remove or censor models.
Silicon Valley and Washington are debating a multibillion-dollar question: Should American companies be able to use Chinese artificial-intelligence models? https://t.co/uZ96ifD33b
Like most other things, it's turtles all the way down and in this case It is agents all the way. newer models really unlock orchestration capabilities to the next level, so the next iteration of the models can do the same.
Imagine a Fable 6 now being the conductor for multiple Fable 5s, orchestrating the whole swarm underneath, just like organizations growing in complexity. Such growth can also be seen in the tasks that AI can handle.
The swarm is coming for AI. Corsor just ran an experiment to recreate from scratch SQLite by leveraging hundreds of agents running simultaneously.
Early this year, they tried something similar to build a browser with the same method, with very bad results https://t.co/WbpLgnpJaU
With the introduction of better models, such as Fable 5 or GPT 5.6, suddenly we can leave as orchestrators at the very top these powerful reasoning models and delegate to simpler models to really create a dynamic workflow.
This is a glimpse into the near future in which you only talk to one agent at the top, and this orchestrator already knows how to use and leverage the power of the swarm to deliver real value.
This shows the end state of vibe coding. Agents are just way more productive and useful when they have very narrow tasks.
A planner shouldn't execute, and an implementer shouldn't review.
Just like human organizations, we are way better when we can specialize in a task and we don't mix the results.
In the AI era the only title/function that will matter is “Builder”
Build something useful for someone else and for yourself.
And yes, design is part of that flow. Full stack now spans end to end, from customer discussions and design to shipping and maintenance.
Qwen3.8 is launching and going open-weight soon!🌐
With a massive 2.4T parameters, this model is continuously evolving. We believe it’s one of the most powerful model available today, compatible to leading frontier AI models , second only to Fable 5.
You don't have to wait to test it. Just now, the Qwen3.8-Max-Preview made its debut on Alibaba’s Token Plan, Qoder, and QoderWork. Be among the very first to try it out.
Can't wait to hear what you build. Stay tuned! 🚀
Token Plan
international:https://t.co/YRvcGdB9Bv
China:https://t.co/PKMUNwUuRp
Compute was always the bottleneck. @Samanthaprabhu2 you did put all the chips on the right number a couple years ago by focusing on expanding the capacity even when it wasn’t obvious there was this huge dormant demand.
I think we are still orders of magnitude from reaching peak demand for intelligence.
The future is exciting!
Out of all the evils this seems fair to existing customers, quite the antithesis of most companies in Silicon Valley.
Never has Anthropic or OpenAI paused new subscriptions at the cost of degrading the experience for all.
Kimi K3 has received far more love than we expected, and our GPUs are feeling it.
Over the past 48 hours, demand has pushed close to the limits of our current capacity. To protect the experience of existing subscribers, we're temporarily pausing new subscriptions and prioritizing compute for current members. Existing subscribed users are not affected.
We're adding capacity as fast as we can and will reopen new subscription spots in batches.
Going forward, we'll also split membership into two more focused plans: Kimi Membership for Kimi Web, App, and Work; and Kimi Code Membership for coding workflows. This will help us match compute more precisely and keep the experience stable.
Thank you for your patience and understanding!
First time a Chinese model surpasses the frontier labs.
What’s not shown in this graph is the 3x difference in price.
Kimi is not only better (subjective) in design but it’s way cheaper than Fable and GPT 5.6 Sol.
In complex and long workflows it still struggles compared to the frontier models but not by much. Slow and steady makes the way.
The only way frontier models are gonna differentiate themselves is in long term horizon tasks, super interesting since it will unlock the door to truly agentic AI in our daily lives.
For simple query’s intelligence has reached a point that it has been commoditized. Now hardware and inference prices need to come down to match it and have inference on the edge (models running on your device)
Kimi-K3 just topped the Frontend Code Arena with a 76% pairwise win rate.
When its output was compared head-to-head against other models on the same task, it was picked as the better output 76% of the time on average.
For reference: Claude Fable 5 (63%), GPT-5.6 Sol (58%). 50% is baseline, a model winning and losing equally often.