Have to say, the #fable throttling - sometimes all the way back to Sonnet - on very innocuous prompts is a bit of a joke at this point. 5.6 is handling without issue.
Dario Amodei just gave his first interview since the Pentagon blacklisted his company. The toll is visible on his face.
He was asked one question. What would you say to the President right now?
He didn’t hesitate.
Amodei: “We are patriotic Americans. Everything we have done has been for the sake of this country.”
Anthropic built their models to defend America. They were the first AI lab cleared for classified military systems. They wanted to help the warfighter.
But the Pentagon demanded unrestricted access to fully autonomous weapons and mass surveillance of American citizens.
Amodei drew the line.
The government responded with emergency Cold War powers. A supply chain designation normally reserved for foreign adversaries. A six-month federal phaseout ordered from Truth Social.
Amodei: “When we were threatened with supply chain designation and Defense Production Act, which are unprecedented intrusions into the private economy, we exercised our classic First Amendment rights to speak up and disagree with the government.”
The administration framed Anthropic’s refusal as anti-American.
Amodei’s response dismantled that framing in one sentence.
Amodei: “Disagreeing with the government is the most American thing in the world.”
Here is the deeper paradox nobody in Washington wants to say out loud.
We are in a geopolitical race against autocratic adversaries who use AI for mass surveillance of their own citizens and autonomous weapons with no human oversight.
The Pentagon demanded that Anthropic build those exact capabilities for America.
Amodei: “The red lines we have drawn, we drew because we believe that crossing those red lines is contrary to American values.”
You cannot defeat authoritarianism by adopting its methods.
You cannot defend the open society by forcing private companies to build its antithesis under threat of wartime emergency powers.
Anthropic held the line. Got blacklisted for it. And came out the other side saying the same thing they said going in.
That is what it actually looks like to mean it.
Nvidia, Amazon, Google will have to divest from Anthropic if Hegseth gets his way. This is simply attempted corporate murder. I could not possibly recommend investing in American AI to any investor; I could not possibly recommend starting an AI company in the United States.
As a lawyer who uses LLMs every day at work, I feel qualified to respond.
First, hallucinations are no longer a problem. Consistent with the prediction you quoted from 2023, GPT-5.x almost never hallucinates. And overall, the percentage of inaccurate responses I get from GPT-5.2 Pro is lower than the percentage of inaccurate responses I would get from a competent junior associate (yes, fully accounting for hallucinations).
Second, people wildly overestimate the difficulty of most tasks performed by lawyers. The vast majority of the things we do are not nearly as challenging intellectually as solving an Erdos problem. Key skills for a lawyer are attention to detail, ability to synthesize and reason through precedent, ability to construct logical arguments, writing, research. LLMs are *very* good at most of these things even today, and top-tier LLMs (GPT-5.2 Pro) are excellent at them.
Put in another way, I feel that the biggest barrier to widespread adoption of AI by lawyers today is connectivity, interfaces, harnesses - *not* intelligence of the best models, and certainly not hallucinations. Unclear to what extent these issues will be resolved in the next 12-18 months, but given how economically valuable lawyers' work is, I wouldn't be surprised to see significant progress on that front. It's also worth considering that, given the general trend of rapidly falling costs of running reasoning models, it is likely that a model as intelligent as GPT-5.2 Pro, but *much* cheaper and faster, will be publicly available in the next 12-18 months.
Note that the above assumes (conservatively) that the next 12-18 months in AI will be relatively boring: no continual learning, no drop-in virtual employees, not much further progress in agentic AI (Codex), no significant progress in intelligence possessed by the best models. Relaxing these assumptions would mean that we should expect even faster progress.
@emollick So true. Can only squeeze cost and productivity gains on “current” process so far - facading larger “system” and ways of working opportunities.
TL;DR: We built a transformer-based payments foundation model. It works.
For years, Stripe has been using machine learning models trained on discrete features (BIN, zip, payment method, etc.) to improve our products for users. And these feature-by-feature efforts have worked well: +15% conversion, -30% fraud.
But these models have limitations. We have to select (and therefore constrain) the features considered by the model. And each model requires task-specific training: for authorization, for fraud, for disputes, and so on.
Given the learning power of generalized transformer architectures, we wondered whether an LLM-style approach could work here. It wasn’t obvious that it would—payments is like language in some ways (structural patterns similar to syntax and semantics, temporally sequential) and extremely unlike language in others (fewer distinct ‘tokens’, contextual sparsity, fewer organizing principles akin to grammatical rules).
So we built a payments foundation model—a self-supervised network that learns dense, general-purpose vectors for every transaction, much like a language model embeds words. Trained on tens of billions of transactions, it distills each charge’s key signals into a single, versatile embedding.
You can think of the result as a vast distribution of payments in a high-dimensional vector space. The location of each embedding captures rich data, including how different elements relate to each other. Payments that share similarities naturally cluster together: transactions from the same card issuer are positioned closer together, those from the same bank even closer, and those sharing the same email address are nearly identical.
These rich embeddings make it significantly easier to spot nuanced, adversarial patterns of transactions; and to build more accurate classifiers based on both the features of an individual payment and its relationship to other payments in the sequence.
Take card-testing. Over the past couple of years traditional ML approaches (engineering new features, labeling emerging attack patterns, rapidly retraining our models) have reduced card testing for users on Stripe by 80%. But the most sophisticated card testers hide novel attack patterns in the volumes of the largest companies, so they’re hard to spot with these methods.
We built a classifier that ingests sequences of embeddings from the foundation model, and predicts if the traffic slice is under an attack. It leverages transformer architecture to detect subtle patterns across transaction sequences. And it does this all in real time so we can block attacks before they hit businesses.
This approach improved our detection rate for card-testing attacks on large users from 59% to 97% overnight.
This has an instant impact for our large users. But the real power of the foundation model is that these same embeddings can be applied across other tasks, like disputes or authorizations.
Perhaps even more fundamentally, it suggests that payments have semantic meaning. Just like words in a sentence, transactions possess complex sequential dependencies and latent feature interactions that simply can’t be captured by manual feature engineering.
Turns out attention was all payments needed!
@Starbucks Significant customer service opportunity: waiting 40 minutes for a mobile order at PDX C6, while we watch people coming through the line getting their drinks 20-30 minutes ahead of us??
When picking among the 9 AI models that are now available from OpenAI, the rules are easy:
1) The model with the biggest number is mostly not the best
2) Mini means worse, except for the mini that is the second best
3) o1 pro beats o3-mini-high beats o1 beats o3-mini, naturally
Last Friday, I had one of the most intellectually amazing experiences of my career:
I got to do the following Idealcast interview (yes, they're coming back!) of Dr. Carliss Baldwin, the William L. White Professor of Business Administration, Emerita at the Harvard Business School.
Among many things, she is the researcher who pioneered the study of modularity and how it increases option value — and that there are cases such as IBM and Amazon that it creates so much surplus value it can "blow entire industries apart."
Her mentor was Dr. Robert C. Merton. He worked with Drs. Myron Scholes and Fischer Black, who the Nobel Prize in Economics in 1997. Their insights showed how to precisely value options, which are the right but not the obligation to take an action in the future.
Dr. Baldwin used the same principles of option theory to explain value creation in modular systems and organizational design.
In my quest to understand how to see what it looks like when option value is created (especially for GenAI!), and how one would measure it, I was able to ask her, as well as Dr. Steven Spear (who had Dr. Baldwin as his advisor when he worked on his doctoral dissertation at HBS), and Steve Yegge, famous for his 20 years of work at Amazon and Google.
My goal for this amazing 2 hour interview was to explore the following:
- Option Value in Manufacturing: How the Toyota Production System creates and measures value through modularity — what does creation of option value look, how does one measure it? How does that relate to things like doing 4,000 daily andon cord pulls through localized line stops and rapid experimentatio?.
- Option Value in Hardware Development: How did the IBM System/360 project generate 25x value creation through 25 modules and 25 parallel experiments, revolutionizing computer architecture. How do we replicate the calculations she did to get 25x higher value accreditation?
- Option Value in Software Architecture: How did Amazon's transformation from monolith to microservices in the early 2000s create massive option value through team independence and rapid deployment capabilities?
- Option Value in Modern Development: How GenAI is creating new forms of option value by giving developers "more swings at bat" and enabling rapid exploration of alternatives.
- Option Value Theory: How Merton's work on temporal options and Baldwin's work on spatial modularity combine to explain value creation across domains.
It was such an amazing conversation, to hear how their collective experiences give life to theory and vice versa. The dialogue between manufacturing floors, software architectures, and financial models was unflippingly amazing.
But the coolest part was that the simple formula that concretized everything! I think this is something that every technology leader needs to know!
** Understanding Option Value Through NK/T and σ
Incredibly, there’s a simple formula that ties all of these concepts together. It’s NK/T and σ
N = number of modules that can be worked on independently
K = number of parallel experiments that can be run on each module
T = time required for each experiment cycle
NK/T represents how many independent experiments you can run in parallel divided by how long each takes. For example, in the IBM System/360 case, they had ~25 modules (N) and could run ~25 experiments per module (K), massively accelerating their ability to innovate compared to a monolithic design.
(Note that K is within one module. So at IBM, the total number of experiments possible was actually much larger - potentially 25 × 25 = 625 experiments across the whole system. Note how number of modules multiplied by the total number of parallel experiments rises exponentially!!)
Similarly at Amazon, they went from one module (the monolith) to tens of modules, to hundreds and eventually thousands. The deployments per year went from hundreds in 1999 and almost ground to a halt, doing only tens of deployments per year in the early 2000s. This led to the "Thou shalt use APIs" Jeff Bezos memo which Steve Yegge told the world about. This:
- Increased N: The number of independent modules grew exponentially
- Increased K: The number of parallel experiments that could be performed per module
- Massively reduced T: Going from quarters to do an experiment to maybe days or maybe even hours
Given the hyper-competitive e-commerce marketplace in the early 2000s, σ was high. We did a back of the napkin calculation and guess that the option value created was much higher than even the System/360 project in 1960s. (Some argue that AWS was a byproduct of the modularization effort.)
** The Role of Uncertainty (σ)
σ (sigma) represents volatility or uncertainty, ranging from 0 to potentially infinite, where:
σ = 0 means perfect knowledge/certainty
In this case, option value is zero because you know exactly what to do
You don't need the "right but not obligation" to decide later. You can just make the optimal choice now
Example: If you knew tomorrow's stock price with certainty, you wouldn't need options - you'd just buy or sell the stock directly
As σ increases, so does option value
σ = 0.2 represents low volatility
σ = 0.4 represents medium volatility
σ = 0.8 represents high volatility
The higher the uncertainty, the more valuable it is to have options
This explains why options are more valuable in uncertain domains:
- In manufacturing with established processes: traditionally assumed to have low σ (but see the next section for Toyota’s big insight!)
- In new product development: higher σ
- In software/technology innovation: very high σ
- In completely new domains (like early GenAI): extremely high σ
The combination of these metrics helps explain why modular systems can create such enormous value - they let you run many parallel experiments (high NK/T) to capture value in uncertain environments (high σ).
## Toyota's Big Insight
Toyota made a revolutionary discovery that challenged conventional wisdom: even in seemingly "repetitive" manufacturing, σ (uncertainty/volatility) is actually quite high. While traditional mass production assumed standardization and rigidity, Toyota recognized that there is so much variance in high volume manufacturing. Quality issues, supplier issues, customer demand, fluctuations in cost, etc.
Instead of trying to eliminate this uncertainty, they built a resilient system that can create value from it.
Their response was three-fold: they expected and embraced uncertainty, created cheap options to respond (like the andon cord system pulled 4,000 times daily), and made exercising these options inexpensive through modular line segments that could stop independently. This created extraordinary capabilities: they could run multiple model years simultaneously, perform 60 line-side store changes per day, and implement rapid die changes (SMED) - all while maintaining high quality and efficiency.
This success can be understood through option value metrics: they achieved high NK/T through multiple independent modules (N), many parallel experiments (K), and quick cycle times (T), while recognizing and exploiting high σ (uncertainty). While other manufacturers focused on copying visible tools like kanban and andon cords, they missed this fundamental insight about uncertainty and option value creation, making Toyota's system difficult to replicate and leading to their sustained competitive advantage in global manufacturing.
** Bonus: Visualizing Option Value Creation
As a bonus, I asked ChatGPT-4 to make me a visualization of how N*K/T and σ interact with each other. This was to try to understand and replicate Dr. Baldwin's calculation of how 25 modules * 25 experiments created 25x value creation at IBM. Amazingly, it gave me this incredible JavaScript visualization which you can rotate in 3D. We live in an age of miracles.
Today's "DeepSeek selloff" in the stock market -- attributed to DeepSeek V3/R1 disrupting the tech ecosystem -- is another sign that the application layer is a great place to be. The foundation model layer being hyper-competitive is great for people building applications.
I’ll tell you and this is one of my favorite topics - it’s not sexy so it’s rarely discussed
I wish there were books on this subject
Companies have both a direction (vector) and a speed (magnitude). Either can be hand tuned at any moment. Getting the direction right is fundamental - if it’s wrong, no matter how fast you move, you won’t reach the right destination
You can see Figures speed, but what you don’t have visibility on is our direction
What you’re not seeing is a playbook to nail the direction:
> goals planned to the year and then down to the day
> 9:00am daily team stand-ups in every group, obsessing over being correct and moving fast
> product roadmap (multi-year and annual)
> mountain of planning and frequent engineering design reviews at different maturity milestones
> trade studies exhausting and testing every possible engineering outcome
> subsystem testing at every area of maturity
> team culture of “making the correct decisions” (the company cares deeply about not being wrong)
I’m personally paranoid about executing in the wrong direction - it’s the reason most of all startups fail and its not so obvious in the short term
Hope this helps
I believe that if you understood intelligence, you could build it in software form on a $1M budget, including training. It would not need to be trained on the entire Internet, nor on the thought process of thousands of experts.
Despite all the twitter hype there still hasn't been public proof that the "reasoning" models have any emergence. I.e. is there a class of problems that are solvable with "advanced reasoning" that were not under GPT4o with search under some computational budget?
Everything you love about generative models — now powered by real physics!
Announcing the Genesis project — after a 24-month large-scale research collaboration involving over 20 research labs — a generative physics engine able to generate 4D dynamical worlds powered by a physics simulation platform designed for general-purpose robotics and physical AI applications.
Genesis's physics engine is developed in pure Python, while being 10-80x faster than existing GPU-accelerated stacks like Isaac Gym and MJX. It delivers a simulation speed ~430,000 faster than in real-time, and takes only 26 seconds to train a robotic locomotion policy transferrable to the real world on a single RTX4090 (see tutorial: https://t.co/bEkIlCKqdf).
The Genesis physics engine and simulation platform is fully open source at https://t.co/DhBv7NdyqH. We'll gradually roll out access to our generative framework in the near future.
Genesis implements a unified simulation framework all from scratch, integrating a wide spectrum of state-of-the-art physics solvers, allowing simulation of the whole physical world in a virtual realm with the highest realism.
We aim to build a universal data engine that leverages an upper-level generative framework to autonomously create physical worlds, together with various modes of data, including environments, camera motions, robotic task proposals, reward functions, robot policies, character motions, fully interactive 3D scenes, open-world articulated assets, and more, aiming towards fully automated data generation for robotics, physical AI and other applications.
Open Source Code: https://t.co/DhBv7NdyqH
Project webpage: https://t.co/SBNyhFB0yn
Documentation: https://t.co/3yuBoaealV
1/n