Another remarkable DeepSeek innovation slashes memory use per agentic token even as it hugs massively more expensive closed frontier...who needs a CFO or IR when you have full stack engineering at this level?
We benchmarked DeepSeek V4.1 Flash by @deepseek_ai .
It reached 98% of GPT-6 Astra’s score at 1.4% of the cost on everyday design tasks based on user requests.
Every model except Astra scored lower AND cost more.
Are open models overtaking closed ones?
Full results below ↘️
I think it’s due to 3 core reasons,
1. The Chinese political leadership sees AI as a process, one that will generally diffuse to actual industrial machinery and robotics, and less as an EVENT. Americans have imbued the general AI revolution with now religious overtones(the singularity), and recursive ASI as one massive day of reckoning that changes everything. So it’s also a broader metaphysical difference between east and west.
2. The Compute and Hardware gap, China wants to completely master advanced chips, lithography, and domestic semiconductor tooling before spacing out about models and their total capabilities, and I think they plan to stay on this path as they believe this is where sovereignty actually lies, let it also not be forgotten that via rare earth and certain tooling exports that they could also suffocate the American buildout, so China has literal Trump cards it can still present to Trump lmao.
3. It’s just not CPC culture to literally freak the fuck out about technology or geopolitical competition to the public at large. This is a political difference.
Some will disagree with me, but I am tempted to believe the CPC once again is playing this game entirely right, and when there is a likely massive Capex bubble burst in a year to two in the West, they will look very prescient.
Two extraordinary charts from today’s PISA test results:
1) School test scores continue to collapse internationally, underscoring how this is no longer a Covid effect but sustained decline.
Those falls in reading and maths are equivalent to about two years of lost schooling.
This humanoid brain issue being resolved without the need for a new class of AI 'world models' for perception/motion control would be very bullish for the hardware supply chain from actuators, motors and magnets to batteries and power semis...but almost all in Asia
LLM token output speed increases by 2-7x per year, with Fable-class models doubling every month.
If trends continue, LLMs could meaningfully control robots in real time by end of the year, or by 2029 at the latest.
@ParadisLabs Without access to Chinese rare earth magnets, all the semis in the world can't build these actuator and motor intensive robots at any scale...beyond mining, Western supply chain (mostly Japan, Shin Etsu etc) is adequate to supply about 1% of these projections
Great deep dive engineering analysis on how OpenAI made its Jalapeno custom silicon breakthrough...AI disruption has reached semi compiler design and Nvidia's CUDA moat has sprung a leak
This paper is a brutal reality check for long-horizon AI. Give an agent a year of interconnected decisions, delayed feedback, and consequences from its own past actions, and its performance collapses relative to humans.
The researchers tested eight leading models, including GPT-5.6 Sol and Claude Opus 4.8. Yet the best-performing setup, Qwen3.7-Max with Hermes, ended with only 27.3% as much money as the average human participant.
A system that finishes a year-long task at barely a quarter of human performance is nowhere near dependable long-horizon execution.
ox-Alpha is an exceptional example of why Chinese labs are so impressive with tiny budgets.
Looks how fast ZAI adopted innovations by their peers!!!
1.) They adopted DeepSeek-V4's mHC based residual connection and
2.) also, DeepSeek's DSA like Sparse Attention
3.) also, MoonShot's KDA type linear attention
Result:
1.) Strong performance with 4.44 times less KV cache
2.) 3 times less flops
3.) Kick ass model that can be served 10 times cheaply even compared to their own GLM-5.3
Every lab builds on the innovations by other, so they don't have to repeat all the same experiments themselves. This is extremely economically efficient.
🎙️Full tech report on how to actually achieve SovereignAI, from data to model training, values, infrastructure, to large-scale deep research.
Bridging up to 7 months of Frontier AI development with a $450k training Continual Learning run 🧠
📜 Paper: https://t.co/ooEhgoF0lU
💻Model: https://t.co/DgcWLuqnQX
Partners: @imperialcollege, @datologyai, @LambdaAPI
People had some aggression towards @benthompson for calling out very valid points. Who else has as good of a track record of commentating in public on tech company strategy? Good luck finding them. Great job @InvestLikeBest@patrick_oshag
>Capital is the binding constraint he watches. Hyperscalers have already moved from free cash flow to debt and, in Google’s case, equity issuance. Nvidia is structuring vehicles aimed at pension and insurance capital. The buildout can stay technologically rational while running out of investors willing to finance the next increment before returns arrive.
>Today’s compute shortage was locked in by underinvestment decisions in 2023–2025 @ TSMC. Most of the capital being committed now only becomes usable chips and operating data centers in 2028–2029. Scarcity-era utilization and pricing are therefore being used to underwrite assets that will enter a potentially more abundant market. Next few years looks WORSE and shortages MORE ACUTE before it gets better.
>Commodity-market mechanics dominate and tech investors/participants do not understand them (or refuse to!). Data centers, ships, memory fabs, and other high fixed-cost assets keep operating as long as revenue covers marginal cost (fuel, crew, power, port fees). Scarcity attracts new capacity with a multi-year lag; when that delayed fleet arrives, prices can collapse even if the original capital never earns an adequate return.
>TSMC’s multi-decade fab discipline protects its own economics but pushes shortage risk onto customers ("risk does not disappear, it moves somewhere else"). Morris Chang’s decision to invest through the financial crisis after recognizing the iPhone opportunity helped create TSMC’s dominance. Acute enough scarcity can finally make the cost and pain of qualifying Intel or Samsung rational for hyperscalers, delivering geopolitical redundancy as a byproduct of commercial necessity.
>Meta’s advertising business is one of AI’s cleanest existing monetization engines. It already runs a global, real-time verification loop for generated creative and prediction. Small improvements in matching and conversion can be worth billions of dollars; Thompson believes the company still under-explains this structural advantage.
>Nvidia’s reported chip margins do not capture the full economics. Guarantees, equity investments, purchase commitments, and customer financing support function like price concessions. Nvidia accepts risk outside the product gross-margin line in order to keep GPU demand flowing.
>The durable residue of an AI overbuild is likely power. GPUs age quickly and data-center economics can reset. Added generation, transmission, and grid capacity can lower constraints across the broader economy for decades after the compute assets themselves have been written down.
*not touching the China-TSMC-proactive strike if we get out ahead in AI implications part - beyond my expertise.
42 GW worth of chips could be stuck in inventory by 2030, according to a recent BloombergNEF report.
This outcome seems almost impossible for the tech industry to avoid. Chips are made in factories, while they need to be powered in physical infrastructure that is very difficult to permit/build.
The same dynamic played out in China’s solar industry. Manufacturers scaled up production dramatically. But global demand for panels was throttled by permitting, labor, etc. As a result there was a huge glut of panels.
This isn’t just a local opposition problem for data center developers. There are plenty of other constraints that could throttle data center capacity: skilled labor, specialized power/cooling equipment, etc.
Launching data centers into space could theoretically solve this. But going from no orbital data center capacity to a few rocket launches per day as @elonmusk proposed on the @dwarkesh_sp podcast also seems pretty hard!
The median company is spending $12 / employee / month on AI
The top 1% are spending $7,500 / employee / month
Not sure we've ever seen an adoption gap quite like this
(h/t @tryramp data, @a16z)
Great post both if you’re driving AI in an enterprise or building for an enterprise.
AI productivity gains are going to wildly vary - much wider than you think - because what you can do at the frontier is so significant if you fundamentally change your workflows to support agents. The challenge is that most people won’t or can’t naturally do this because of the complexity it entails.
Thus, a large amount of automation with agents is going to come from just getting agents wired up into a workflow that the user never ends up noticing or caring about:
“For everybody else, the work has to get done without them changing how they work, which means the AI goes into the background. Why do you need AI to be prompted by a human anyway? Just figure out what the most repetitive processes are, build agents in your existing systems of record that employees are used to”
This is the real work ahead for most enterprises right now that want to get significant gains from AI. A lot less “let a thousand flowers bloom” and more “pick off the 10 highest leverage things in the enterprise and apply automation to them”. There’s no real shortcut here.
I visit China frequently. This is the best roundup of China's rise I've read. "A report by researchers identified almost twenty thousand scientists of Chinese descent who left America between 2010 and 2021." https://t.co/osi9rUr3Dm
An important signal for $MSFT was that product velocity and quality improved several months ago, while the stock completely ignored it. Agent Studio release today continues the strong product cadence.
Anthropic will probably be the primary enterprise agent beneficiary (best models), but after that I think $MSFT has the strongest hand, with access to the best models (OpenAI for free, Anthropic partnership) without model lock-in, integration into Office apps, and all of the critical “plumbing” that enterprises care abt to actually make agents that can run efficiently and securely in the real world.
NVIDIA and Huawei have reached the same conclusion: competitive advantage in AI hardware is migrating from the chip to the system. Their engineering answers point in opposite directions.
NVIDIA packs 72 top-tier GPUs into one rack and lets dense silicon do the work. Huawei spreads 384 weaker processors across a wider fabric and lets system architecture compensate. SemiAnalysis assessed Huawei's design as "arguably a generation ahead" in system architecture, at an estimated 2.5 times worse energy efficiency per FLOP.
Two bets, each exposed to a different failure mode. My latest analysis maps where each one breaks: https://t.co/Or6133RYGW
The 'whole of economy' AI adoption pace in China is now world leading by a huge margin, reflected in mainland token volumes...frontier open source intelligence, massive scale economies and experimental iteration at the application layer are about to make it a global software power challenging the US
The most interesting AI adoption data out of China this week comes from a flea market.
Xianyu, Alibaba's secondhand marketplace, released first-half numbers today: 9.8 million AI service orders, up 157% from a year earlier. The fastest-growing category is AI coding and website building, up 1,732%.
62% of the sellers are women. Average monthly sales are about $130 (Rmb 897), and most treat it as a side hustle. The demand is not coastal either: buyers in tier-2 cities now outnumber those in tier-1.
My read: this is diffusion data you cannot get from enterprise surveys. When AI skills become a tradable commodity on a secondhand app, priced at impulse-buy levels, adoption has moved past the early-adopter crowd. The frontier debate is about model capability. The volume story is a woman in a smaller city selling AI tutorials after work.
Is any Western marketplace producing comparable numbers? Fiverr is the closest analogue I can think of, but I have not seen AI gig data at this scale from it. If you have, I want to see it.
Excited to see Anthropic acknowledge the real problem with open is it competes with their corporate economic strategy. The “awareness” here is revealing.
When you build the biggest private market cap ever, you should expect competition. I’m sure the secondaries have been nice.
Companies ex tech have only just started integrating AI into proprietary databases via tools like MCP, while hardly any are yet reorganising around digital workers as a default e.g. changing job roles to task flows...the employment shock will arrive, but the timeline looks more mid late 30s, not 2030 as Anthropic's Amodei claimed. It's all a reminder of the conservatism of the corporate world ex Silicon Valley, where 'moving fast and breaking things' with radical experimntation is liable to blow up a longstanding business