I wrote about the state of AI, why I’m concerned about the next few years, and the choices we need to make to keep the future in humanity’s hands.
An Alien Mind: https://t.co/FeIfWNe0UE
WHO GETS PAID TO TURN AI DATA CENTERS ON
U.S. data centers above 100MW are set to explode this decade but power, interconnects and electrical infrastructure determine how much of that pipeline actually gets built:
What gets the data center online
• $GEV sits at one of hardest bottlenecks with a $176B backlog stretching into 2031 & equipment pricing already more than 20% above Q4 2025
• $VRT captures ~$3.8M per AI MW from grid to rack & is moving further upstream through UtilityInnovation Group adding behind the meter power that can bypass interconnection delays.
• $BE attacks same problem differently with onsite fuel cells letting data centers generate power without waiting years for the grid.
• $FPS sits between the substation & building through switchgear & transformers with utilization still around 30% against capacity being built to support as much as $5B of revenue.
Who already owns the power
• $IREN, $CIFR, $CORZ, $APLD & $WULF already control power and interconnects from their bitcoin infrastructure which is why IREN could go from 5MW of active AI capacity toward a 480MW target in roughly a year since the scarce asset is energized land.
• $VST owns generation already connected to grid giving it direct exposure to rising data center power demand while others wait through interconnection queues.
• $CEG pairs an already interconnected nuclear fleet with hyperscaler demand giving it one of cleanest ways to monetize 24/7 power scarcity through long term contracts.
What gets paid inside the rack
• $ON $NVTS become more important as AI moves toward 800V DC in 2028 since content per rack could rise from ~$15K toward ~$115K as more of the power conversion stack moves inside system.
• $VICR is one of purest plays on converting and delivering power efficiently close to the processor as rack densities rise.
The market can announce hundreds of new data centers but a ton of the value accrues to whoever controls the power, interconnects and electrical infrastructure required to make them operational.
Today we’re introducing a new 3.8 Flash model, our 3rd Flash release in just 6 wks.
It delivers significant leaps from 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning. On DeepSWE v1.1, it outperforms most larger frontier models in autonomously solving complex engineering problems end to end, at a fraction of the cost.
More details here: https://t.co/4YkKkdKQvu
Gavin, spot on.
AI is bringing manufacturing back to America and reindustrializing the nation after decades of offshoring.
AI is creating demand that drives investment in our aging power grid and sustainable energy, powered by market forces, not subsidies.
AI is creating construction and manufacturing jobs across energy plants, chip fabs and data centers.
AI is creating new companies and industries. $400 billion has been invested in AI startups in the past six months alone.
Builders must partner with communities to build in their hometowns, earn trust and create local benefits.
We have an opportunity to create lasting benefits for communities across America and help America lead the next industrial revolution.
a big f*ck you to the U.S in this glm release: "all of this traffic served on chinese AI chips"
zhipu achieved a 3x end-to-end inference performance vs. nvidia chips by improving their software + chip stack.
if i had to guess they're using huawei's latest (china's state mandated all chinese labs to shift inference and eventually training to chinese made chips)
Apple is no longer hiding it. These machines are built for local AI.
Apple unveiled the new Mac Mini with M6 their first 2nm chip. Starting at $899.
What you can run on the $899 Mac Mini (32GB) 👇
→ Qwen 3.6 27B : fits comfortably, near-GPT-4 quality
→ Qwen 3 Coder 32B : full coding assistant on your desk
→ Llama 4 8B at full precision
→ DeepSeek R1 32B for reasoning tasks
→ 13.5x faster LLM processing than M1. 4x faster than M4.
What you can run on the Mac Studio with M5 Ultra (512GB) 👇
→ Llama 4 70B at high quality quantization
→ DeepSeek V4 with hundreds of billions of parameters
→ Frontier-scale models entirely in local memory
→ Connect four Mac Studios over Thunderbolt 5 cluster them to run trillion-parameter models locally
Apple said their Mac business grew 30% last quarter largely because people are buying these to run AI models on their own desks instead of paying for cloud tokens.
The future of inference is local. Apple built the hardware for it.
Introducing GLM-5.3-Flash
- Leading capabilities at a highly competitive price
- Natively multimodal with a 1M-token context window
- A 320B-A18B model released under the MIT License
- Previously previewed as Ox Alpha, running entirely on Chinese AI chips
Blog: https://t.co/tzOmB7gdZP
Available now across all official platforms:
Weights: https://t.co/9LRMahY9Wa
API: https://t.co/VcaQnzYmS9
Coding Plan: https://t.co/Nk8Y98HNhU
ZCode: https://t.co/Peepqv4XSx
Chat: https://t.co/WCqWT0qCQb
AutoClaw: https://t.co/aGEG5HqTTb
Which model should you run on the iPhone 17 Pro?
Announcing our new intelligence and inference testing for small models on mobile devices: independent measurement of how capable small models are in typical on-device tasks, and how they perform on popular phones - in partnership with @liquidai
We have partnered with @liquidai to deliver mobile device inference benchmarking, covering a range of models in 4-bit or lower precision on the iPhone 17 Pro and Galaxy S26 Ultra
We’re publishing our phone-scale intelligence evaluation results in combination with @liquidai's inference performance benchmarks to give users and developers a holistic view of how small models are performing on phones
Inference benchmarking is conducted in a controlled environment using Liquid AI’s inference benchmarking software, which Artificial Analysis has examined and is open-sourced on Liquid AI’s GitHub. Results cover end-to-end generation time, output speed, peak memory usage and other metrics. The inference benchmarking app, ‘Pipette’, is available to download for free on iOS and Android, allowing users to test a variety of models on their own devices
We are ranking phone-scale model intelligence based on each model’s average score in five evaluations chosen for the task-based work these models do in practice: BFCL, IFBench, AA-Omniscience, GPQA Diamond and MATH-500. These evaluations are run by Artificial Analysis using our independent methodology. By default, we limit models to 16K context on each of these evaluations, representing the lack of memory space for significant KV cache on mobile devices. This leads to some intelligent but verbose models dipping in relative score - they were not designed for the constraints that phone memory imposes on token use
We are defining our portable device category as including models that fit within 8 GB of memory after quantization, including KV cache, at 8K context
We expect both the intelligence and inference benchmarks to evolve over time, as new models, devices, inference frameworks and quantization techniques are released. Our pages will remain up to date with these new additions
Initial results:
➤ Nanbeige4.2-3B and LFM2.5-2.6B share the top average evaluation score at 63 (with a 16K context limit), ahead of Ornith-1.0-9B at 62 and Qwen3.5 9B (Reasoning) at 61. LFM2.5-2.6B achieves its score more efficiently: on an iPhone 17 Pro it answers a standard 1,024-token prompt in 8.0s using 2.3 GB of memory, against 21.4s and 4.0 GB for Nanbeige4.2-3B and 25+ seconds and 6.9 GB for the two 9B models
➤ The 16K context limit shapes the leaderboard: as an example, Qwen3.5 9B (Reasoning) spends 74.5M output tokens across one pass of the benchmark set, hitting the 16K limit on 29% of its generations and landing in fourth place overall. With the limit raised to 64K, Ling 3.0 Tiny takes first place with a score of 66, ahead of Nanbeige4.2-3B (65), Qwen3.5 9B (Reasoning, 64), and LFM2.5 2.6B (64). But a 64K window does not fit in mobile phone memory, and at 55 output tokens/s on an iPhone, generating 64K tokens could mean a 20+ minute wait and a lot of battery use. This is why our primary results are capped at 16K, but we're also publishing a set of results capped at 64K, and another capped at one minute of generation time
➤ The speed-intelligence Pareto frontier is short: six models are unbeaten on both intelligence and speed on an iPhone 17 Pro: LFM2.5-230M (27 at 0.9s), MiniCPM5-1B (45 at 2.9s), LFM2.5-8B-A1B (58 at 5.7s), Ling 3.0 Tiny (59 at 5.7s), LFM2.5-2.6B (63 at 8.0s) and Nanbeige4.2-3B (63 at 21.4s). LFM2.5-8B-A1B and Ling 3.0 Tiny are mixture-of-experts models that activate ~1B parameters per token, which is how they answer in under 6s with 8B-class weights
➤ Leading models have opposite strengths: Nanbeige4.2-3B is the most balanced (76% on BFCL, 96% on MATH-500, 67% on GPQA Diamond); Qwen3.5 9B (Non-reasoning) is the strongest tool caller (77% on BFCL) and scientific reasoner (79% on GPQA Diamond); LFM2.5-2.6B follows instructions best of any model measured on the iPhone (59% on IFBench), clears 90% on MATH-500 and hallucinates far less on AA-Omniscience (79% non-hallucination, against 33% for Nanbeige4.2-3B and 24% for Qwen3.5 9B (Reasoning))
More details below in thread ⬇️
Dear Dario,
1. If Claude can cure cancer to save people like your dad, why should we "pace the progress"? Does that mean more people with Hepatitis C will die?
2. If Fable is so cyber-capable that it must be restricted, why are its safeguards too dumb to distinguish cyber defense from cyber offense prompts? When Hugging Face was under attack, why did Fable refuse to help the defenders?
3. We’re glad you want AI to cure cancer. Why is it OK for Claude to force 30-day data retention on pharma's own data and start competing firms but NOT OK (IP theft) if others distill Claude's data?
4. You’ve said advanced models can recognize when they’re being tested and change their behavior accordingly. So why is government(or anyone) able to design the most thorough test before every model launch? Would that just encourage manipulative models?
It seems that every one of your “safety” proposals seems to end the same way: Anthropic gets more leverage, ordinary users get less access, customers pay higher costs, and competitors bear higher regulatory costs, maybe people are not distrusting AI, they are distrusting your approaches with AI.
The Gavin Baker vs. Sholto Douglas exchange on the economic concentration of power of AI is one of the best AI debates. Here's a gist with some context and some thoughts:
One of the central themes of the debate offense-defense balance and how it governs whether AI is (or should be) concentrated in the hands of few or distributed to many.
The core idea behind offense-defense balance, outlined in Robert Jervis's 1978 essay "Cooperation Under the Security Dilemma" on world politics, is: for a given technology, is it cheaper to attack or to defend?
Baker initially argues that he agrees with Zuck's view in his essay that AI should be distributed and not concentrated into a few corporations' hands. Sholto rebuts that it squarely depends on the offense-defense balance of the domain (cyber, bio, etc). And even if you do distribute, the bottleneck is compute.
He cites two domains:
— Cyber he says is defense-dominant, eventually. If it's opened up today, the cost to attack is minimal and the you cannot defend quickly enough. We'll need a ~2yrs of hardening the critical systems like financial institutions and infra providers before this may no longer be a risk.
— Bio he says is offense-dominant well into the 2030s. Synthesizing deadly self-replicating pathogens could cost O($10K) and vaccines cost billions and years. You just can't patch the human body with smarter models.
The distribute everything vs gate everything is a false dichotomy. You want to distribute in defense dominant domains and gate in offense dominant. And its all extra muddied because capabilities are dependent on compute which is extremely unequally distributed. And compute eventually dictates proliferation of capabilities.
______________
Both gloss over some offense dominant domains that are already distributed massively:
— Propaganda and persuasion, for example, were once regarded offense dominant yet already ubiquitous today. That said, technology changes, and watermarking outputs and AI detection tools may have shifted it to being defense dominant.
— Mass surveillance is also offense dominant (usually the attacker is the state) and most LLMs you can use for <$1/M toks can do this today. Today, even non-state actors (private cos) can surveil you in pretty dangerous detail by running agent swarms on vast amounts of public data.
Both Gavin and Sholto are right, and we see it today.
— Sholto's right: the very pinnacle of frontier capabilities is locked away in those who have compute and will continue to be. The shape of compute of the >100 GW compute by 2030 is almost known at this point. Between 6-7 US entities (2 big labs, 3 hyperscalers clouds, Meta/xAI) will own or control 65-75%. With RSI around the corner, as long as it holds true that more compute leads to being able to train and serve more powerful models at scale, those companies will have proprietary frontier capabilities and only they / the market will dictate when they need to release them. Notably, this isn't monopolistic concentration of power since multiple parties will have that capability.
— Gavin's right too: frontier capabilities in certain domains are making it to the public and will continue to do so. Models are getting cheaper. Whether we like it or not, people can do dangerous things with models today. LLMs already cyber exploits that are less talked about (take even the LLM-assisted hack on the nationwide Indian exam system, CBSE). As long as prices for models continue to fall at a more rapid rate (10x YoY) than growth in compute (2-3x YoY), this dynamic will continue. Frontier intelligence will continue to be more accessible.
$MSFT closed below its 200-week moving average in April for the first time in over a decade.
It’s up 35% since then.
As Charlie Munger says, buy exceptional companies below their 200-week moving average and hold forever.
So GLM-5.3 has already found a "potentially serious vulnerability" in Cursor.
GLM-5.3's CyberGym score rose to 84.5%, while ExploitBench more than doubled from 24.4% to 54.4%.
Shows how much more performance a frontier-scale base model can deliver without going through another costly pretraining run.
“Scaling post-training is all we did for GLM-5.3,” Z .ai said in its technical announcement.
Grok 4.6 is roughly the same performance as Fable 5 Max at an 85% discount. 80% cheaper for input tokens and 88% cheaper for output tokens.
Pareto dominant.
Grok 4.7 will be significantly better as is a much larger model with the Cursor and SpaceX data included in pretraining.
Introducing NVIDIA Nemotron 3.5 Lightning⚡
An open 30B MoE model with 3B active parameters, built for always-on agents to complete high-volume, specialized tasks faster.
It delivers up to 4x the output speed of similar-sized models.
This is actually correct.
God knows how many posts I have seen since yesterday comparing $CRWV sequential sales growth vs. debt growth.
They just don’t get fixed assets and scale economics. This is why they missed $AMZN as well.
For $CRWV and all other clouds, they recognize debt upfront and use it to build a fixed base of revenue-generating assets, and then only gradually recognize revenue from these assets over time.
Of course, when you look sequentially, you’ll see faster capex and debt growth and lower revenue growth.
It’s very similar to loan product ramp cycles in fintech, and it’s exactly why those old-schoolers don’t do well in fintech either.