I’ve been using 5.1 Thinking as my default for day2day work, but 5.2 Thinking has been a step backward in my experience. It takes longer to respond, tends to wander, and frequently over-delivers beyond what I asked for. Additionally, the added content isn’t high-quality. 5.2 is surprisingly better in that regard so I've been finding I'm using it more than Thinking.
xAI gave us early access to Grok 4 - and the results are in. Grok 4 is now the leading AI model.
We have run our full suite of benchmarks and Grok 4 achieves an Artificial Analysis Intelligence Index of 73, ahead of OpenAI o3 at 70, Google Gemini 2.5 Pro at 70, Anthropic Claude 4 Opus at 64 and DeepSeek R1 0528 at 68. Full results breakdown below.
This is the first time that @elonmusk's @xai has the lead the AI frontier. Grok 3 scored competitively with the latest models from OpenAI, Anthropic and Google - but Grok 4 is the first time that our Intelligence Index has shown xAI in first place.
We tested Grok 4 via the xAI API. The version of Grok 4 deployed for use on X/Twitter may be different to the model available via API. Consumer application versions of LLMs typically have instructions and logic around the models that can change style and behavior.
Grok 4 is a reasoning model, meaning it ‘thinks’ before answering. The xAI API does not share reasoning tokens generated by the model.
Grok 4’s pricing is equivalent to Grok 3 at $3/$15 per 1M input/output tokens ($0.75 per 1M cached input tokens). The per-token pricing is identical to Claude 4 Sonnet, but more expensive than Gemini 2.5 Pro ($1.25/$10, for <200K input tokens) and o3 ($2/$8, after recent price decrease). We expect Grok 4 to be available via the xAI API, via the Grok chatbot on X, and potentially via Microsoft Azure AI Foundry (Grok 3 and Grok 3 mini are currently available on Azure).
Key benchmarking results:
➤ Grok 4 leads in not only our Artificial Analysis Intelligence Index but also our Coding Index (LiveCodeBench & SciCode) and Math Index (AIME24 & MATH-500)
➤ All-time high score in GPQA Diamond of 88%, representing a leap from Gemini 2.5 Pro’s previous record of 84%
➤ All-time high score in Humanity’s Last Exam of 24%, beating Gemini 2.5 Pro’s previous all-time high score of 21%. Note that our benchmark suite uses the original HLE dataset (Jan '25) and runs the text-only subset with no tools
➤ Joint highest score for MMLU-Pro and AIME 2024 of 87% and 94% respectively
➤ Speed: 75 output tokens/s, slower than o3 (188 tokens/s), Gemini 2.5 Pro (142 tokens/s), Claude 4 Sonnet Thinking (85 tokens/s) but faster than Claude 4 Opus Thinking (66 tokens/s)
Other key information:
➤ 256k token context window. This is below Gemini 2.5 Pro’s context window of 1 million tokens, but ahead of Claude 4 Sonnet and Claude 4 Opus (200k tokens), o3 (200k tokens) and R1 0528 (128k tokens)
➤ Supports text and image input
➤ Supports function calling and structured outputs
See below for further analysis 👇
Last week I spoke to the Columbia entrepreneurship group about building and scaling startups.
My path to CPO was all but linear. I made tons of mistakes, got stuck, and experienced many failures.
Here are 7 principles I wish I knew when I was in college that I learned along the way.
Innovation doesn’t start with spreadsheets; it starts with stories. Data reflects reality, but it’s not reality. To uncover true opportunities, immerse yourself in the messy context, ask why, and solve the puzzle behind the 'Job to Be Done'.
As organizations grow, it’s easy to become overly focused on the data and lose sight of the actual jobs we’re solving for our customers.
And let’s be honest, data has a way of bending to support whatever narrative we want it to.
As @NateSilver538 has said, "The most calamitous failures of prediction usually have a lot in common. We focus on those signals that tell a story about the world as we would like it to be, not how it really is."
Even with the best intentions, companies often drift into thinking their business is defined by the products and services they sell when in reality, it’s defined by the jobs they’re solving.
To build meaningful solutions, we need to keep our focus where it matters most. Focusing not just on the metrics, but on the people behind them. It’s the customer’s struggle that should guide our innovation.
Introducing our first set of Llama 4 models!
We’ve been hard at work doing a complete re-design of the Llama series. I’m so excited to share it with the world today and mark another major milestone for the Llama herd as we release the *first* open source models in the Llama 4 collection 🦙. Here are some highlights:
📌 The Llama series have been re-designed to use state of the art mixture-of-experts (MoE) architecture and natively trained with multimodality. We’re dropping Llama 4 Scout & Llama 4 Maverick, and previewing Llama 4 Behemoth.
📌 Llama 4 Scout is highest performing small model with 17B activated parameters with 16 experts. It’s crazy fast, natively multimodal, and very smart. It achieves an industry leading 10M+ token context window and can also run on a single GPU!
📌 Llama 4 Maverick is the best multimodal model in its class, beating GPT-4o and Gemini 2.0 Flash across a broad range of widely reported benchmarks, while achieving comparable results to the new DeepSeek v3 on reasoning and coding – at less than half the active parameters. It offers a best-in-class performance to cost ratio with an experimental chat version scoring ELO of 1417 on LMArena. It can also run on a single host!
📌 Previewing Llama 4 Behemoth, our most powerful model yet and among the world’s smartest LLMs. Llama 4 Behemoth outperforms GPT4.5, Claude Sonnet 3.7, and Gemini 2.0 Pro on several STEM benchmarks. Llama 4 Behemoth is still training, and we’re excited to share more details about it even while it’s still in flight.
A big thanks to all of our launch partners (full list in blog) for helping us bring Llama 4 to developers everywhere including @huggingface, @togethercompute, @SnowflakeDB, @ollama, @databricks and many others👏 This is just the start, we have more models coming and the team is really cooking – look out for Llama 4 Reasoning 😉
A few weeks ago, we celebrated Llama being downloaded over 1 billion times. Llama 4 demonstrates our long-term commitment to open source AI, the entire open source AI community, and our unwavering belief that open systems will produce the best small, mid-size and soon frontier models. Llama would be nothing without the global open source AI community & we are so ready to begin this next chapter with you. 🦙
Read more about the release here: https://t.co/7mbK3uggjO, and try it in our products today.
AI OVERPRODUCTION
China seeks to commoditize their complements. So, over the following months, I expect a complete blitz of Chinese open-source AI models for everything from computer vision to robotics to image generation.
Why? I’m just inferring this from public statements, but their apparent goal is to take the profit out of AI software since they make money on AI-enabled hardware. Basically, they want to do to US tech (the last stronghold) what they already did to US manufacturing. Namely: copy it, optimize it, scale it, then wreck the Western original with low prices.
I don’t know if they’ll succeed.
But here’s the logic:
(1) First, China noticed that DeepSeek’s release temporarily knocked ~$1T off US tech market caps.
(2) Second, China’s core competency is exporting physical widgets, more than it is software.
(3) Third, China’s other core competency is exporting things at such massive scale that all foreign producers are bankrupted and they win the market. See what they’re doing to German and Japanese cars, for example.
(4) Fourth, China is well aware that it lacks global prestige as it’s historically been a copycat. With DeepSeek, becoming #1 in AI is now something they actually consider possibly achievable, and a matter of national pride.
(5) Fifth, DeepSeek has gone viral in China and its open source nature means that everyone can rapidly integrate it, down to the level of local officials and obscure companies. And they are doing so, and posting the results for praise on WeChat.
(6) Finally, while DeepSeek was obscure before recent events, it’s now a household name, and the founder (Liang Wengfeng) has met both with Xi but also the #2 in China, Li Qiang. They likely have unlimited resources now.
So, if you put all that together, China thinks it has an opportunity to hit US tech companies, boost its prestige, help its internal economy, and take the margins out of AI software globally (at least at the model level).
They will instead make their money by selling inexpensive AI-enabled hardware of increasing quality, from smart homes and self-driving cars to consumer drones and robot dogs.
Basically, China is trying to do to AI what they always do: study, copy, optimize, and then bankrupt everyone with low prices and enormous scale.
I don’t know if they’ll succeed at the app layer. But it could be hard for closed-source AI model developers to recoup the high fixed costs associated with training state-of-the-art models when great open source models are available.
Last, I agree it’s surprising that the country of the Great Firewall is suddenly the country of open source AI. But it is consistent in a different way, which is that China is just focused on doing whatever it takes to win — even to the point of copying partially-abandoned Western values like open source, which seemed like the hardest thing to adopt.
On that point: they did build censorship into the released DeepSeek AI models, but in a manner that’s easily circumvented outside China. So, you might conclude they don’t really care what non-Chinese people are saying outside China in other languages, so long as this doesn’t “interfere with China’s internal affairs.”
Anyway —this is an area I’ve been watching, and my reluctant conclusion is that China is getting better at software faster than the West is getting better at hardware.