I looked at 13 different providers for even 1 node of B200/B200s while I wait for my order to get delivered.
Zero availability.
I’ve never seen GPU capacity scarcity like this. Prices are also headed towards $6.50-7/gpu/hr. Expect inference to get more expensive.
From CNBC:
In an internal meeting with employees on Wednesday, finance chief Sarah Friar and board chair Bret Taylor touted OpenAI’s revenue growth and addressed competition with Anthropic, CNBC has learned. Friar said OpenAI’s annualized recurring revenue in July exceeded the entire second quarter.
“And Q2 was no slouch,” Friar said, according to a partial transcript of the meeting that was reviewed by CNBC.
Friar and Taylor said momentum was driven by the release of the company’s GPT-5.6 series of models, its new enterprise agent called ChatGPT Work, and growing adoption of its AI coding tool, Codex
Despite major launches from 5+ labs this month, OpenAI occupies most of the token efficiency Pareto frontier
We measure the number of output tokens models produce per task in the Artificial Analysis Intelligence Index. Output tokens consist of answer tokens (can be thought of as how verbose the model is) and reasoning tokens (how much the model thinks before giving an answer). Reasoning tokens in particular offer a way for models to use compute at inference time to improve responses.
Output tokens are an important determinant of both cost and time per task. Various effort levels of GPT-5.6 Sol dominate the frontier - Terra and Luna produce comparatively more tokens for any level of intelligence.
Anthropic remains well ahead in ARR, but OpenAI has recently accelerated.
TickerTrends is now tracking Anthropic at $74.1B versus OpenAI at $41.3B, with OpenAI’s run rate rising from $33.0B in May to $41.3B in July.
The first-ever measured silicon numbers for @NVIDIA Vera Rubin NVL72 are in 😲
First measured performance shows 10x more tokens per megawatt than Blackwell.
No projections. Real results from live hardware.
From our data checks, @AnthropicAI % spend from developers/engineers came down in June, while @OpenAI spend went up. This is contrary to much of the 3rd party data out there. If this is true, on the margin, it benefits $MSFT
Kimi K3 has received far more love than we expected, and our GPUs are feeling it.
Over the past 48 hours, demand has pushed close to the limits of our current capacity. To protect the experience of existing subscribers, we're temporarily pausing new subscriptions and prioritizing compute for current members. Existing subscribed users are not affected.
We're adding capacity as fast as we can and will reopen new subscription spots in batches.
Going forward, we'll also split membership into two more focused plans: Kimi Membership for Kimi Web, App, and Work; and Kimi Code Membership for coding workflows. This will help us match compute more precisely and keep the experience stable.
Thank you for your patience and understanding!
Introducing a limited preview of GPT-5.6 Sol, our next generation frontier model, as well as GPT-5.6 Terra, a balanced model for efficient, everyday work, and GPT-5.6 Luna, a fast and affordable model for high-volume work.
https://t.co/OoM83SyISN
Show Codex a workflow once. Reuse it as a skill.
Record & Replay lets you show Codex a recurring task, like filing an expense report or submitting a time-off request.
Codex turns that demo into an inspectable, editable skill.
You control when recording starts and stops.
Interestingly, Google no longer has a public frontier model. They have a very good flash model, but a very good flash model can't do frontier work without a good frontier orchestrator.
I am sure this will change soon, but Gemini 3.1 Pro is very clearly lagging at this point.
I’m excited to share that I’ll be joining OpenAI and look forward to working with the exceptional team there.
It was a difficult decision to move on. I’m incredibly proud of the amazing team at Google and everything we’ve built together. It has been an honor and a pleasure to work with all of you.