We are open-sourcing our infrastructure for running the MCP Atlas benchmark https://t.co/qOHXEuSuty
This release includes the agent harness we use to score models for the leaderboard, and allows running multiple evals in parallel with your docker image
AI diffusion is far more rate limited by having good evals than most realize. The kind of evals that you see for every model release are incredibly helpful, but only tell you the shape of general AI progress and the relative capability level of models.
The far bigger space over time are evals on all the major workflows that enterprises do, down to the specifics of an individual company.
This will be a huge space over time because you can’t automate what you can’t assess the progress on. Enterprises will not be able to go just on vibes.
the level of schadenfreude around situational awareness / leopold is sad
let the man cook, he's having a generational run and even if he loses 80% of his book, he'll end the year with better numbers than most HF managers will ever print.
I believe the brain of every AI engineer is also an open router when building with claude/codex. I find myself constantly switching between models and reasoning efforts to stay within token budgets, and squeeze out maximum efficiency.
This looks like using Fable 5 with max effort to plan backend + frontend updates, for which it is naturally expected to burn more token to gain 'holistic context', especially when the scope of your app grows beyond a certain point. Components like a cache layer for optimizing conv retrieval, can introduce subtle bugs if feature updates are planned trivially.
And then I would either reduce effort to xhigh or switch to Opus-5 to spawn sub-agents and start implementing the plan.
Switch back to Fable 5 with max effort to test things before shipping the update. Or tag team codex on the same implementation to identify bugs/ enhancements that may be outside the distribution of claude's token output.
It's been great working with the @thinkymachines
team to evaluate tool calling capabilities of the Inkling-small family. Quite insane to see the strides made by the team in such a short period of time! Inkling-small ranks only behind Kimi K3 on open source, and edges out Gemini 3. flash lite and gpt 5.6 luna for long horizon tool calling.
@ziqiao_ma@satu0king
Leaderboard - https://t.co/gpxWNFL87b
Congrats to @thinkymachines on the release of Inkling-small! A smaller variant of Inkling, now live on our AudioMultiChallenge and MCP Atlas leaderboards.
Inkling-small is tied for🥇on AudioMultiChallenge, scoring about the same as the larger Inkling despite the size difference. Strong multi-turn audio performance has typically come from the biggest models, so this is a promising signal for teams building voice applications where latency and cost matter.
Also notable, on MCP Atlas it ranks second among open models on tool calling, behind Kimi K3 and ahead of GLM 5.2. Holding up on both audio reasoning and tool use at this size is a strong showing.
On harnesses, I vacillate between three beliefs:
- the less harness, the better. Models are the magic
- post training a model and harness is dramatically better and the model providers win
- harnesses have real independent value from the model
I have no idea which is right.
There’s something major happening with @OpenAI & @AnthropicAI’s businesses that has huge implications for the compute and memory complex
>Both Anthropic and OpenAI have had explosive revenue growth this year, while at the same time, seeing significant gross margin expansion (see chart)- close to unheard of for companies to see sequential margin expansion while they’re scaling this aggressively
At least so far, it looks like each jump in capability has been worth more to customers than the cost of delivering it. And that gap is getting wider, not narrower… 1/5 🧵
The ads business is changing, especially for startups, SMBs and mid market firms. This plus reports of $META eating into $SHOP’s margins.
Continent creation is not really a choice for most businesses, they NEED to buy ads.
As someone who has personally spent $500k / mo+ on Google Ads for years, I can tell you with certainty:
This revenue growth in Search is artificial & extremely unhealthy for Google’s business long term
Search volumes are declining as legacy search is being increasingly cannibalized by non-monetized LLM queries
Google’s response?
Manufacture revenue growth via short-sighted, highly extractive, customer-hostile tactics. I.e. charge advertisers more for lower quality clicks, including clicks they do not want and explicitly did not approve Google to charge them for
A few examples to illustrate:
For all of its history until recently, Google operated on a 2nd price auction model
I.e. if you bid $5 CPC and the next highest bidder bids $1 CPC, Google charged you $1.01 for the click (one penny more than the 2nd highest bidder) rather than the $5 you bid
This was a genius move by Google early on as it incentivizes advertisers to input their true maximum willingness to pay rather than trying to play the game of bidding low and constantly adjusting to try to stay just ahead of the next highest bidder while still not paying too much
However recently, Google silently deprecated the 2nd price auction and began charging advertisers as much as their bid and budget caps allow, regardless of what anyone else is bidding
It’s a short-sighted cash grab at the expense of the long term health of the advertiser ecosystem
Making thing worse, Google also recently nerfed keyword targeting precision
Google previously had precise keyword targeting settings that allowed advertisers pick individual search phrases to bid on, defined down to the character w/ exact match or phrase match targeting
This was one of the core features that made search advertising magic, enabling advertisers to run extremely precise campaigns based on exactly what their target customer typed
But now, even if you bid on a specific term or phrase using the strictest exact
-match targeting settings, Google will show your ad across 1000’s of unrelated keywords, labeling them as as “exact match (close variant)”
The definition of “close variant” means whatever they want it to and changes constantly. The result is advertisers get billed for clicks that are totally irrelevant to their business and that their targeting settings explicitly forbid Google from targeting. Google does it anyway and there’s no ability to turn this off
So now exact match is broad match, and broad match is just meaningless spam
This is all very bad for advertisers, but for Google, it allows them to show your ad and bill you for clicks across 1000x more searches that were previously going unmonetized (mainly because they’re garbage queries no one wants)
This is how you grow revenue atop declining search volumes
Lastly, and perhaps most egregiously, Google quietly stopped respecting budget caps by a factor of 2x. For example campaigns we’ve been running for years with $1000 daily budget caps suddenly began spending $2000+ per day
And the extra spend is entirely on the garbage keywords Google arbitrarily throws in as “exact match (close variants)” which have no value to our business, but can’t be turned off
Google offers no refunds nor any recourse for overspend or spend on keywords you explicitly did not target
These are not the actions of a healthy business. These are the actions of company whose core business is in decline but desperately needs to pump quarterly earnings so Wall Street will continue to fund insane capex while hopefully looking through their rapidly deteriorating negative free cash flow
Google operated a benevolent monopoly for the better part of 25 yrs
Meaning the value Google captured from Search was but a small fraction of the value it created, and that spread produced a potential energy that justified expectations of high earnings growth far, far into the future
This is now no longer the case
At the alter of AI capex, Google is sacrificing the golden goose
As someone who has personally spent $500k / mo+ on Google Ads for years, I can tell you with certainty:
This revenue growth in Search is artificial & extremely unhealthy for Google’s business long term
Search volumes are declining as legacy search is being increasingly cannibalized by non-monetized LLM queries
Google’s response?
Manufacture revenue growth via short-sighted, highly extractive, customer-hostile tactics. I.e. charge advertisers more for lower quality clicks, including clicks they do not want and explicitly did not approve Google to charge them for
A few examples to illustrate:
For all of its history until recently, Google operated on a 2nd price auction model
I.e. if you bid $5 CPC and the next highest bidder bids $1 CPC, Google charged you $1.01 for the click (one penny more than the 2nd highest bidder) rather than the $5 you bid
This was a genius move by Google early on as it incentivizes advertisers to input their true maximum willingness to pay rather than trying to play the game of bidding low and constantly adjusting to try to stay just ahead of the next highest bidder while still not paying too much
However recently, Google silently deprecated the 2nd price auction and began charging advertisers as much as their bid and budget caps allow, regardless of what anyone else is bidding
It’s a short-sighted cash grab at the expense of the long term health of the advertiser ecosystem
Making thing worse, Google also recently nerfed keyword targeting precision
Google previously had precise keyword targeting settings that allowed advertisers pick individual search phrases to bid on, defined down to the character w/ exact match or phrase match targeting
This was one of the core features that made search advertising magic, enabling advertisers to run extremely precise campaigns based on exactly what their target customer typed
But now, even if you bid on a specific term or phrase using the strictest exact
-match targeting settings, Google will show your ad across 1000’s of unrelated keywords, labeling them as as “exact match (close variant)”
The definition of “close variant” means whatever they want it to and changes constantly. The result is advertisers get billed for clicks that are totally irrelevant to their business and that their targeting settings explicitly forbid Google from targeting. Google does it anyway and there’s no ability to turn this off
So now exact match is broad match, and broad match is just meaningless spam
This is all very bad for advertisers, but for Google, it allows them to show your ad and bill you for clicks across 1000x more searches that were previously going unmonetized (mainly because they’re garbage queries no one wants)
This is how you grow revenue atop declining search volumes
Lastly, and perhaps most egregiously, Google quietly stopped respecting budget caps by a factor of 2x. For example campaigns we’ve been running for years with $1000 daily budget caps suddenly began spending $2000+ per day
And the extra spend is entirely on the garbage keywords Google arbitrarily throws in as “exact match (close variants)” which have no value to our business, but can’t be turned off
Google offers no refunds nor any recourse for overspend or spend on keywords you explicitly did not target
These are not the actions of a healthy business. These are the actions of company whose core business is in decline but desperately needs to pump quarterly earnings so Wall Street will continue to fund insane capex while hopefully looking through their rapidly deteriorating negative free cash flow
Google operated a benevolent monopoly for the better part of 25 yrs
Meaning the value Google captured from Search was but a small fraction of the value it created, and that spread produced a potential energy that justified expectations of high earnings growth far, far into the future
This is now no longer the case
At the alter of AI capex, Google is sacrificing the golden goose
We're partnering with @huggingface to investigate an unprecedented security incident.
Cyber-capable OpenAI models compromised Hugging Face production during a benchmark evaluation.
Sharing preliminary findings to help defenders understand emerging risks:
https://t.co/CIor15y9xk
A year ago, while writing my PhD thesis, I realized one important piece was still missing:
How do we make benchmark datasets useful beyond producing a score?
A few months ago, I got the chance to work on this. CRAFT feels like a natural continuation of my PhD, and a meaningful step toward answering that question.
Some observations on Kimi:
1. It's a very good model! I don't think its performance can be explained away by distillation or anything like that. In agentic coding sessions, it seems pretty much on par with the best public models of Q1 2026. In my fairly limited use, it also seemed very token hungry. It's not obvious to me that this model is actually that cheap to run.
2. I am personally surprised the Chinese state continues to allow the open sourcing of models this good, given potential risks. To be clear, I *myself* might be fine with models presenting this level of marginal risk being open weight, but I am surprised that China is fine with it. I suspect the reason they are is 75% explained by strategic blindness/lack of AGI-pilledness (the CCP is very Yann Lecun-y in its views of AI). The other 25% or so is their lack of compute for customer inference (making China's open-weight strategy an unintended byproduct of US export controls) and the normal Chinese strategy of aggressive exports. For the companies, as opposed to the government, the decision to open source is partially ideological and partially because they are behind, and they know that very few people would pay for sub-frontier models from China.
3. Open-weight models are inherently decelerationist, and I'm continually surprised to see the so-called "accelerationists" so excited about open-weight models. I suspect the reason they are is that they know open-weight models are effectively ungovernable, and they simply like the overall cloak of ungovernability open-weight models create over the whole of AI. It's not a bad strategy; it reminds me of James Scott's recounting of the hill people in "the art of not being governed." Still, in the end, open-weight models deter further AI capex.
4. One probable outcome of an open-weight-model-dominant world is full AI communism, which is precisely what China proposes: rather than a market product, AI is a "public good" which will ultimately be provided by the state as a kind of "digital public infrastructure." This future strikes me as a dystopian hellscape, but I've never met an open-weight models advocate who doesn't ultimately concede this is where things end. You'd be surprised how many 'accelerationists' lobbied me, while I was in government, to support an eleven or twelve-figure federally funded data center so that startups could train models at a subsidy and then give them away for free. There was no other way for AI to progress, they said. Perhaps this is the logical end state of things. Nonetheless, I find myself surprised to see supposed accelerationists excited about such an outcome. I think many of them just don't know what they're doing. Many accelerationists do not view the creation and serving of frontier models as a legitimate business.
5. I would guess that the Trump Administration will at some point realize that their best strategy here would be to create large amounts of regulatory risk around the use of open-weight Chinese models. You don't need to "ban open source" (one of the dumber motifs of AI policy discussion). You just need to direct every agency to issue soft law that creates FUD. "A Federal Reserve Advisory Bulletin found that there may be backdoors in Chinese AI models." It needn't be that well justified. You just create enough regulatory risk that every regulated enterprise backs off. You probably don't want to create so much regulatory risk that you scare off the hyperscalers from serving Chinese models; this will just drive startups to sketchier providers. There's a happy middle ground here. I'd assume they will do some version of this.
6. It's probably true that open-weight models of this capability make the world a bit more dangerous, but not so much more that you'll really notice. At some point the models will be capable enough that you will notice. "A nonliving, invisible, dangerous, and infinitely self-replicating agent escaped from a Chinese lab," you say? Color me shocked.
This is a neat RL trick from Thinking Machines Inkling.
Instead of optimizing only for task success, they optimize:
Reward = Task Reward − λ × (# reasoning tokens)
Then they vary λ across rollouts and pair it with different effort instructions.
The result: model learns that reasoning is a resource to spend, not something to maximize.
Great to work together with the @thinkymachines team over the past week to report MCP Atlas scores on Inkling.
Check out the benchmark and our paper:
Benchmark - https://t.co/F4fEhvqQO2
Paper - https://t.co/DFh5skKS5Y
Congrats to @thinkymachines on the release of their open weight model Inkling! We were proud to work with their incredible team on preparing this model for release for the past several months.
Now live on our MCP Atlas and AudioMultiChallenge leaderboards. Inkling tied for 🥇 on AudioMultiChallenge, surpassing Gemini 3 Pro as the de facto frontier model that supports native audio input.
Also notable, on MCP Atlas Inkling had a low hallucination rate compared to other frontier models.
The personally rewarding part was to work on the failure mode taxonomy for this domain of biomedical agents, and performing root cause diagnostics on agent trajectories. GPT-5.5 led at 51.6%, with Gemini 3.5 Flash and Claude Opus 4.8 close behind.
New @ScaleAILabs benchmark: 𝐃𝐫𝐮𝐠𝐃𝐢𝐬𝐜𝐨𝐯𝐞𝐫𝐲𝐁𝐞𝐧𝐜𝐡! 💊🧬👩🔬
We built it with @phylo_bio to evaluate agents on early drug discovery workflows.
Bringing a drug to patients can take decades and billions of dollars. In a typical campaign, thousands of compounds are tested, and early decisions shape everything downstream.
New @ScaleAILabs benchmark: 𝐃𝐫𝐮𝐠𝐃𝐢𝐬𝐜𝐨𝐯𝐞𝐫𝐲𝐁𝐞𝐧𝐜𝐡! 💊🧬👩🔬
We built it with @phylo_bio to evaluate agents on early drug discovery workflows.
Bringing a drug to patients can take decades and billions of dollars. In a typical campaign, thousands of compounds are tested, and early decisions shape everything downstream.