The AI compute flywheel we are building at YAI breaks down into three connected sides:
1. Hardware owners plug in idle GPUs to run real, verified AI tasks and earn steady yield.
2. Builders and teams consume that capacity through a low-cost, secure API with zero data logging and complete privacy.
3. Users and businesses manage multi-agent workflows in our smart workspace with strict budget caps.
Supply earns yield, builders cut compute bills with private execution, and the chat acts as the live proving ground.
If you are at @token2049 Singapore and tired of paying absurd cloud bills for your model calls, keep an eye out for us.
Our team will be roaming the floor looking for companies to become our initial design partners. You get full access to our compute layer, legacy status, and our undivided engineering focus.
Drop a DM if you want to grab coffee, compare benchmarks, or just talk shop.
Spot the tee with the [Y].
A lot of teams discover mid-sprint that their inference provider logs prompts by default.
Then they go hunting through docs for an opt-out that may or may not exist.
YAI retains nothing after the response is delivered, that's the architecture, not a settings toggle.
Model competition prevents one form of concentration. It does not solve the next one.
Even with open models, inference can still depend on a small number of centralized providers.
An open model layer needs an open execution layer beneath it.
Kimi K3 may be an important inflection point for AI. Potentially negative for Anthropic and OpenAI while being net positive for essentially every other company in the world. I mean that very literally. Although the real “Sputnik moment” would be an open-source frontier model that was also token efficient unlike Kimi K3 which is 50-70% more expensive to run than GPT 5.6 per Artificial Analysis.
Rationale:
A world where there are only 2-3 dominant frontier labs with 90% inference margins is net negative for every other layer while being awesome for those 2-3 labs. Those labs would become monopsonies for power, data centers, semiconductors and hyperscalers and would obviously vertically integrate over time into all those layers while also completely subsuming the application/software layers.
Anything that lowers margins and increases competition at the model layer is good for every other AI layer: power, semiconductors, hyperscalers, neoclouds and yes even software.
This is why Jensen is so supportive of open-source. An open-source model requires the *exact* same amount of compute to run as a closed frontier model of similar size and architecture. Kimi K3 is roughly the same price as GPT 5.6 Terra on a per token basis, which actually suggests that it is less computationally efficient as I am sure that GPT 5.6 is priced to a higher margin than K3. And given that K3 is a token wastrel, i.e. token inefficient, it is significantly more expensive per task than GPT 5.6 and Grok 4.5, which are much more token efficient. Cost per token and token efficiency (i.e. intelligence density per token) are the drivers of intelligence per unit of cost. The winning AI companies will be those that offer the most intelligence per $ over time.
Lower margin % at the model layer = more margin $ at every part of the infrastructure layer and is a godsend for software. This can happen either through open-source models like K3 at the frontier *or* having a vertically integrated model company like Meta, SpaceX or Google at the frontier. Both outcomes result in a lower margin % at the model layer as vertically integrated model companies don’t really care where the margin $ come from. This is why it was so painful for OpenAI and Anthropic when Google was right there with them from a model competitiveness perspective and why Grok 4.5 and Muse 1.1 were just as important as Kimi K3.
The reason Kimi K3 is only *potentially* negative for Anthropic and OpenAI is 1) the @ericvishria point that the Claude and ChatGPT products and harnesses may be more important than their models today and 2) the hypothesis that they have much more advanced model checkpoints internally that are already being used for RSI. In the latter scenario, reaching RSI even a few months ahead of other labs might be enough to cement a permanent lead.
Time will tell on both points. And likely fairly quickly.
Caveat would be that since Kimi K3 is not token efficient and thereby actually more expensive than ChatGPT 5.6, we may need to see a more token efficient open-source model at the frontier or see Grok 5/Composer 4/Muse 2 at multiple points on the Pareto frontier for this potential risk to Anthropic and OpenAI to play out. And I am sure they will both vertically integrate as quickly as possible while continuing the product/harness strength they have shown over the last 8 months.
Open models are only half the shift.
As model margins compress and token demand expands, value moves toward the infrastructure that can execute each token at the lowest reliable cost.
The next cost curve starts with compute that is already there.
The mega bull case for AI infrastructure would be *if* market share shifted away from certain frontier labs with 90%+ inference margins toward cheaper models, whether open-source or closed.
It would increase the ROI on AI spend for end customers by increasing intelligence per dollar, which would drive incremental token demand. Margin dollars would effectively get redistributed from the frontier labs to AI infrastructure providers. The infra winners would be those with the lowest per token cost and the winners at the model layer would be those with the highest token efficiency.
There are many reasons Jensen is so focused on open source, but this is likely the most important one as I think he is probably less worried about a monopsony these days. Lower margin % at the model layer = more margin $ at the infra layer all else equal.
With SpaceX and Meta being vertically integrated and possessing the #3 and #4 models respectively it is more possible than ever. Note that Grok 4.5 is ahead of Fable for some useful tasks at a much lower cost, so ranking them #3 is conservative.
This is not happening yet. Cheap, mostly open source tokens are likely the majority of volume today but the majority of economic value is still accruing to the most intelligent models. Might change though.
We will see.
Karp is right. It’s not only about which model is best. It’s about who controls the compute, who owns the data and who can verify what actually ran.
These cannot be solved by trust alone.
These are the same issues that shaped YAI from inception.
This is why we’ve placed privacy and proof into the execution layer.
Palantir CEO Alex Karp on what customers actually want, the real business of frontier labs, and the importance of open source models:
“What the technical customers want is control over their compute, their models, their data stack, and their alpha. They want to know they own the means of production, and it's not being transferred to someone else.”
"Who owns the data? Are the prompts secure? Is this being transferred to you?"
"If it was so valuable, and I can make you a billion dollars, wouldn't I say I'll make you a billion dollars and I want 30%? Why are they charging for tokens if it's so valuable?"
A lot of AI infra stops at one side of the market.
Either demand aggregation, or compute supply.
YAI is focused on the full loop.
Connecting demand to useful compute, verifying the work, and making sure compute operators get paid.
Inference demand is scaling faster than centralized infrastructure can efficiently absorb.
Underused GPU capacity already exists.
YAI turns that mismatch into verifiable useful work.
Goldman is describing the next AI margin question.
Agentic AI multiplies token volume while unit costs fall.
The question is whether that efficiency becomes hyperscaler margin or lower-cost execution for the market.
YAI is designed to push that efficiency back into the market.
As consumers and enterprises adopt agentic AI technology, token consumption could see a 24-fold increase by 2030, according to Goldman Sachs Research. Read more: https://t.co/Mu0Cfgi9ZF
Brian is pointing at the real shift.
When demand for intelligence keeps expanding, the bottleneck becomes execution.
Where it runs. How it is verified. How privacy is preserved.
YAI chose decentralization because this is where the future of AI is heading.
Good take
My guess is
- demand for intelligence is near infinite
- but 80% of workloads will be running on 99% cheaper models within 12-18 months
- 20% of workloads will still run on latest gen models where IQ maxing is important (scientific breakthroughs, higher level ochestrator agents?)
- rough analogy might be what % of macbooks or gaming PCs sold have the maxed out specs for CPU/GPU, prices are falling much faster than Moore's law here though
- this leads me to think the limiting factor will be energy and compute, not better models
At Coinbase we're working hard on routing prompts to cheaper models where appropriate, and in some cases have been able to keep costs roughly flat, while token usage continues to grow exponentially.
The YAI whitepaper is live.
It lays out our approach to efficiently coordinating inference demand with decentralized AI execution.
Available now
https://t.co/wd4wKvazIS
AI inference is becoming the cost center of the AI economy.
Not just bigger models. Not just more GPUs.
The next layer will coordinate decentralised GPU supply into better economics, private execution paths, and proof of useful work.
We are building YAI to make it real.