The headline is a 16B-parameter model, but the model size isn't the hard part.
The hard part is everything beneath it: coordinating hundreds of independent GPUs, in different datacenters, over the public internet, while machines join and leave whenever they want.
That's the iota engineering that makes Orion's 16B run work.
Here's how it holds together.
"The world is not short of installed GPUs. What it is short of is a way to turn imperfect, scattered, intermittent supply into useful training, at a cost and reliability a buyer can plan a roadmap around. That is the system iota is being built to be, and Orion is the first real evidence that it works."
"That gives us two results that mark the ground already covered:
Orion-100B — efficient hundred-billion-parameter training on globally distributed commodity hardware.
Orion-16B — efficient heterogeneous decentralised training across hundreds of GPUs.
The first proved the approach holds at scale; the second, live now, proves it holds across mixed hardware. Between them they answer the two questions that matter most: can a model this large be trained this way at all, and can it be done on the imperfect, varied fleet the real world actually offers.
What remains is the widest version of the problem: less controlled, more open participation, where the interruptibility and heterogeneity arguments still lean more on mechanism and testing than on a finished, demonstrated run, that is the next thing to prove."
"TAO is the same monetary skeleton pointed at a more valuable problem. The world is full of stranded compute the way it was full of stranded energy, and TAO gives every idle chip value. The difference is the output. It is not a hash whose only virtue is being expensive. It is intelligence. The scarcity model with a 21 million hard cap is deliberate. No board can dilute it and each halving is coded into the blockchain. The monetary policy, like the neutrality, is entirely structural.
So how does the token capture value? Not as equity, because it is not equity in anything. TAO is the unit of account, the incentive reward, and the settlement asset of the whole exchange. Miners are paid in it, validators must stake it, subnet entry is priced in it, and every subnet's token trades against it. Every function creates a reason to hold it and the supply that can be held shrinks on a set schedule. You value it the way the market eventually learned to value Bitcoin: by the size of the market it settles, discounted by the probability that the market matters. It’s important to note one of the most important lessons in commodity markets. When products commoditize, the producers compress toward utility returns and the durable asset is the layer that prices, settles, and collateralizes production. Wheat farmers earned commodity margins while the Chicago Board of Trade became one of the great franchises in American finance. If intelligence is becoming a commodity (and I argue it is), then the scarce asset is not any individual model. It is the exchange."
"The run's real significance is architectural. Every previous decentralized effort required each participant to host the entire model, which means the biggest model you can train is capped by the smallest machine you allow in. IOTA shards the model itself across the network, one slice per peer, so capacity is bounded by the sum of the network rather than the minimum of it. Macrocosmos is not training one model. They are building the architecture on which anyone can run frontier-scale training over the open internet. When the rails exist, you do not get one open model. You get a proliferation of them."
"Then, in June, the ceiling itself moved. Macrocosmos, one of the network's largest research groups, published Orion-100B: the largest distributed pretraining run ever conducted over the open internet, a hundred-billion-parameter model split across sixteen stages, each hosted by a single commodity GPU, achieving roughly 65% of datacenter training speed at a fraction of the cost. When their IOTA architecture launched in June 2025, the prevailing view among researchers was that this was, as one leading skeptic put it, "like fighting gravity." Eleven months and seven hundred experiments later, they scaled model size 67x in a single month, and the known optimizations they have not yet applied would push compute utilization toward 98%."
@Old_Samster Why has @IOTA_SN9 not thought of this or, better yet, what prevents iota from building models using this method? Seems to me like something they can do on top of the work they have already done. thoughts @WSquires@macrocrux@MacrocosmosAI ?
The second question was harder: could the system get better on its own?
IOTA is highly sensitive to its launch configuration — cache sizes, queue sizes, batch sizes, accumulation counts and nodes-per-pipeline-stage are all tightly coupled, and each one moves throughput, efficiency and convergence together.
Our ablation runs proved we could build a system that systematically searches this space and self-improves, rather than relying on manual tuning run after run.
That was proof #2: Orion doesn't just work — it learns how to work better.
https://t.co/2RJ9sNQ3mY
First, we needed to prove the core idea worked: that training split across a distributed network could match a model trained in a single data centre.
Our 2B-parameter run, trained to 16B tokens, did exactly that, converging to a quality comparable to a centralised baseline.
That gave us proof #1: distributed training isn't a compromise, it works.
Our co-founder @macrocrux post provides more insight into this first proof.
https://t.co/SHfOWMObCm
this system is very fault-tolerant, but we had a LOT of unexpected faults in launching it (some planned, some provider based, some unexpected):
- we had problems with one provider's network being bottlenecked during weight merge, so needed to build automated monitoring to account for improve this (and also added roadmap to look at dynamic / more effective ways of using relative local connections)
- we managed to completely overload our caching service because we were running something like 18 runs at once at scales we hadn't done before - our fallback covered this, but we improved the autoscaling logic on a lot of the infra
- we did a series of planned shutdowns, including a 48 hour one where we did full checkpoint restarts
all this and more, but they are important steps for us making this thing bombproof for production. we have some tweaks to improve on this, and can occasionally be perfectionists, but thought it was best to get this into the wild!
this run is going to be something of a gym for us - we're gonna do different exercises on it (trying to break it like the above), and test different parts of our system on it throughout august.
we will also be onboarding more and more bittensor compute! first tests went with this, so we are paying out based on contribution now - exact details on our first pass approach for this later this week, but we do want lots of community feedback!
The activation routing algorithm used in Orion-16B was developed by our subnet 1 miners.
When nodes drop out during training the remaining nodes quickly become congested and create bottlenecks, which brings training to a halt. These kinds of failures are common and unpredictable, but you can statistically beat them by using the route planning algorithm that we saw our miners create. Even better, our approach improves fault tolerance and provides optimal load balancing across heterogeneous devices.
Sustained training speed of 50-60% of a datacenter, despite the model being split into 10 slices, distributed across 3 continents, and run on cheap commodity hardware is a truly impressive feat 👏 great to see our competition creating real value for @IOTA_SN9!
A 16-billion parameter model is now training live on @IOTA_SN9; globally distributed over three continents, heterogeneous, permissionless, owned by no single entity.
Today, we present the next stage of Project Orion - Orion-16B.
Over the coming weeks, we'll show what it means to train across distributed, heterogeneous and "unpredictable" compute: the engineering, research and systems beneath the model, the providers powering it, and the moments that prove the infrastructure holds.
More details below.
literally you're not on the same type shit we on fr we visualizing real pipeline parallel distributed activation transfer over p2p internet connections you actually not built like @IOTA_SN9
Recent changes to Bittensor emissions incentivizes subnets to reduce their miner burn to zero.
In light of this, we are rolling out zero-burn initiatives across our subnets which balance productivity gains with price stability.
We are commencing today by reducing miner burn to 0% on @Apex_SN1.
These changes support a positive flywheel for our subnets:
1. Miners can earn more alpha for their work, enabling higher quality and more ambitious solutions. Overall competitiveness will increase, selecting for top performers.
2. Subnet outputs will improve due to increased competitiveness, which supports our products and services
3. Zero burn increases chain buys (deepening pool tao reserves), which supports prices for holders
We will continue to adopt to changes in Bittensor protocol in a measured, auditable way that drives value for holders, and fair economic returns to participants.
Every frontier lab ships some version of the same structure: a cost effective tier for routine work, a mid tier for harder reasoning, and, in some cases, a further premium tier.
Anthropic's version: Sonnet handles routine work, Opus takes the harder reasoning, Fable provides the premium finish, but the logic only holds if each step-up buys real capability.
@Kimi_Moonshot K3 doesn't threaten that by beating the top tier. It threatens it by handling everything below it well enough that the top tier only gets called in for the hardest edge cases and final sign-off.
K3 bakes most of the cake, the premium model now applies the icing.
It can still be the best model in the world and still become a much smaller share of total consumption.
This is often framed as a sovereign AI problem, nations and companies pushing for more control over the models and infrastructure they depend on.
That's part of the answer, but not the deepest layer of it.
For AI to be genuinely resilient, openness has to extend to the training layer itself: models trained across distributed infrastructure, supplied by multiple operators in multiple locations, with no single hyperscaler or national compute base as a chokepoint.
That's the problem @IOTA_SN9 is built to solve.