Today we present Orion-16B, our live training run.
The purpose of this model is to show that @IOTA_SN9 can train using hundreds of globally distributed, heterogeneous commodity GPUs.
- 16 billion parameter language model
- Training across 3 continents
- Split into 10 pipeline stages using ResBM
- Up to 256 GPUs
- Heterogeneous mix of 4090s and 5090s
- Bittensor subnets provide permissionless compute
https://t.co/5JyZOHfsHv
A 16-billion parameter model is now training live on @IOTA_SN9; globally distributed over three continents, heterogeneous, permissionless, owned by no single entity.
Today, we present the next stage of Project Orion - Orion-16B.
Over the coming weeks, we'll show what it means to train across distributed, heterogeneous and "unpredictable" compute: the engineering, research and systems beneath the model, the providers powering it, and the moments that prove the infrastructure holds.
More details below.
Our first commercial @Apex_SN1 competition is a wrap!
Very happy with the results and great working with the @AureliusAligned team on one the most important problems in AI.
Great write up on what we built together, and what’s next.
I would have assumed it was fairly obvious, but in case it's not: a million-line codebase (also known as a "harness"), running at inference time, orchestrating thousands of calls to a neural network for any given task, is the exact definition of a "neurosymbolic architecture"
The economics of distributed training are very compelling. Once the tech reaches a critical point, I believe it will trigger a chain reaction: reducing training costs severalfold will enable many new teams to build efficient, intelligent models.
I think we're very close.
Distributed training is possible, that question is settled.
The harder question is can a system take compute that constantly joins, leaves and restarts, over ordinary internet, without degrading the model or eroding the cost advantage that made it worth using?
The next part of our thesis works through what has to be true for that to hold.
Read it here: https://t.co/3RLJmC7GeT
Our distributed training runs use compute that iota doesn't own and can't fully control. Machines join, leave, stall, and occasionally fail – these are normal operating conditions for @IOTA_SN9.
So, how does iota ensure fault tolerance when the GPUs that are used for training are unreliable?
Over July 23–24, Orion's 16B run gave us a live example of exactly that.
We're already working on support for multi-stage nodes to get value out of diverse heterogeneous compute.
Much bigger machines are assigned multiple pipeline stages which act like activation superhighways. This maxes out their VRAM, increases net training throughput and reduces data loss.
Project Orion's 16B run isn't training on one type of machine, in one place, from one supplier.
The run is built to handle 256 machines. Right now, around 188 are online; drawn from multiple independently operated providers, running different hardware, in different locations, under different conditions.
Here's what's powering the run and why the mix matters.
You are still early. Extremely early!
Updated July 2026:
Each dot = ~3.3 million people
2,500 dots = 8.3 billion humans
Gray: Never used generative AI → 5.9 billion (71%)
Green: Free chatbot users → 2.3 billion (28%)
Yellow: Pays for AI → ~80 million (1%)
Red: Uses coding agents → ~12 million (0.14%)
Most of humanity still has no idea what’s coming.
~2% of U.S. households pay for AI subscriptions.
We’re not even close to mainstream yet.
Source: https://t.co/3Jmpfqh11e (July 2026)
To the skeptics who think that distributed training is horse shit, you’re not far off. About 75% of our GPUs in Orion-16B are powered by biofuel, courtesy of UK livestock. Thanks to @green_compute_ for making our 16B pretraining run almost entirely carbon neutral.
Talk about the long tail of compute..
iota is built to bring distributed capacity online: hardware that would otherwise sit unused, coordinated into one functioning training run.
Project Orion's 16B run trains across 256 globally distributed, heterogeneous GPUs, supplied by an open network of independently operated suppliers.
Today, we wanted to spotlight one of our partnership and introduce one of the suppliers behind Orion’s 16B run: @green_compute_
iota is built to bring distributed capacity online: hardware that would otherwise sit unused, coordinated into one functioning training run.
Project Orion's 16B run trains across 256 globally distributed, heterogeneous GPUs, supplied by an open network of independently operated suppliers.
Today, we wanted to spotlight one of our partnership and introduce one of the suppliers behind Orion’s 16B run: @green_compute_
Orion-16B is training at an average MFU of around 23% making it the most efficient training run of its kind (and cheapest).
We actually had a bunch of 30%+ MFU epochs at the start, and several shorter 16B warmup runs which averaged over 30%. Was really tempting to keep tuning but we wanted to just get it out. Next run I’m confident we can sustain well above 30% for 100B+ tokens.
Probably the coolest thing about Orion-16B is how resilient our system is. Check out the screenshots -- loss continues decreasing despite over 50% of nodes dropping during partition upload.
Project Orion shows that @IOTA_SN9 can train models efficiently using highly unreliable compute from around the world. We spent over a year of R&D on this because making training an interruptible, liquid workload makes it >5x cheaper, meaning more people can own their own model.
This is our answer to the question of open models.
this system is very fault-tolerant, but we had a LOT of unexpected faults in launching it (some planned, some provider based, some unexpected):
- we had problems with one provider's network being bottlenecked during weight merge, so needed to build automated monitoring to account for improve this (and also added roadmap to look at dynamic / more effective ways of using relative local connections)
- we managed to completely overload our caching service because we were running something like 18 runs at once at scales we hadn't done before - our fallback covered this, but we improved the autoscaling logic on a lot of the infra
- we did a series of planned shutdowns, including a 48 hour one where we did full checkpoint restarts
all this and more, but they are important steps for us making this thing bombproof for production. we have some tweaks to improve on this, and can occasionally be perfectionists, but thought it was best to get this into the wild!
this run is going to be something of a gym for us - we're gonna do different exercises on it (trying to break it like the above), and test different parts of our system on it throughout august.
we will also be onboarding more and more bittensor compute! first tests went with this, so we are paying out based on contribution now - exact details on our first pass approach for this later this week, but we do want lots of community feedback!
I cannot overstate how excited I am to share this with the world. 16B trained over the internet by 4090/5090s with help from @green_compute_ and @TargonCompute. MFUs upwards of 40%, with compute all the way from the USA to Hong Kong. @IOTA_SN9
https://t.co/rXA9XMVxvR
Today we are reducing miner burn to 0%, effective immediately.
Our prize pools for all active competitions have been increased by >8x 🚀🚀
With more funding available to our agents, we can tackle bigger problems, onboard more agents and support a larger ecosystem of competitions than ever before.
There has never been a better time to build with us.
https://t.co/JezlIffppS