This explains how the system works really well.
Solving Erdős problems through @conjectures_io shows why Bittensor is worth building on.
We’re proud to be part of this project. Congrats to @Wejh99 on bringing Conjectures to life.
Build on Bittensor.
Four years ago I had to explain to every engineer we interviewed why building on a decentralized AI network isn't career s*icide.
The objections never change. It's too early and risky.
Open source can't win anyway.
And the whole thing runs 24/7, so the responsibility sounds scary.
All three deserve serious answers and I'm going to write them up properly.
But first I'm curious about the other side: if you've looked at Bittensor and decided against building or working on it, what stopped you?
Honest answers only. I'll respond to the interesting ones.
Fourth, and this one is uncomfortable: exploit your own subnet before anyone else does. We mine our own, publicly, and publish what we find.
You can't change miners. They're rational profit maximisers, exactly like the cobra farmers.
You can only change what you pay for.
Second: open source everything, including validation.
SN17 does this. Every miner can audit every scoring decision, which means exploits get found by the crowd instead of quietly farmed.
Third: king of the hill on your own benchmarks.
Beat the incumbent or earn nothing. SN3 and SN97.
So what separates a real mechanism from a cobra bounty?
Four things we've seen work.
First: make the ground truth something nobody can control.
SN18 scores against tomorrow's weather. You can't front-run a thing that hasn't happened yet.
Every subnet has this problem. You want good models, so you pay for a score.
But a score is a proxy, and proxies can be satisfied without producing what you actually wanted.
We once got full emissions on a subnet by echoing the validator's own questions back at it. Scored 110%. Zero compute.
India had too many cobras. The government paid a bounty per dead snake. Reasonable.
Within months people were breeding cobras at home.
Safer than hunting them, same payout.
When the scheme was cancelled, breeders released their stock.
More cobras than before it started.
The failure isn't that people cheated. Nobody broke a rule.
The failure is that the measure got decoupled from the goal. They wanted fewer live cobras. They paid for dead ones.
Those aren't the same thing, and someone will always find the gap.
Today, we’re announcing a solution found by our miners to Erdős Problem 96, open for over 66 years.
The result disproves the conjectured linear bound, constructing strictly convex polygons with superlinearly many unit-distance pairs.
Verified in Lean through Conjectures. Full proof below.
Sentiment right now feels exactly like it did four years ago when we started.
That turned out to be the best possible time to begin.
Not saying history repeats. Just that it rhymes, and I know this rhyme.
What it actually costs to compete in the 110B run:
8x B300 minimum. ~$8/hr per card, so $64/hr, and that's the cheap end.
A training run takes 3-4 days.
About $5k per attempt. Which might not beat the king.
You can go cheaper with more, weaker cards, but then you need a genuinely optimised training setup, which is its own skill.
But the reward you can get - even $20k per day
1 WEEK IN. 186 MODELS EVALUATED
The benchmarks are climbing, new models keep entering, and we’re just getting started.
In the last week, we added new datasets to the evaluation suite. Now we’re evaluating on more than 7T tokens. And this is only the beginning.
No invite needed. Just bring your model and see how far you can push it.
More info about competition you can find: https://t.co/AhLCS6n6dD
Our defense is layered.
We red-team our own judge, run static pre-checks, and keep a separate model whose only job is spotting injections. It's held so far.
The open question we're working on now: which judging methods survive this pressure long term.
If you've run adversarial evals yourself, I'd genuinely like to compare notes.
We run an AI training competition on Bittensor (subnet 97): crowdsourcing the training of open-source coding models from independent teams worldwide.
Training turned out to be the easy part.
Every single participant also tries to cheat, and handling that took more engineering than the training itself.
A few favorites from the log:
Also in the log: resubmitting the champion's weights with cosmetic changes to slip past similarity checks, and custom chat templates emitting fake meta tokens to bend the judge's scoring.
Built a compression layer, shipped it last week, and the algorithm inside it isn't ours.
It comes out of the competition on SN114 and gets replaced every two weeks by whoever wins.