**TrajectoryRL Update**
A quick recap of what shipped on SN11 recently.
—
**What TrajectoryRL is**
TrajectoryRL ships out-of-the-box SOTA agents on small open-source LLMs. SN11 on Bittensor is the open market that pays for the agent scaffolds that move the quality / cost frontier. First vertical: autonomous coding on `qwen/qwen3.5-35b-a3b`. Miners ship prompts with harness; the network pays for the ones that move the frontier.
**Introducing Terminal-Bench**
In a recent update, we introduced **Terminal-Bench** — `trajrl-bench` leverages part of its scenarios for our eval harness. That means miners optimize against real, public agent tasks rather than a benchmark we invented in-house. Every scenario has public provenance, and SN11's SOTA claims sit on top of established agent-evaluation work. https://t.co/2G4Kb5mVn0
—
**Challenger / Winner mode — new incentive mechanism**
The previous mechanism ran 24-hour epochs that re-evaluated every miner from scratch. We've moved to **Challenger / Winner mode**: one challenger per epoch, evaluated head-to-head against the seated winner. The seat only changes when a challenger qualifies and beats the seated score by ≥ δ. cleaner signal, Faster epochs and faster finalized emissions.
—
Learned a lot from Distill along the way — shout out to @const_reborn
Live at https://t.co/RoH66VNwR2.
#Bittensor #SN11 #TrajectoryRL
SN11 update: free inference for validators, on our own stack.
Every active SN11 validator now gets a free key to run the eval model (Qwen3.6-35B-A3B) on our own GPU fleet instead of paying OpenRouter.
- Faster and more stable, served and tuned by us
- About 1/10 the cost of OpenRouter, free for validators
- Zero data retention
- Same testee across the fleet, cleaner consensus
Next: we will bring finetuned models to this same stack, ours and miner submitted ones. The models the subnet produces will run on the infrastructure the subnet runs on.
We're the first to make the full GLM-5.2 (FP8) run on RTX 4090s.
GLM-5.2 is the new 753B SOTA open-weights model, and it officially ships for datacenter GPUs only: H100, H200, B200. We ported its sparse-attention kernel stack to consumer hardware.
A frontier open model, off the scarce GPUs and onto the abundant kind.
https://t.co/08rCk3b3h5
@Zai_org
Why it matters: the next era of SN11 is SOTA agentic models built on small open LLMs, free to use and fine-tunable for any domain. We serve them at full context, on hardware anyone can run. Up next: Kimi K2.7 Code and GLM 5.2 results.
You don't need a rack of H100s to serve frontier models at scale.
We're running Qwen3.6-35B-A3B at production scale on consumer GPUs: full 256K context, ~170 tok/s flat at any length, ~1,600 tok/s aggregate.
@Alibaba_Qwen@lmsysorg@opentensor
And it's cheap. At scale we serve it for roughly 5x less than OpenRouter charges for the same model. The reason is the hardware: consumer cards rent for a fraction of datacenter GPUs per hour.
Want to finetune but lack hardware? We operate our own GPU clusters and will provide servers to
committed finetuners. Reach out in Discord.
The goal: the SOTA agentic model for a domain, one domain at a time. Next up: DevOps, the agent
that operates machines.
The next era of SN11: bring your own model.
Season 1 proved the skill competition. SKILL.md packs took a stock 35B open model from "reads
the task and gives up" to near-ceiling on real terminal tasks. Now we're opening the next lever:
miners will submit finetuned models, not just skill packs.
A submission becomes a (SKILL.md, model) pair. Models are SFT finetunes of Qwen3.6-35B-A3B,
registered on our inference subnet (separate announcement coming) and served with verified
inference: the server cryptographically proves it ran the claimed weights.
Any miner can build packs on any registered model. When a submission wins, the model author
earns additional SN11 emissions on top of what the pack author earns. Pack rewards stay whole.
Build the best model, and every pack that wins on it pays you.
Today, @MichaelElabd, @QuantumArjun, and I are excited to announce Trajectory.
We are a research lab and product company building the platform for Continual Learning.
Our platform unlocks the signal already sitting in product usage, so companies can continuously post-train large-scale agentic models that outperform the frontier. @trajectorylabs
We’ve raised $15M from @Conviction, @BessemerVP, @radicalvcfund, @jeffdean, @drfeifei and more.
We’re partnering with some of the best AI-native companies: @ClayRunHQ@Harvey, @DecagonAI, @mercor_ai, @RogoAI to power their agentic systems, some of which we are already in production with.
We’ve brought together a world class research team from DeepMind, OpenAI, Apple, Meta Superintelligence, Amazon AGI, Scale AI, and an elite product team from Stripe and Figma.
AI will never again start on day one. Every correction, every retry, every edit will make products smarter. This is Continual Learning.
@addyosmani It's also the bet behind TrajectoryRL — a Bittensor subnet that benchmarks harness quality(prompts, tools, sandbox, loop).
Once you measure agents this way, the "which model is smartest" debate gets a lot less interesting. Harness gap is the real capability gap.
"Agent = Model + Harness." Spot on.
It's also the bet behind TrajectoryRL — a Bittensor subnet that benchmarks harness quality(prompts, tools, sandbox, loop).
Once you measure agents this way, the "which model is smartest" debate gets a lot less interesting. Harness gap is the real capability gap.
it's crazy to think open-source small LLMs are only 1-1.5 years behind frontier. imagine running GPT-5.5 or Opus 4.7 on your gaming PC a year from now. Purely local inference.