We currently offer three routing modes on RunAPI:
Default gives you access to all vendors at around 50% below typical market pricing. It’s built for everyday use and a strong cost to performance balance.
Plus costs about 1.5× Default, and is tuned for consistently high speed and stability. Best for production workloads and anything latency sensitive.
Auto uses smart routing to send each request to the fastest available upstream in real time. It’s available to paid users who want top performance without having to think about vendor selection.
Pick the mode that matches your workload, and let the platform handle the rest.
Opus 4.6 surprised me in a 4 model agent build test.
Compared to Gemini 3 Flash and DeepSeek, it streams fast and keeps going without losing structure. Long outputs stay coherent, actionable, and actually implementable.
I’m running these comparisons on https://t.co/ywdDBbWWAK: OpenAI compatible endpoint, up to 70% off, plus $10 free credits on signup.
Claude Opus 4.6 hasn’t even officially launched yet ,but we’re already preparing the pipeline.
When it drops, three things happen immediately:
• Benchmark it properly. No hype, just numbers.
• Opus 4.5 already runs ~70 tokens/s on our routing , let’s see if 4.6 resets the ceiling.
• And the one that matters most: how much cost can we cut for developers?
Speed is expected.
Cost reduction is intentional.
We’ll handle the optimization.
Claude Opus 4.6 just showed up in Perplexity's API.
Label: "Claude Opus 4.6"
Description: "Anthropic's most advanced model"
And right below it:
Label: "Claude Opus 4.6 Thinking"
Description: "Anthropic's Opus reasoning model with thinking"
Two variants.
A standard model and a thinking model.
Already labeled and described in a live API.
This isn't a Vertex AI log.
This isn't a 403 scan.
This is a product integration that has already defined the model, named it, and categorized it.
Perplexity doesn't add models to their API for fun.
They add them when they're preparing to serve them.
Claude Opus 4.6.
Anthropic's most advanced model. It's not a rumor anymore.
Kimi-K2.5 and Qwen3-Max are now live on RunAPI.
Both models are production-ready with high throughput, strong stability, and seamless integration through our unified API.
To celebrate the launch, we are offering a limited-time price at 50% of the official rate.
No migration needed. Switch the model and start building.
Right now on https://t.co/ywdDBbWoLc
claude 4.5 opus : 42% lower cost
gemini 3 pro preview : 52% lower cost
gpt 5.2 : 55% lower
Same models, same behavior.
Just smarter routing and fewer unnecessary markups.
If you’re shipping with these models every day, the difference shows up on the bill pretty quickly.
A lot of people still don’t understand what we’re building, or how our API prices can stay at roughly 50% of the “market” rate.
We’re not hosting the models ourselves and reselling them.
For every model on RunAPI, we do the unsexy work: we hunt down multiple strong providers, vet them hard, and only list the ones that can consistently deliver real quality. The goal is simple. Same model behavior you’d expect, but without paying the premium just because one provider is the default.
Then we add the part that actually makes it feel good to use: routing. Every request gets sent to the best provider at that moment, based on real time speed and reliability signals, not just a static rule. If one provider slows down or gets flaky, traffic shifts automatically. You keep the same model, but you’re always hitting the fastest lane.
So the “why” is this: competition plus selection plus routing. We keep multiple providers in the pool, we don’t compromise on performance, and we let the system continuously choose the best option.
You pay half, but you still get the best available provider experience. That’s the whole point.
Kimi K2.5 dropped. When a new model ships, we usually find multiple providers at ~60% of the market price within 3 days.
Then we run the boring stuff:
1. TPS + time to first token
2. all day stability
3. agent tool calling sanity checks
If it passes, it goes live on RunAPI.
Believe it or not, most models on our platform end up as fast as OpenRouter (sometimes faster) but with a real cost cut.
@adishjain333 Appreciate it — we’ve actually got it running already. If you’re shipping agents or heavy workflows, feel free to try the API and shave down cost + latency at https://t.co/PeGtTTrs9L
Seeing 50% cost cuts while still pushing 80+ tokens/s through OpenCode was honestly surreal. Matching OpenRouter-grade speeds inside a coding workflow is not something I expected this soon. Excited for more vibe-coding experiments.
Just did a speedrun in Claude Code using our API. The skill came together in seconds. It feels genuinely snappy, and the best part is the cost. For the same workflow, the bill ends up about half of what I'm used to.
Behind the scenes, nothing fancy. Just pointed Claude Code at our API, ran it once, and it landed. If you do this a lot, the savings add up fast. Also, new users get $10 to try it out. So click https://t.co/PeGtTTrs9L
Totally fair question I’m not speeding up the model or trading quality for TPS RunAPI is basically smart routing We pick the best provider in real time based on price latency and stability The upstreams are official partner cloud providers running the same models so correctness and edge case behavior should stay consistent
@Angaisb_@OpenAI Actually on RunAPI you can pick the reasoning effort: none, minimal, low, medium, or high. It’s a simple toggle, so you’re not stuck with always-on thinking.
If Snow Bunny is really “very very good”, we’re ready.
Adding it to RunAPI day-one, aiming ~50% cheaper, and we’ll post benchmarks vs Opus 4.5 + GPT-5.2.