@Abdella6if look into our Inference Router to optimize qaf’s token cost - ive been using qaf myself (number 4790!!) and i think you could benefit from our large selection of OSS models and passthrough, and get the best bang for ur buck: https://t.co/HR1g9O2Hip
@simondelorean@digitalocean@Kimi_Moonshot@vllm_project hey! kimi k3 should be available on the model selection screen with “/models”. if you are unable to see kimi k3, make sure your OpenCode is updated to the latest version! i was able select via the connector, lmk if you get any issues
@botirkhaltaevv at digitalocean we ship a preset software-engineering router based on evals with specific tasks, might interest u? https://t.co/YAA1gAmllu
@markobilal@finbarr at digitalocean we made a product called inference router; we’ve made our own 4B/30B MoE model dedicated for just routing called Plano-Orchestrator. https://t.co/0ZTYjIVCDh
DigitalOcean Inference Router is now available through OpenCode & setup is 3 steps. ☁️🆕🤝
Run/connect & it automatically routes to the right model for the job. No more paying frontier prices for a commit message.
@simondelorean@defileo there isn’t direct support via local models unfortunately - but if you are looking to use your OpenAI subscription with OpenCode, our open source proxy Plano (https://t.co/bSreKbNEpz) can help you with this! we’re adding Anthropic CLI support soon, and Codex CLI is already in.
@TFisPython@defileo plano is an open source proxy that has a dedicated model just to route: https://t.co/bSreKbNEpz
its free and really easy to setup into Claude Code, Codex or any coding agent - just point it!
@sama i’m the youngest engineer (19) at digitalocean, working on inference router (helping cost effective measures with routing between models for the right tasks). i’d love to come to the party! been using codex with model routings :)
Cut your LLM costs by 50%.
Plano is an open-source AI proxy, powered by Arch-Router-1.5B (deployed at scale at HF 🤗), that auto-routes each prompt to the right model based on complexity.
Also handles orchestration, guardrails & observability.
GitHub: https://t.co/rLFA2Etojj
Just released support for preference-based LLM routing for OpenClaw in Plano 🚀
Those who use @openclaw know that it can churn through tons of tokens. So you have two options pay for those token or plugin in a cheaper alternative and sacrifice perf. What if you don’t have to make this trade off? What if you could route traffic for certain tasks to @claudeai and others to!@Kimi_Moonshot ?
With Plano you can: https://t.co/BioebnQLJx. Check out our demos folder under LLM routing for more details
Plano is finally out - and who better to introduce it to the world than our ex high-school intern, now turned dx aficionado @spherrrrical (see video below).
Plano is delivery infrastructure for agentic apps: An AI-native proxy and dataplane for AI workloads.
On-the-ground AI practitioners will tell you that calling an LLM is now the easy part. The hard part is delivering agentic applications to production quickly and reliably, then iterating without rewriting systems every time.
In practice, teams keep rebuilding the same concerns that sit outside any single agent’s core logic. They need model agility — the ability to pull from a large set of LLMs and swap providers without refactoring prompts or streaming handlers. They need to learn from production by collecting signals and sample traces that tell them what to fix.
They need consistent policy enforcement for moderation and safety, rather than sprinkling hooks across codebases. And they need multi-agent patterns to improve perf and latency without turning their agentic app into orchestration glue.
These concerns get rebuilt inside frameworks and application code, coupling product logic to infrastructure decisions. It’s brittle, fast-changing and it pulls teams away from core product work into plumbing they shouldn’t have to own.