Boundless set out to build a zero-knowledge proving network. We ended up building something bigger: a global network that coordinates GPU capacity at scale.
Today, we’re expanding into AI.
@sshankar on what’s next:
Most public inference benchmarks are wrong for production. I analyzed more than 55 million public request records across chat, code, RAG, reasoning, and agents, plus modeled batch-job examples.
All skewed heavily toward prefill before cache reuse: long inputs followed by short outputs. That changes the bottleneck. Decode is often memory-bandwidth bound, while prefill can process many tokens in parallel and become compute-bound.
Agent harnesses are pushing this pattern further. They send large tool schemas and histories while managing state client-side, repeatedly passing context instead of relying on a persistent model-side session. This makes overlooked, compute-skewed GPUs, I.E. 5090s, more useful than most people think.
At @boundlesshq, we’re improving inference across the stack, from routing and serving engines to matching each production workload with the right GPU.
Boundless CEO @sshankar spent his early career at Lyft and Grab, where the problem was idle cars and drivers with no way to convert time into income.
GPUs have the same problem: the workloads come in spikes, and even well-run infrastructure teams top out at 70–80% utilization.
Idle compute is sitting in homes, datacenters, clouds, and mining operations around the world. @sshankar on how we track it down and put it to work, and what that does to the cost of intelligence ↓
Going live in an hour with @rkbaggs on @Cointelegraph to talk about compute becoming a tradeable commodity and the demand for AI inference rising.
Join us live at 3pm BST