@MaratDukhan@cHHillee@soumithchintala tok/s is using end to end duration which include TTFT. But the portion of TTFT over end to end duration will affected by number of output tokens
Two months ago, I left Annapurna Labs, ending a five-year chapter at Amazon. Amazon treated me exceptionally well and gave me more opportunities to grow than I can count. I’m grateful for the people I worked with and everything I learned there.
I followed my heart to Prime Intellect. Its mission to build an open superintelligence stack deeply resonated with me. I believe every person and every business should be empowered to build their own intelligence. Your AI agent should be yours, not a third-party stranger.
Walking into Prime’s SF office, I felt ten years younger. These past two months have been the most exciting of my career.
Three things stand out:
1. The people.
Prime’s team doesn’t all come from famous AI labs, but I’ve rarely seen this level of curiosity and determination.
When we made our high interactivity GLM-5.3 serving endpoint available internally, everyone switched their agents to it. They never went back.
The speed improved productivity. Just as importantly, the team kept sharing feedback, helping us refine performance and tool-calling reliability. Soon, RL rollouts, synthetic data generation, and long-horizon agent research were running on the same infrastructure.
As of this week, our internal inference usage is 600–700 billion tokens per day. That experience gave us the confidence to open the service to the public.
2. The full stack.
Prime brings together hosted training with prime-rl, verifiers, sandboxes, an RL environment hub, the Prime Agent harness, and now Prime Inference.
You can train, deploy, and run agents in one place, without stitching together infrastructure through a massive engineering effort. And most of these components are fully open source.
Our inference stack builds on that same foundation, with deep collaboration across NVIDIA Dynamo, Inferact, and the broader vLLM community.
3. The compute.
For anyone working on LLM infrastructure, access to compute at this scale is a privilege. Building and operating inference across large scale of GB200 racks and B300s is not something I take for granted. Vera is already in our hands, and Rubin is on the horizon. NVIDIA’s latest hardware gives us the foundation to push inference performance further.
Alongside the public launch, we published a technical post on how we built Prime Inference: https://t.co/zsEVO51jsL
Inference has moved incredibly fast over the past year, and it can be hard to separate useful advances from the noise. We use GLM-5.3 optimization to show the work end to end: how we serve agent workloads with low latency while also supporting high concurrency, dependable availability, and reliable tool calls. Recommend for a read if you are interested in learning about P/D, topology choice, NVFP4 KV compression kernel fusion, tool call reliability.
Fast inference matters. Serving real agents reliably at scale is what counts.
we’ve known early on that inference is a crucial component of the open superintelligence stack
Prime Inference began as the serving platform we needed ourselves. excited to release it today
Fast inference doesn't necessarily convert into value. At Prime, we ran rigorous Eval to ensure every quantization, inference optimization technique preserve the original model authenticity
In the post, we provided a deep insight on how we serve GLM-5.3 in production with vLLM and Dynamo, From topology design decision to NVFP4 KV compression to low latency KV transfer.
Fast single-user responses matter, but the real test is keeping agents responsive and meeting our SLA when many are running at once.
Prime Inference processes close to a trillion tokens a day across RL rollouts, synthetic data, long-horizon agents, and coding tasks. Now we’re bringing our open-source inference stack to the public—a key piece of Prime’s mission to build open intelligence for continual learning agent loop.
Introducing Prime Inference:
We've served trillions of tokens for RL and dedicated customer deployments
To own your intelligence, you need to own your inference
Unpacking our inference stack
I joined AMD right after graduating from MIT. Two stints there and a decade building @cerebras later, I still lose sleep over how much memory and compute to put on a chip.
I’ve spent my entire career in computer architecture, and I’ve never seen the problems change this quickly. Models keep getting larger. The demands on the hardware keep growing. Decisions about where to put memory and how to connect it have become central to what people can build with AI.
We chose SRAM for its bandwidth. Now we’re working on wafer-scale stacked DRAM to give larger models more capacity while preserving fast inference. That brings us back to some very physical problems: how to power it, cool it, and manufacture it reliably.
This is the work I love. In part two of our AMA, I answered your questions about those decisions and where we’re going next.
0:30 What changes between CS4, CS5, and CS6?
2:19 HBM already stacks memory. What’s different here?
3:34 We chose SRAM for speed. Why add DRAM?
4:52 Drawing the stack is easy. Making it work isn’t.
7:23 The memory-versus-compute decision that keeps me up at night
lots of categories emerging around the AGI stack:
- open model RLaaS
- data & evals
- inference
- coding harnesses
- GPUs
- agent tracing
- sandboxes
at @primeintellect, we agree. we do all of these things. people used to ask why we do so many things. they ask that less now.
Introducing Prime Sandboxes:
MicroVM sandboxes purpose-built for RL training.
Model training requires running tens of thousands of concurrent sandboxes, leading to complex and costly configuration. We built Prime Sandboxes for our own team. Today we're releasing them publicly.