the stack should be optimized
everyone's screaming "more compute" "better chips" while we're squeezing everything out of existing hardware
infrastructure upgrades won't magically fix your inference costs
@SiliconFlowAI Hey @SiliconFlowAI Our product optimizes your entire serving stack around real agentic workloads. Way more margin on the same hardware, costs are predictable. Think you'd find it interesting.
@nikolaborisof Hey @nikolaborisof we built a simulator that optimizes your entire serving stack around real workloads. Way more margin on the same hardware, costs are predictable. Think you'd find it interesting.
We ran GLM 5.1 on MI350X -ย much faster and cheaper than MI355X and B200.
SwarmOne optimization delivered off-the-chart agentic performance on 150-200K context windows.
Single node. 8รMI350X. 140 tok/s/user. $1.44/Mtok.
34โ55% faster. 54โ64% cheaper. No synthetic benchmarks.
Most configs on the chart sit on frustrating or constrained territory for agentic inference. Below ~50 tok/s/user with variable ISL/OSL and tool calls, your agents are bottlenecked - no matter how low the cost looks on paper.
Respect to @AnushElangovan, @roaner and HaiShaw for pushing MI355X forward. We're building on that momentum and showing what full-stack agentic optimization unlocks on ANY silicon.
If you're scaling agentic workloads and cost-per-token matters, let's talk.
@nvidia@AMD@SemiAnalysis_
Every LLM benchmark tests chatbot performance at 2K context.
Agentic workloads runs at 40-400K context with tool schemas, file contents, and error traces growing every turn.
No benchmark tested this. So we built one.
Introducing AgenticSwarmBench ๐งต๐
Every LLM benchmark tests at ~2K tokens of context.
Claude Code and Cursor send 40โ400K token requests.
The real-world performance gap? 3โ5x.
Today we're launching AgenticSwarmBench -
the first open-source benchmark for real agentic workloads.
- 110 tasks, 5 languages
- Context up to 400K tokens
- Prefix cache defeat (true cold-start)
- Record & replay real Cursor/Claude Code sessions
Tested on @nvidia@AMD@intel@tenstorrent
https://t.co/N5tnFKpV0O