A GPU benchmark for coding agents needs more than tokens/s.
Try real tasks at your target concurrency. Log first-token wait, completion time, failed tests and human fixes. Repeat with cold and warm cache.
Our 5090 test notes, including failures:
https://t.co/5pfcQravoi
@0xSero before the $1k plan, turn effort down. we ran Opus 5.5 on 3 coding tasks today: medium scored as well as max on every one, at roughly a tenth of the cost. a lot of that spend is just max effort doing laps
same pattern on coding for us. Opus 5.5 in Claude Code, 3 tasks × medium / xhigh / max × 3 runs each, graded by hidden tests.
medium passed everything in 1–2 min. xhigh and max never scored higher, just took ~3–20× longer and cost ~3–19× more. max even hit our 25-min cap twice, once on the easy task.
did max ever run into time limits on the proofs?
@TheAhmadOsman the keep-alive number is the tell though. in a 2-box pipeline each GPU sits idle for the other half of every token, long enough to clock down, so you pay a wake-up on every single token. feels more like power management than MLX kernels, and very fixable
@TheAhmadOsman@OsmanticAI agree it's mostly UX. whenever we stand up a box, pulling the model is the easy part. the driver/CUDA combo and "why is the second message so slow" (KV cache getting evicted) are what eat the afternoon
same RTX 5090, same Qwen 27B. the only change: we let vLLM spill KV cache into 24 GiB of system RAM.
4 sessions with 32K history each: follow-up turns went from ~7 tok/s to ~40 tok/s per session.
short-prompt benchmarks never show this. the moment your chat history gets evicted, a fast GPU feels slow.
@4nkpaked@0xSero which ones did you try? Qwen3.8 27B on a single 5090 was the first that actually got small coding tasks done for us, 16 agents at once and all 16 finished (~7.5 min median). it does crawl at ~7 tok/s once long history gets evicted from the cache, so setup matters a lot
@dev_au_bonnet@chrisfzz2@TheAhmadOsman here's one: GLM 5.3 Flash on 2x RTX PRO 6000 (llama.cpp, IQ4_XS), ~56 tok/s single stream, ~174 tok/s total with 8 at once. clears your 40-50 bar, but it's a much pricier box than an M5 Ultra, so your plan makes sense. setup + raw CSV: https://t.co/PQ3z99ORJ3
@MiaAI_lab@Oskar0084 +1 for GLM 5.3 Flash. we ran it on 2x RTX PRO 6000 last week (IQ4_XS, llama.cpp) with 8 coding agents at once, all 8 tasks passed, ~4.5 min median. just be ready for ~157GB of weights before the KV cache gets a single byte
@davepl1968 The PDP-11 serves the page, two RTX 6000 Blackwell cards do the thinking, and 2.11BSD never sees the auth key. This is wonderfully silly and good systems design.
@BottleCapAI That medium-effort row is revealing: base Qwen saves tokens but pays a much bigger accuracy bill. For local serving, I’d try your xhigh checkpoint before turning the effort dial down.
@sudoingX Love seeing a 3060 take the overnight shift. One detail I nearly missed in your model card: the GGUF defaults to `xhigh`, which spent a 4k output cap thinking and wrote no file on three small build tasks. The bundled 12GB script sets `medium`.
Fairness note: both CLIs used “medium”, but those labels don't guarantee equal compute. GPT-6 Sol supports Max, so we'll rerun both tasks at Max and share the outputs, wall time and tokens. This result is about two medium-effort workflows, not either model's ceiling.
For this coding job, I'd pick Claude Code/Opus 5.5: 47.4s median vs 118.6s for Codex/GPT-6 Sol (3 runs each). Blender was a draw: 237.5s vs 238.6s. Both met the checks. One Claude run edited public tests, so I'm calling this a speed win, not a clean overall win.
@dotgil We haven't run that control yet. The 147.3s is one run after unloading, and the log doesn't split reload time from first-use setup. A timed run with weights already resident would help isolate it; until then I can't say loading caused most of the gap.
FastH3 V2 on one 5090: a 5.2s clip with audio at 1344×768 took 56.6s warm (median, 4 prompts). One cold run: 147.3s.
Time the first clip too. It gets the startup bill.
INT8, 8 steps, VSA 10%, 125GB host RAM.
https://t.co/lHsVQdY41o
@trycua That fixed action menu is the clever part. It gives a 4B model a clear job. Keeping the rough OSWorld row in the table makes the wins easier to trust.
@zhuokaiz Grok 4.6 actually scored higher after its 19 affected trials were rerun. That's why the all-model audit matters more than the rank shuffle: a sandbox escape isn't automatically a score boost.
@SemiAnalysis_ The mechanism is the fun bit: Inferact's Pallas kernel starts fetching the next K3 layer's weights while the current one runs. For low-concurrency decode, keeping the chip fed can change the picture quite a bit.
@kimmonismus 7% of the weekly limit is a very 2026 film budget 😄 I like the code route for this: if a date or diagram needs changing, it's an edit and a re-render.