1/ Excellent, careful article. Some matched data from a close cousin of the card: 4× RTX 2000E Ada (16GB, 224 GB/s, 50W) on a WRX90 / Threadripper PRO 9955WX, every card at PCIe Gen4 x8, vLLM 0.29.
2/ On section 5 (TP without P2P): our first CUDA P2P test ran at 0.95 GB/s and returned NaNs, with AMD-Vi IO_PAGE_FAULTs in the kernel log. Setting amd_iommu=off fixed it: 12.2 GB/s and verified data, then ~13.15 GB/s on every GPU pair. The NCCL_P2P_DISABLE workaround may be masking an IOMMU translation issue rather than a hardware limit — worth checking dmesg.
3/ With P2P working, TP4 scaled well at batch 1 (tok/s, TP1 / TP2 / TP4):
• Llama-3.1-8B FP8: 24.2 / 44.8 / 80.0 (TP4 = 3.30× one card)
• Gemma-4-12B AWQ: 23.0 / 41.4 / 68.9 (3.03×)
• Gemma-4-12B AWQ + MTP depth 3: 135.7 at TP4
NCCL 8 KiB all-reduce: 10.6 µs at TP2, 18.0 µs at TP4.
4/ A simple model predicted every point within a few %: time per token ≈ bytes read ÷ (TP × ~205 GB/s) + 2 × layers × all-reduce latency + ~1.5 ms fixed. Applied to your Qwen3.8-27B NVFP4 TP4 row, it gives ~53 tok/s against the 49.0 reported.
5/ Agree with your TP2-vs-TP4 reversal under load: our TP4/TP2 gain fell from 1.78× at 1 request to 1.33× at 64. Faster cards should scale less well at TP4, because communication and fixed overhead don't shrink with card speed. Different card from yours (Ada, not Blackwell), so treat these as a comparison point, not a substitute.
@Hikari_07_jp It is worth trying Multi-Instance GPU (MIG) for 2 or 4 splits per GPU if model size allows in order to delay contention issues while maintaining a form of concurrency.
Thanks, a few more questions and I leave you in peace ! What is the MTM of your P8?
(e.g. 30HHxxxxx)
What BIOS version are you running? I bought three of these machines at $8k each and they are now over $40k, I was beginning to despair about not being able to get the Max-q(3) cards I have working.
@GPTWare Thanks, will try again (I got no support from Lenovo as I bought generic bulk Max-Q cards) how many cards did you manage in the one workstation, I guess two is the limit?
@GPTWare I tried this but struggled with non-Lenovo rtx 6000 gpus- did you connect the power cables directly to the motherboard. Would be great to know as I will try again.
@ivanfioravanti With MIG you can split 1 rtx pro 6000 into 4 x 24 gb or 2 x 48 gb logical cards - gives better results than concurrent users on 1 model