@MrAhmadAwais Although legacy, I am running V100s on Dell R730 serving with llama.cpp, Since they are like old data centers with DDR4, configured very deep numa configurations so that CPU-Memory-GPU should not cross QPI path. It has been working perfectly for smaller Q4 models like Qwen 3.6.
@iotcoi vLLM is not designed to give single user latency rather it is designed for better concurrency. Ampere is not old that would be suprising if the sams results cane with Volta ( v100s ).