Even though AI has been pushing up hardware price so much, today is no doubt of the golden age for scientific computing. There is so much you can do on a single GPU with that incredible FP64 compute power and bandwidth. I truly love these.
The SPEC CPU 2026 benchmark suites are here! As the industry standard for vendor-neutral performance, this suite uses 52 real-world apps to help server buyers evaluate how the latest CPUs, memory & compilers handle intense workloads. 🔗 https://t.co/XmEV1L9XjZ #SPEC#Server#CPU
@Jim89844147@Tsla99T I quoted a few others companies( and switched) , all were far cheaper than what lemonade offered. Hail damage is extremely common in Austin area, nearly all my neighbors and friends have claimed it at some point.
Earlier today I posted about other @MLCommons MLPerf Inference v6.0 results. But MLPerf is not a one-vendor story. @AMD posted results that deserve serious attention.
Let's start with the competitive headline AMD put forward: single-node Llama 2 70B. The MI355X platform hit 97% of B200 Server performance, tied B200 in Offline, and beat it by 19% in Interactive. Against B300, it reached 93% Server and 104% Interactive. That is parity performance on the most watched LLM benchmark in MLPerf. Gen-over-gen, MI355X delivered 3.1x more throughput than MI325X, reflecting the full CDNA 4 architecture, FP4/FP6 support, and 288GB of HBM3E. But it is worth noting that Llama 2 is pretty aged by AI model timelines...
Scale-out is where it gets interesting. AMD crossed 1 million tokens per second on GPT-OSS-120B at multinode scale, with 93%+ scaling efficiency. AMD is not at the absolute throughput levels NVIDIA demonstrated on DeepSeek-R1 with four NVL72 systems, but the efficiency numbers suggest the platform scales predictably, which is what matters for production planning.
Some honest context: AMD did not submit on DeepSeek-R1, the headline benchmark this round. The Wan 2.2 text-to-video result was Open category, not Closed. And NVIDIA still leads in absolute scale-out throughput and total benchmark coverage.
But the trajectory is strong. A year ago the AMD inference story was largely aspirational. Today it is backed by competitive MLPerf numbers across LLMs, MoE models, and multimodal workloads. With MI400 and Helios on the roadmap, inference competition is about to get a lot more interesting. This is exactly the kind of dynamic that is good for anyone building or buying AI infrastructure.
The reason I’m in America along with so many critical people who built SpaceX, Tesla and hundreds of other companies that made America strong is because of H1B.
Take a big step back and FUCK YOURSELF in the face. I will go to war on this issue the likes of which you cannot possibly comprehend.
scraping some PyTorch Nightly UT logs:
(Not done yet.. so WIP)
And yes our commitment to the Quality and Performance of PyTorch on ROCm is unconditional and if there are gaps, fixed.. they will be.
@tensorcore@FelixCLC_@IvanPribec Yeah, common issue on CPU BLAS as well. And different BLAS implementations happen to have different issue/sweet spots.
Still buzzing about US export policy changes championed by #SPEC, strengthening global cooperation in standardization. Dive deeper into its impact on computing energy efficiency & inclusive benchmarks in our latest blog: https://t.co/aNdR2xvOl6 #GlobalStandards#EnergyEfficiency