@HPC_Guru Low expectations of the dozen+ 100,000+ H100 clusters to list huh? If the top 5 doesn’t have at least 1 AI “proof of concept” cluster the list will be fully irrelevant
@FelixCLC_ This is what, 5 months after the announced wse3. Presumably they had many more months internally to work on this. Is that the speed and efficacy we can expect when developing for their platform
@glennklockwood Not sure how I feel about it as the current roadmap for DAOS meets the Aurora need. Compute Node local storage and better dynamic cluster configuration feel needed for it to be competitive with lower performance enterprise options
@glennklockwood DAOS is practically up for grabs by any company that wants it. ANL put their IO wishlist down and Intel filled a lot of them as a sorry for delivery failure. Intel doesn’t seem interested in keeping it alive now that they don’t really have storage hardware
@inerati@jayyylma0 Are you sure that was a single computer and not just a tightly coupled NUMA memory sharing system? The OS can lie to you about reality
@CUDAHandbook@FelixCLC_ Given I interviewed with that firm just a few months ago and they were still running large amounts of skylake without an update plan it wouldn’t shock me if they didn’t t even try to validate the claim and were optimizing against the architecture. Diggin a hole
@CUDAHandbook@FelixCLC_ Mostly HFT folks from what I’ve been told, a few shops lamented being on Intel due to their back testing being dependent on it. But one of those firms seemed less interested in their tech stack and more interested in minimizing technical work and ran very old hardware