c’est la vie.
Some good Art & Music festivals
Grad student @virginiatech.
Ex-@NERSC intern @Berkley_lab
Researching on Supercomputers #HPC for #AI#bigdata.
How AI empowered Paul Conyngham to create a custom mRNA vaccine to cure his dog’s cancer when she had only months to live. The first personalized cancer vaccine designed for a dog:
First, he burned $50 billion on a cartoon metaverse that no one uses. Now the AI hype train is derailing, and the "downsizing" begins.
All those LinkedIn thought leaders who were "humbled and excited to be part of the AI journey" are about to be humbled and excited to update their resumes.
Are you passionate about weather & climate science 🌧️? Join a Summer School near Hamburg to explore ICON! Learn meteorology basics and tackle code challenges using #HPC and modern programming under expert guidance.
Sign up at https://t.co/zuzhVJYSkn #SummerSchool 💻⚡
#DeepSeek just dropped a file system for distributed #AI training over SSDs
on #RDMA networks! 🧠 Check it out here:
https://t.co/vlbN1lYVu5
Does their system support clairvoyant prefetching 🎇 ?
131,072 Nvidia GPU cluster coming from Oracle “Oracle today announced the first zettascale cloud computing clusters .. Oracle Cloud Infrastructure(OCI) is now taking orders for the largest AI supercomputer in the cloud—available with up to 131,072 NVIDIA Blackwell GPUs.”
I must've missed this earlier this year, but thrilled to see that Charm++ is no longer under a noncommercial license.
Since v8.0.0, Charm++ is using the Apache-2.0 license, with LLVM exceptions. #hpc#opensource
https://t.co/OjH3HgzlRy
Llama 3.1 405B could be the catalyst for much greater #AMD adoption for AI inference 📈
@AMD's MI300X may be uniquely suited to cost-effective Llama 3.1 405B inference. Its 192GB of memory allows a single 8xMI300X node to serve Llama 3.1 405B in its native FP16 precision - whereas two 8xH100 nodes are required on @nvidia.
As we have previously covered, a single NVIDIA 8xH100 node only has 640GB of memory - not enough to hold Llama 3.1 405B’s full 810GB of FP16 weights in memory at once. This means that providers are forced to deploy two 8xH100 nodes with interconnect to serve 405B in FP16 precision, forcing them to accept a significant cost and complexity penalty.
Nvidia’s future H200 and B100 come with 141GB and 192GB of high bandwidth memory respectively - but unlike those, AMD MI300X is available now. @LisaSu noted on AMD’s Q2 earnings call that AMD was demand-limited on MI300X for the remainder of 2024. Will Llama 3.1 405B alone flip that narrative?
We are starting to see adoption and support increase. Both @FireworksAI_HQ and @LeptonAI are hosting Llama 3.1 405B on AMD MI300X chips. They stand out as the lowest cost providers of Llama 3.1 405B. However, it is important to note they are serving the model at FP8 and INT8 precision respectively.
Furthermore, projects like GPU.cpp from @answerdotai (@jeremyphoward, @austinvhuang) are making it easier than ever to write and run portable code across different chip (hardware & software) architectures - decreasing the CUDA lock-in.
What is your view? Long #AMD?
What is SoftBank going to do with Graphcore? Has the time passed for Arm to resell AI accelerator blocks as it does CPU blocks? $ARM $NVDA $AMD $INTC
https://t.co/tRJkDwgpm7
TSMC won Google, Qualcomm 3nm orders away from Samsung Foundry as the latter struggles with yield and power efficiency issues, South Korea media report, noting 7 major firms have chosen TSMC 3nm: Nvidia, MediaTek, Intel, Apple, AMD. Google’s 4th gen Tensor processor was made by Samsung Foundry, but the 5th gen will be made on TSMC 3nm. 1/2 $TSM $GOOGL $QCOM $NVDA $AAPL $AMD $INTC #Samsung #semiconductors
For those wondering the # of parameters in GPT 4. Jensen mentions GPT is 2Trillion parameters and 8Trillion tokens at Computex 2024 (1:05:50)
https://t.co/vAb856adku
#hpc#ai#llm#generativeai
🎉 @neuralmagic & @cerebras unlock the future of AI with the first highly sparse, foundational LLMs! We removed 5B parameters from a 7B parameter model, maintaining accuracy. Slash costs, boost performance, save energy: https://t.co/XVOO7vSLrG…
#ai#llms#generativeai
#OFFMareNostrum4
📴Así fue cómo apagamos el supercomputador MareNostrum 4.
🫶💻Ofreciendo servicio a la ciencia y a la sociedad desde julio de 2017.
▪Gracias
▪Gràcies
▪Thank you
We’re sharing Project Astra: our new project focused on building a future AI assistant that can be truly helpful in everyday life. 🤝
Watch it in action, with two parts - each was captured in a single take, in real time. ↓ #GoogleIO
It's here!
Introducing Llama 3 by Meta. 8B and 70B pretrained and instruction-tuned models are available.
https://t.co/68fc6TNN9E
Details in the thread: