Two weeks ago, we got early access to @AMD's Instinct MI455X.
432 GB of HBM4 and up to 23.3 TB/s bandwidth on a single GPU🤯.
We achieved 99.5% success rate across our core🤗 Transformers model suite after working closely with AMD.
Learn more about it👇
https://t.co/vhxQzcATNR
AMD GPU users: the longer your context, the more this attention kernel beats PyTorch SDPA - and it's one line in Transformers.
I ported ROCm/aiter's Composable Kernel (CK) FlashAttention to the 🤗 Kernels Hub; our first CK-based kernel on the Hub. 🧵
attn_implementation="kernels-community/aiter-flash-attn-ck" - fetched, cached, and loaded automatically.
• Built for AMD Instinct (MI300X / gfx942, MI355X / gfx950)
• Attention-bound regime: up to 1.25× vs. SDPA at 16K on Qwen3-0.6B, and the lead grows with sequence length
This PR i've been working on for a while (6 months) finally got merged, you can now run Transformers on your graph engine/runtime of choice because all of Transformers can be exported directly to @PyTorch Dynamo and Executorch, @onnxai/@onnxruntime (and @OpenVinoAI very soon)
The MSA kernel from @MiniMax_AI team is available through the 🤗 Kernels lib and boy, is it speedy 🔥
It currently powers the MiniMax-M3-VL in 🤗 Transformers!
Go to [hf dot co / kernels / MiniMaxAI / msa] to get started!
Congrats https://t.co/dfGBgP72pk for winning 🥇 at the AMD Hackathon - this Reachy Mini Wireless is on its way to meet you!
@Debdeep_Banerje built Lumi on @AMD Instinct GPUs - A Compassionate AI Companion for Dementia & Alzheimer's
Try it at https://t.co/icWj6xTehN
🔥 AITER's FA3, RoPE, and MegaBlocks kernels are now on the 🤗 Hub and integrated into Transformers.
No model code changes required. Just enable Hub kernels and go.
🚀 Up to 1.63× faster GPT-OSS-20B inference on MI300.
#AMD#ROCm#HuggingFace
not appreciated enough but it’s honestly crazy how reliable huggingface infra is, in the age of agents, jobs and builders pulling models, datasets and kernels from all around the world !
this is the biggest wake-up call to protect and nourish open source AI
if you don't build out sovereign and independent models+infra closed labs will patronize you to an insulting degree
yesterday we brought deepgemm fp4 and megamoe fp4 support to transformers https://t.co/uGmGsP2Ihq
today we are bringing the same level of support to xpu and rocm through triton
https://t.co/5KstIMBmeq
I'm excited about the new @amd Ryzen AI Halo because we need more local hardware for AI builders!
There's something fun and exciting about building on your own machines rather than sending to the cloud! Should we do our own @huggingface hardware for AI builders at some point?