@gaunernst@maharshii Although I'm still unsure -- like peak compute and BW aren't that different from an A100 (~ 312 TF/s @ fp16 tensor, 1.5-2 TB/s)...
But thanks for explaining. And awesome job!
@fleetwood___@Guangxuan_Xiao If you introduce dilation into sliding windows, the receptive field can grow up to exponentially, instead of linearly.
They can eventually "see" globally given enough depth.
https://t.co/ahNV32VY4G
Watch my talk about NATTEN on @GPU_MODE this Saturday at 3PM ET / noon PT.
I'll go over all the exciting new features we shipped very recently, especially our Hopper and Blackwell FNA kernels, now speeding up video / world models by up to 2.6X e2e!
https://t.co/Wn2wwekOVZ
https://t.co/UjezOE9WEJ marks the start of a short series of blogposts about CUTLASS 3.x and CuTe that we've been meaning to write for years. There are a few more parts to come still, hope you enjoy!
NATTEN 0.21 ships Hopper and Blackwell FNA backward kernels, enabling much faster training on those architectures.
Accelerate your training workloads with NATTEN today!
https://t.co/JXYgkGIwZo
All of it is open source, and you can simply use `--natten` when running T2W or V2W:
https://t.co/S8qHyfcCTS
For more information on NATTEN, see the project website: https://t.co/uUV7vx2frp.
(5/5)
Cosmos-Predict2 meets NATTEN.
We just released variants of Cosmos-Predict2 where we replace most self attentions with neighborhood attention, bringing up to 2.6X end-to-end speedup, with minimal effect on quality!
https://t.co/S8qHyfcCTS
(1/5)
Sparsity, and the specific pattern were selected through layer-wise profiling. Some layers can tolerate as high as 98% sparsity, some as low as 55%, and in only one case we retained one self attention layer.
(4/5)
We are releasing a major NATTEN upgrade that brings you new Hopper & Blackwell sparse attention kernels, both capable of realizing Theoretical Max Speedup:
90% sparsity -> 10X speedup.
Thanks to the great efforts by @AliHassaniJr & @NVIDIA cutlass team!
https://t.co/g1QBlrYeZN
NATTEN 0.20.0 brings your our Hopper and Blackwell FNA kernels, Strided NA, improved user experience, profiling toolkit, and more!
Oh, and we have new docs: https://t.co/uUV7vx1HBR.
Run your sparse local attention at the Speed of Light today!
Wondering what's happening with NATTEN in 2025?
Check out Generalized Neighborhood Attention!
Spoiler: NATTEN gets a new stride parameter, we made a simulator for all your analytical studies, AND a Blackwell kernel!
Keep reading for more...
(1 / 5)
🚨🔥 CUTLASS 4.0 is released 🔥🚨
pip install nvidia-cutlass-dsl
4.0 marks a major shift for CUTLASS: towards native GPU programming in Python
slidehelloworld.png
https://t.co/pBLMpQAXHW
@YouJiacheng Block sparsity, when used correctly, can come at no additional cost. You just need to build the block-sparse mask on top of the SoL/SOTA kernel.
See GNA where we get perfect FLOP-wise speedup (i.e. 90% sparsity -> 10X speedup).
https://t.co/ruMvuWECMk