🦆🚀QuACK🦆🚀: new SOL mem-bound kernel library without a single line of CUDA C++ all straight in Python thanks to CuTe-DSL. On H100 with 3TB/s, it performs 33%-50% faster than highly optimized libraries like PyTorch's torch.compile and Liger. 🤯
With @tedzadouri and @tri_dao
🎉CUTLASS 4.0 is here-bringing native #Python support for device-side kernel design, for ops like GEMM, Flash Attention, and more, powered by the new CuTe DSL. For the first time, you can write high-performance GPU kernels in Python with the same abstractions, APIs, and performance as CUTLASS C++-no compromises.
The learning curve for writing optimized kernels is flattened: no more wrestling with C++ templates or long compile times.
CUTLASS 4.0’s Python support delivers: 👀
🏎️ Performance on par with C++ kernels
⏱️ 100x+ faster compile times
🤔 Intuitive, Python-native syntax
⚒️ No need for NVCC installs-just pip install nvidia-cutlas-dsl and go
🤝 Seamless integration with PyTorch and the broader Python ecosystem
📚 Improved documentation and a better debugging experience: https://t.co/Ji6iVDtDOA
Key features in #CUTLASS 4.0:
✅ CuTe DSL: Python-native, low-level programming model mirroring CuTe C++ abstractions (layouts, tensors, thread/data hierarchy)
✅ Supports for NVIDIA Ampere, Ada, Hopper, and Blackwell Tensor Cores
✅ Examples and Jupyter notebooks for rapid onboarding
✅ Further improved Blockwise and Groupwise GEMMs on Hopper and Blackwell
Whether you’re a researcher, student, or ML engineer, CUTLASS 4.0 with Python lowers the barrier to high-performance GPU programming and accelerates the path from prototype to production.
📝 Examples: https://t.co/eU9rlxkcVu
📗 Jupyter notebooks: https://t.co/SdGe6xFKDV
We’re excited to see what you build-feedback and contributions welcome. 🙌
(Note: CuTe DSL is currently in public beta and will evolve with community feedback. C++ APIs remain fully supported for existing workflows).
🚨🔥 CUTLASS 4.0 is released 🔥🚨
pip install nvidia-cutlass-dsl
4.0 marks a major shift for CUTLASS: towards native GPU programming in Python
slidehelloworld.png
https://t.co/pBLMpQAXHW
CUTLASS is in the center of the CUDA Blackwell release blog. As always, we work hand in hand with CUDA team to deliver the next level performance. https://t.co/UnZHT2WesG
🔥🚨 CUTLASS Blackwell is here 🚨🔥
3.8 release is loaded with support for new features of Blackwell, even an attention kernel 👀
Go check it out here: https://t.co/bII66C6oKc
Can't wait to see what y'all end up cooking with this over the next few moths and years 💚
@llllvvuu@hyhieu226 The goals here are :
1. Don't materialize intermediates in HBM
2. Optimally load / store tensors == ~1 time each of A, B, C from HBM
3. Ensure you can keep the GPU compute bound via efficient fusion
Dual GEMM attempts to do all 3.
FlashAttention is widely used to accelerate Transformers, already making attention 4-8x faster, but has yet to take advantage of modern GPUs. We’re releasing FlashAttention-3: 1.5-2x faster on FP16, up to 740 TFLOPS on H100 (75% util), and FP8 gets close to 1.2 PFLOPS!
1/
Find Carbon interesting?
Want a modern approach to language design?
WITH a compiler you can play with today?
AND is prioritizing safety?
AND has C++ interop?
WHY haven't you looked at https://t.co/wHOkiuAKS9 from @jntrnr and @awesomekling ?
I'm part of the pro bono litigation effort planning to quickly file a lawsuit challenging the onerous DOL wage rule impacting H-1Bs and PERMs. We're needing employers, employees and membership organizations to volunteer as plaintiffs. If interested, go to https://t.co/k9kyzRVRyC.
v1.6: native mixed-precision support from NVIDIA (~2x perf improvement), distributed perf improvements, new profiling tool for memory consumption, Microsoft commits to developing and maintaining Windows PyTorch.
Release Notes: https://t.co/7LMXQYlaAd
Blog:https://t.co/0cwKG82iOB
New @ICEgov policy regarding F-1 visa international students is horrible & will hurt the US, students, and universities. Pushes universities to offer in-person classes even if unsafe or no pedagogical benefit, or students to leave US amidst pandemic and risk inability to return.
A very sad day for US science and innovation. We will pay a hefty price for this demagogic insanity.
90% of my lab, myself included, is made of immigrants.
Immigration has contributed immensely to America’s economic success, making it a global leader in tech, and also Google the company it is today. Disappointed by today’s proclamation - we’ll continue to stand with immigrants and work to expand opportunity for all.
People are already so stressed out, stranded in the US with no Visa and No medical Insurance - and booking Evac flights via @airindiain is a nightmare !.
No clarity, horrible customer service, dead website links and phone numbers !.
FIX IT !
@PMOIndia @airindiain #AllowPvt
Trying to book evacuation flights via Air India is probably the worst experience one can ever had dealing with any business !
If you are incapable of providing ANY level of service - don't do it ! Zero leadership, Zero Service, Zero transparency !
#AirIndiaSucks