#Qwen 3.8 Preview and #Kimi K3 are here, 2.4T and 2.8T parameter models pushing the frontier.
The gap to frontier models on benchmarks keeps shrinking. Now waiting for #GLM 5.3 to see where it lands.
#AI#LLM#OpenSourceAI#huggingface
New research: FlashAttention-4
FlashAttention-4 achieves up to 1.3x speedup over cuDNN 9.13 and 2.7x over Triton on B200 GPUs with BF16.
FlashAttention-4 co-designs algorithms and kernel pipelines for Blackwell GPUs, where tensor core throughput doubles but memory bandwidth and exponential units scale more slowly.
The techniques include fully asynchronous MMA operations, software-emulated exponential rescaling, and leveraging tensor memory to reduce shared memory traffic.
FlashAttention-4 achieves up to 1.3x speedup over cuDNN and 2.7x over Triton on B200 GPUs, reaching 1613 TFLOPs/s at 71% utilization.
Implemented entirely in Python via CuTe-DSL with 20-30x faster compile times compared to C++ templates.
Paper: https://t.co/wBiS51m8Bm
Learn to build effective AI agents in our academy: https://t.co/LRnpZN7deE
Spa: 21s gap, that too against a Redbull.
@Max33Verstappen pure 🦁 breed! 🛐
Hopefully, he will tie the record for consecutive race wins at his home circuit. 🤞🤞
@MisterDallas @FastestPitStop I guess he was concerned about the addition of sprint as it’s comparatively tedious for the drivers as they gotta travel every other week. Though sprints are fun to watch and a good way to catch the attention of the audience, but obviously tough/exhausting for the drivers.