Congrats to @GoogleDeepMind on the launch of DiffusionGemma.
The model generates 256 tokens in parallel per step, delivering 150+ TPS on DGX Spark, and 1,000+ TPS on a single H100.
We're supporting it from day one with:
• BF16 and NVFP4 checkpoints on @huggingface🤗
• Free GPU-accelerated endpoints on https://t.co/6T0R9P7EXS
• @vllm_project support with FP8 precision
Get started with DiffusionGemma on NVIDIA: https://t.co/vurk7GCQUs
Yapay zekada bellek duvarı yıkılıyor! Google, TurboQuant ile "ekstrem sıkıştırma" devrini başlattı.
✅ 6 Kat Bellek Tasarrufu
⚡️ 8 Kat Hız Artışı
🧠 SIFIR Doğruluk Kaybı
Llama ve Gemini gibi modelleri artık çok daha uzun bağlamlarla (context), çok daha hızlı çalıştırabileceğiz
Introducing TurboQuant: Our new compression algorithm that reduces LLM key-value cache memory by at least 6x and delivers up to 8x speedup, all with zero accuracy loss, redefining AI efficiency. Read the blog to learn how it achieves these results: https://t.co/CDSQ8HpZoc
@burak_tamac Örneğin senin bulduğun pahalı problemi buldun ve çözdün claude ile. 2 gün sonra diğerleri de claude ile çözerse artık o pahalı problem ucuzlar mı?