Diffusion language models develop 𝐬𝐮𝐛𝐥𝐢𝐦𝐢𝐧𝐚𝐥⏱️
Despite receiving no explicit timestep, mask-based DLMs encode denoising progress in a low-dimensional latent structure. We can steer this clock predictably and watch the model correct perturbations across layers.
(1/6)
Israel planted explosives in a school for disabled children at dawn today in South Lebanon and blew it up.
Not a military target.
A school for disabled children — perhaps the only one in the country.
In what universe is this considered “self-defense”?
This is Israel, world
An Israeli soldier ordered a Palestinian youth to continue walking, then used him as a target for long-range shooting practice before killing him.
Tell us what you think of their actions!
Repost please
🚀 NVIDIA AI-Q Deep Research Agent just topped the DeepResearchBench I & II leaderboards!
Learned a ton working on this w @raja_biswas@DivyanshJain92, @JFPuget David Austin, Ivan Sorokin Ajay Torve and Chantal Rose!
We also just released a detailed blog (link👇)
*Attention Sinks in Diffusion Language Models*
by @MaximoRulli@devoto_alessio@simone_petruzzi@fabreetseo
We study attention sinks in diffusion models, with many findings including "moving" sinks and a strong robustness of the underlying models.
https://t.co/217DMDDVGJ
What do a 5-year-old and a Diffusion Language Model have in common 👦🤖?
Neither can keep their attention in one place!
We show that DLMs shift their attention sink tokens across denoising steps
This anchor flexibility makes the models more robust when masking sink tokens👇👇
We build on prior work on attention sinks (@Guangxuan_Xiao, @tensorqt, @RuscioValeria, @fedzbar...) with an empirical analysis of diffusion LLMs. We find that their iterative denoising process makes sinks move, showing a unique dynamic behavior.
Paper: https://t.co/F6NXq4meol
Attention sinks emerge in Diffusion LLMs too, but they move around! 👀 We find that: sinks shift positions during generation & Diffusion models are surprisingly robust to masking them! ⚡️ w/ @MaximoRulli@simone_petruzzi@s_scardapane@fabreetseo
Expected Attention compresses the KV Cache by estimating the attention score from future queries! Both during prefilling & decoding 🚀 All code released in our KVPress library & Leaderboard!
*Into the land of automatic differentiation*
Material is out! A short PhD course for the CS PhD in @SapienzaRoma covering basic and advanced topics in autodiff w/ slides, (rough) Notion notes, and two notebooks including a PyTorch-like implementation. 😅
https://t.co/EgMWuG7LMc
*MoE Graph Transformers for Interpretable Particle Collision Detection*
by @d_genovese@devoto_alessio@StefanoG
We propose a MoE graph transformer for particle collision analysis, with many nice interpretability insights (e.g., expert specialization).
https://t.co/Rj5mBNxc8B
🌟 Excited to Share Our Latest Work! 🎥🎶
Here we present Stable-V2A: Synthesis of Synchronized Sound Effects with Temporal and Semantic Controls
arxiv: https://t.co/zzqQtTQ3jy
web: https://t.co/WgM1SZPu7p
Our team is hiring Student Researchers @GoogleDeepMind for '25!
🧑🔬 Interested in understanding reasoning capabilities from first principles?
🧑🎓 Currently studying for a BS/MS/PhD?
🧑💻 Have solid engineering and research skills?
🌟 We want to hear from you! Details in thread.
*Alice goes to a differentiable wonderland!* 🔥
I published a short free book on the design of neural networks, from convolutions to transformers, SSMs, and a few other topics.
As a bonus, I tried to make it looking nice - any feedback is appreciated! https://t.co/fBVqrTofug
Ma scusate, la RAI che a novembre si è accordata per pagare circa 1 milione di euro Pino Insegno dopo la cancellazione del suo programma per flop, è la stessa RAI che ha considerato troppo onerosi i 1.800 euro da dare a Scurati per il suo monologo e che per questo l’ha cancellato?
Ma come fa certa gente a farsi prendere così in giro dal Governo Meloni? Sul serio.
Il 25 dicembre 1974, Carlo Ciccioli, allora membro del Fronte della Gioventù, sparò cinque colpi ferendo alla gamba Paolo Tomassoni, appartenente a un movimento extraparlamentare di sinistra. Successivamente, il reato fu amnistiato. Ciccioli è candidato alle europee. #matrice
Remember our leaders claming that sanctions will crush Russia? Russia growth prediction is 2.6% (IMT). the fastest after China and India. Russia trades with the planet; the West is isolated in its fading illusion of world domination. We must collaborate, not fight.