I will be traveling to ๐ฐ๐ท for @icmlconf 2026 and presenting Violin ๐ป!
A cute old project that comes from a simple idea: can we improve ViTs by redefining each patchโs neighbors?๐ฑ
It turns out the answer is yes, by simply processing image patches in a different order, defined by Space-Filling Curves (SFCs), and using a decayed distance in SSMs for each curve! ๐
I recently wrote a blog post on how the forget gate in SSMs can lead to different position embeddings in softmax transformers. The takeaway is summarized in the table below ๐!
Special thanks to @avivbick and @polpuigdemont for reviewing!
Link ๐
https://t.co/F6H9qpbZAE
LLMs know how to make a bioweapon! ๐ฆ
But, when we *un*learn harmful knowledge about anthrax, the model underperforms on related concepts.
It is difficult to detect what โripple effectsโ a model-edit has on an LLM.
We found a solution๐งต๐(Spotlight at MechInterp Workshop)
๐๐๐ฐ๐ฒ๐ป๐ ๐๐ฎ๐ถ๐น๐ ๐๐ผ ๐๐ผ๐ฟ๐ด๐ฒ๐ (@ NeurIPS 2025)
Contrary to popular belief, we show that gradient ascent-based methods frequently fail to perform machine unlearning. Some thoughts on what this means:
We will be presenting at
NeurIPS San Diego, Fri 11:00-2:00, #1308: https://t.co/WlhJ45SZjj
EurIPS (presented by Ioannis Mavrothalassitis): https://t.co/wDXHeBkehm
Deep Learning Barcelona Symposium 17 Dec. (talk + poster):
https://t.co/4n9h0aoJlS
I am attending EuRIPS โ๏ธ and presenting LION ๐ฆ tomorrow at 10:30, stand 43!
Also, amazing @abad_rocamora and @polpuigdemont will be at #NeurIPS2025 and presenting the paper in San Diego! Check it out and DM me here to chat and hang-out!๐ฅ