🚨NEW PAPER OUT 🚨
Excited to share our latest research initiative on in-context learning and meta-learning through the lens of Information theory !🧠
🔗 https://t.co/Tj5cYudDwy
Check out our insights and empirical experiments! 🔍
SAEs fail at OOD tasks. Why?
Features in superposition are linearly representable but not linearly accessible. Instead of discarding sparse coding, we embrace the geometry of superposition and use methods equipped to handle the nonlinearity it induces.
Mechanistic interpretability aims to understand models — and the more superhuman or incoherent they become, the more we need that understanding to be reliable. We propose a framework for this, drawing on established tools from causal reasoning and statistical identifiability:
🧵
Very excited to release a new blog post that formalizes what it means for data to be compositional, and shows how compositionality can exist at multiple scales. Early days, but I think there may be significant implications for AI. Check it out! https://t.co/bxHUgljfK2
Lol, let me tell you, some reviewers are so lazy, they didn’t even wait for AAAI's novel A- powered peer review, and have already been using crappy LLMs to write their reviews for a while
link : https://t.co/qvusTvO32z
@TheXeophon @roydanroy 100% agreed. Those people should be banned and their name publicly exposed. We completely under-estimate the negative the impact of LLM-generated reviews being published in the wild. This behavior is so detrimental to the motivation of good faith researchers.
8/ A proposed system:
- Authors / AC flags suspicious reviews (with help from watermarking tools), that's the easy part.
- Panel of volunteers = the jury decide the guilt of the defendant
- A "judge", part of the conference committee decides the sentence
1/ A Call for Action: Introducing the Court of Scientific Integrity ⚖️
NeurIPS, ICLR, ICML reviews are broken. Many of us have gotten fully generated LLM reviews —zero effort, zero accountability. It's disheartening, and it's only getting worse. Let's talk solutions 🧵
7/ This is about protecting the integrity of research and the mental health of researchers. Getting annihilated by a fake review that make no sense, written in 10s, after hard-working for many months? That's devastating.
@mgostIH A compression algorithm can be derived from any prediction model p.
The amount of bits required to encode a sample is the log-likelyhood of that sample under p, so it totally makes sense that the lower your risk is, the better you can compress data.
https://t.co/FlgEGDxV7j
@scaling01 Not the worst one I've seen, my take is :
- take the pattern on the right and replicate it to the left.
- if along replication, you fall exactly on one of the square on your way, stop here and change the color of the whole line pattern
That's it.
1\ Hi, can I get an unsupervised sparse autoencoder for steering, please? I only have unlabeled data varying across multiple unknown concepts. Oh, and make sure it learns the same features each time!
Yes! A freshly brewed Sparse Shift Autoencoder (SSAE) coming right up. 🧶