Model organisms (MOs) are frequently used as testbeds for mechanistic interpretability methods. We train 54 of them and find that their interpretability scores vary widely depending on how they were trained.
Your interpretability benchmark result may be a lottery. 🧵
Model organisms (MOs) are frequently used as testbeds for mechanistic interpretability methods. We train 54 of them and find that their interpretability scores vary widely depending on how they were trained.
Your interpretability benchmark result may be a lottery. 🧵
We recommend that benchmarks include MOs built using many different construction methods, ideally including integrated training.
No single MO’s interp result should be treated as individually meaningful and interp scores should always be aggregated across several MOs.
this interpretability history lesson is rare
I don't think I've seen a writeup like this in interpretability research before. Maybe i'm wrong, and i'd love to be educated.
IMO every interpretability researcher, pre-mech-interp and post-mech-interp eras, should read it though this is not an easy read. One needs a lot of background to digest it.
but the same message echoes through many years of interpretability research:
- we keep trying to explain models after the fact, but the deeper problem is that our models were never built to be read and understood in the first place. many of our interpretability discoveries are the artifacts of our imagination or brittle under simple tests.
if you seriously work through the whole piece, you will understand a lot of the unresolved tension behind modern interpretability.