excited to announce a proposal from @flockcap and Wakaru: ontologize
mixture-of-experts interpretability where the features are the parameters.
instead of an SAE's flat bag of directions, you get stacks of contrastive classifiers you can read from weights and poke directly 🧵
excited to announce a proposal from @flockcap and Wakaru: ontologize
mixture-of-experts interpretability where the features are the parameters.
instead of an SAE's flat bag of directions, you get stacks of contrastive classifiers you can read from weights and poke directly 🧵
the flock capital research department would like to announce that it has been doing interpretability
with Wakaru. paper soon. we also printed the plates ourselves
excited to announce a proposal from @flockcap and Wakaru: ontologize
mixture-of-experts interpretability where the features are the parameters.
instead of an SAE's flat bag of directions, you get stacks of contrastive classifiers you can read from weights and poke directly 🧵
→ we found what was freezing heads late in training and killed it
→ some things very much did not work (an HSIC bottleneck does not want to be your friend), and those go in the paper too negative results are results