How can we exploit feature universality between models?
We learn affine mappings (stitches) between two different sized models, allowing us to:
1. speed up SAE training
2. transfer steering vectors and probes
3. and study transferability of features
(1/9)
This work would not have been possible without the wonderful mentorship from @jack_merullo_, @alesstolfo, and @Brown_NLP. We'll be presenting the work at the ICML AIW workshop!
paper: https://t.co/3Mq8RCxL1m
(9/9)
How can we exploit feature universality between models?
We learn affine mappings (stitches) between two different sized models, allowing us to:
1. speed up SAE training
2. transfer steering vectors and probes
3. and study transferability of features
(1/9)
We also find that functional features (Gurnee et. al 2024 (https://t.co/B68p2o2WSr), Stolfo et. al 2024 (https://t.co/BEzSGCHajP)) that are known to be universal indeed preserve their role after transfer e.g. entropy features. (8/9)