AI safety with @sheimersheim at @LASRlabs / On sabbatical from the National Renewable Energy Lab @NatLabRockies / @pomonacollege computer science ’23 / he/him
Model organisms (MOs) are frequently used as testbeds for mechanistic interpretability methods. We train 54 of them and find that their interpretability scores vary widely depending on how they were trained.
Your interpretability benchmark result may be a lottery. 🧵
Hey @SunCountryAir, left my laptop in a black padded case on a plane from Belize to @mspairport yesterday, SY712. Already filed a lost item report. Any advice?