We trained a decoder to read the internal activations of an LLM and answer questions about what the model will think about or do next.
We find that this decoder can understand LLM behaviors, even when the model itself is confused! (for instance, if the model has been jailbroken)
@vvhuang_ i think clothes are a statement of self expression! i am a believer that clothes (and all other forms of self presentation) should be intrinsically motivated so if deliberately dressing with 0 effort is what you truly want to do then you should do it
and... I am turning down the NSF GRFP to pursue a one-year Fulbright research fellowship at the AEI Potsdam in Germany! I'll be starting at Princeton in 2025 :)