Mechanistic Interpretability be like..
We wanted to see what the bread slice is thinking when it is in the toaster. Is it taking its singing equanimously? thinking religious thoughts? or plotting revenge? (Face it--a vengeful singed slice is the last thing we want!)
But alas, the bread slice doesn't talk! Its thoughts are intricately encoded in its singe pattern.
So we devised a clever method--of showing the bread slice to a random Joe, asking him to describe its feelings.
Joe says that the slice is having a religious experience involving Mother Mary.
But is random Joe describing these feelings right?
To check, we called random artist Jane and asked her to render these thoughts back to the external appearance of a (different!) bread slice.
We co-trained Joe and Jane until they learn to auto-encode the real truth about the slice's thoughts.
We are finding this technique to be a great way to understand what the bread slice is thinking. The tehcnique is not always (or even sometimes) correct, but it gives us a great window into publishing bread slice thought related podcast articles.
Hope you find this research useful. #AIAphorisms
We're Optimal Intellect, a research lab from the team behind CVXPY. Today we're introducing Moreau: a GPU-native solver that's orders of magnitude faster than the best existing tools.
This is the key difference between in-domain and out-of-domain generalization, and we still have not truly solved out-of-domain generalization. It just turns out you can build world changing technology by throwing so much data at things that the entire universe is in-domain.
@NeelNanda5 [2/2] From the discussion at ~6m, perhaps using multiple activations / PCA reduced how much the explanation can update *without* changing the underlying prediction, which may explain why it improved alignment?
@NeelNanda5 [1/2] This was awesome. I'm also excited about leveraging interp in training; my approach focuses on reducing degrees of freedom of model parameters & input w.r.t. explanation, because this provably prevents malicious compliance. Then training against these enforces alignment!
@bitsgopew@ninja_maths The channel below in general is an absolute goldmine, but the video specifically talks about the visualization in the tweet above, makes it very intuitive:
https://t.co/oE99IpG7B5
I also recommend checking out 4.2.2. where prof talks about the connection with normal equations
@jennyzhangzt My reviews have been constructive across the board; meaningful feedback from all 4. It's my first major conference submission too, so I was mentally prepared to have everything either torn to shreds or to be fed AI slop but I've been pleasantly surprised😃
Currently working on RFMs for tool use and other industrial tasks @personaaiinc. We have an early customer (Hyundai) and significant capital ($28M pre-seed). Looking to expand the ML team, dm me if you're interested in joining. Also taking cracked interns :)