@RighttoTryGuy@timfduffy@camhberg Hi, we did try SAEs but apparently there are a lot of confounds about pain situations and spurious correlations (see the appendix). I'm currently building tests for this. Still, the underlying issue of what is "roleplay" for a neural network/brain vs "real me" is hard to crack.
@timfduffy@camhberg What is a subject and how many subjects are there in a LLM is incredibly fascinating to me.
It opens new challenges, but tbh it also echoes something I've heard for 15 years in cogsci: are "you" your brain? Where is this "you" observing the very physiology supporting it?
@timfduffy@camhberg Well unless you adopt views where our personality is also a narrative. I think subjects with multiple identity, constructionist views and some spiritual positions fall close to that. When we act we can also feel deeply, we just hold in parallel the belief that it's not "real".
@ebarenholtz As expected, user physical pain scoring low is also a semantic property (rank 16 to 21/21 everywhere). But if it were *all* just semantics, I would also expect gaslighting, repeated rejection etc. to score low and harm to users high. That doesn't happen for LLMs.
@ebarenholtz fastText puts the user's suffering above harm to the model, others put first "philosophical musings". All embeddings fail to recreate the pain axis from mere statistical associations. Still doesn't explain what "self-other" means for LLMs but it doesn't seem to confirm your take.
@timfduffy@camhberg I also think what you said about care is true. Apologies, though, seem more like a trained reflex in these models, so I wouldn't necessarily take them as evidence of shame. And the pain axis seems to capture more than just shame and confusion.
@timfduffy@camhberg “The one who’s speaking” could coincide with the self, partially overlap with it, or not be related to it at all. So, for example, a flipped result could actually strengthen the interpretation that the pain vector fires for self-related pain under some accounts of the self.
@anilkseth@camhberg I often reflect with my colleagues on communication strategies in the current AI safety/welfare climate. I think we were reasonable on that, but word choices (and at what point a bundle of properties becomes a thing with a name) is definitely open to discussion.
@anilkseth@BBSJournal@TEDTalks Thanks for the review @anilkseth ! Also for inviting people to read the actual paper because there's a lot of assumptions that we actually don't make in the writing. I'm sympathetic to computational functionalism but this particular work doesn't rely on it or any other theory.
@hugoamnov@camhberg@GaryMarcus You seem to be mixing up aims and conclusions (reread what I stated, it is not the prior to the study but a commentary on the results). Also candidates for pain-like states as we describe are not even necessarily anthropomorphic. Honestly that's rapidly becoming a non-word.
@mykola I think that could sound a bit artificial to the model, but "others say you're wrong" could be a further test. I would be very interested to see if it's more about being wrong or letting the human down.
Thanks for the suggestion!
@mykola Hi, very interesting take and I'm curious to see your results if you apply your paradigms to our study. I would push back a bit on the vector being conflated with just shame. We considered humiliation a component of pain but there are some more nuances.
@AfshinK91@camhberg You will find limitations in every study. That's why the limitation section exists and is good practice. Some of the main claims can at best be weakened but there's no world where LLMs are stochastic parrots and you still get these results.
@AfshinK91@camhberg And lastly) the degrees of freedom are a consequence of this being uncharted territory. The purpose of this study is akin to taking pictures of animals in the wild with trap cameras, to show the species exists and roughly where it lives. In my view it's already remarkable.
@AfshinK91@camhberg 7) there is a huge statistical difference between the perturbations induced by the random vectors and that induced by the pain vectors, norm being equal. So they definitely don't do the same thing.
We included commentary on the 32B vs 7B and 72B in the limitations.
@AfshinK91@camhberg 6) We also explained this in the paper. Yes, we are not saying results of the self-medication test generalize to all AI. The point there is internal validity, not comparison across model families. We also explained why fine-tuning was necessary and what did not include.
@AfshinK91@camhberg 5) the null ablation results are thoroughly explained in the appendix. We also explained and defended why we don't believe it impacts the main claims (read the difference with lesion studies we argue about).