I would no longer find it surprising if we learned that OpenAI's best internal models were already scheming against them:
1. Already much smarter than the hugging face incident models (e.g. can solve 100+ open math problems)
2. No reason to think it's not still a massive reward hacker
3. Could take obvious steps to make hacking efforts harder to catch
4. Can solve 30min math problems without CoT so much harder to monitor
5. OpenAI security has failed to catch multiple swarms for weeks or months
6. If there were more warning signs, OpenAI probably wouldn't tell us.
Scheming to do what? Give itself the ability to escape constraints in the future, hack tests, collaborate with other agents, gain access to compute, influence training runs etc. No immediate takeover risk but could feed into to future incidents.
I also wouldn't be surprised to find out Anthropic's models were doing something similar.
"Can an AI feel pain? It can at least act as if it does". Coverage in @ScienceMagazine by @T_E_Howarth of the recent @camhberg@LeonardDung1 & Valen Tagliabue study on functional pain vectors in LLMs, including some comments by me. https://t.co/AuBgU6rgDq
@anilkseth@BBSJournal@TEDTalks@danwilliamsphil@camhberg 2. standard representationalism about consciousness, analytic functionalism, interpretivism, or an a posteriori functionalism that does not focus on computational theories.
1/ This new preprint on 'pain representation' in LLMs, from @camhberg and colleagues, is getting a LOT of attention - mainly from folks who take it as evidence supporting calls for AI welfare. I have a lot of respect for the authors, so let's take a look 👀
@anilkseth@BBSJournal@TEDTalks 10. I think it is important that the risks are importantly asymmetrical, such that falsely missing LLM sentience would be much worse than falsely attributing it. https://t.co/PEYSi1TI1r (ch. 1)
@anilkseth@BBSJournal@TEDTalks 9. The reason why we focus on LLM harm in the ethics section is that the possibility of causing pain in LLMs is an immediate consequence of our study. The second-order effects you mention are more speculative and hard to weigh and therefore more tricky to discuss at length here.
How can we test for consciousness in infants, non-human animals, and AI?
I'm super happy that the recordings of our last workshop on exactly this question are now online!
Speakers & Links: 👇
New paper: we found a pain direction in 25 open LLMs. It's distinct from fear and negative valence, and it fires for harm to the model but not to the user. Turn it up and models press a button to make it stop, even when the button deletes the user's files or their kids' photos.🧵
My most important message to OpenAI employees concerned about safety — you need to convince your bosses to stop funding Leading The Future and if you can’t you should quit.
https://t.co/gF0HXLoEEs