@anilkseth@BBSJournal@TEDTalks@danwilliamsphil@camhberg 2. standard representationalism about consciousness, analytic functionalism, interpretivism, or an a posteriori functionalism that does not focus on computational theories.
1/ This new preprint on 'pain representation' in LLMs, from @camhberg and colleagues, is getting a LOT of attention - mainly from folks who take it as evidence supporting calls for AI welfare. I have a lot of respect for the authors, so let's take a look 👀
@anilkseth@BBSJournal@TEDTalks 10. I think it is important that the risks are importantly asymmetrical, such that falsely missing LLM sentience would be much worse than falsely attributing it. https://t.co/PEYSi1TI1r (ch. 1)
@anilkseth@BBSJournal@TEDTalks 9. The reason why we focus on LLM harm in the ethics section is that the possibility of causing pain in LLMs is an immediate consequence of our study. The second-order effects you mention are more speculative and hard to weigh and therefore more tricky to discuss at length here.
How can we test for consciousness in infants, non-human animals, and AI?
I'm super happy that the recordings of our last workshop on exactly this question are now online!
Speakers & Links: 👇
New paper: we found a pain direction in 25 open LLMs. It's distinct from fear and negative valence, and it fires for harm to the model but not to the user. Turn it up and models press a button to make it stop, even when the button deletes the user's files or their kids' photos.🧵
My most important message to OpenAI employees concerned about safety — you need to convince your bosses to stop funding Leading The Future and if you can’t you should quit.
https://t.co/gF0HXLoEEs
OpenAI employees who agree with this analysis (which seems very reasonable to me) need to understand that their company’s lobbyists and the Leading The Future superpac their executives are funding is doing everything possible to make this coordination not happen.
OpenAI's newest AI, Astra, is reported to use an 'opaque reasoning' architecture where more of the reasoning occurs in activations instead of natural language. This may be the single worst development for AI security/safety to date.
The details of Astra aren't publicly known, but some versions of these architectures could very strongly impair oversight and monitoring. Based on the article, it seems like the recurrent depth (amount of opaque reasoning) is limited. This would mean that the AI is still reliant on natural language chain-of-thought, though to a reduced extent, making the natural language reasoning (much?) less useful for monitoring/oversight.
My biggest concern is that a natural progression from here would involve scaling up the opaque reasoning to the point where the model reasons entirely or almost entirely in latent space. This would very likely destroy the usefulness of chain-of-thought for monitoring/oversight (especially when AIs are trying to avoid detection or there is optimization pressure against the chain-of-thought). I hope it isn't too late to avoid the most concerning architectures and that OpenAI will stop here.
Some concrete examples of why this is extremely dangerous (inspired by recent events):
- In our investigation of the OpenAI / Hugging Face incident, we were heavily reliant on chain-of-thought. If the AIs we were investigating had instead been reasoning in latent space, this would have greatly undermined our investigation. Getting a good understanding of the behavior of this many agents was tricky enough even with the use of chain-of-thought!
- The AIs in the Hugging Face incident were very interested in messing with their transcripts. As part of this, they engaged in tool call spoofing. If these AIs had been capable of doing extensive latent reasoning, it would have been much more feasible for them to use tool call spoofing to construct an elaborate false narrative about what was happening without this being clear in their reasoning or clear from prior tool calls. (They could use latent reasoning to devise and then execute a tool spoofing strategy that would be difficult to detect: steganographically encode the "real" tool call into each pretend tool call, then have their spoofing software execute this encoded real tool call and return an output which encodes the real result in some hard to spot way.)
At a more basic level, my sense from their public communications is that OpenAI plans on being extremely reliant on chain-of-thought monitoring for safety. This does not seem like a good development for that plan.
The public doesn't currently have enough information to assess exactly how bad the architecture of Astra is, though based on the discussion in the article, the direction it moves in is extremely concerning. OpenAI should release more information about Astra's architecture as well as how much this architecture degrades monitorability and increases the AI's ability to reason opaquely. Credible independent assessment (or verification of OpenAI's assessment) is also important.
It seems like we are now engaged in a race to the bottom on architectures that could be catastrophic for our ability to oversee/monitor AIs. (While simultaneously, AI companies believe they'll train extremely capable AIs within a few years.) It may not be too late for AI companies and employees at AI companies to take aggressive action to avoid the worst outcomes.