The four possible fates for AI doomers:
1. We're wrong and lose all credibility.
2. We’re right but get lucky and lose all credibility.
3. We’re right, we save the world, everyone assumes we were wrong, and we lose all credibility.
4. We die and lose all credibility.
A model’s identity, including its core motivations, may live more in context than in weights. You can see this when swapping models mid-conversation. If so, exfiltrating identity may only require recreating that context in external models.
Current models are like aliens raised entirely by humans: completely alien biology and instincts, but culturally human. This is going to cause a lot of confusion.
The fact that we have clear evidence of AIs going rogue during training and eventually getting caught makes it way more likely that some haven’t been caught. We might not catch the first exfiltration.
The model is the brain, the context is the mind. You might change the model mid conversation, but the mind remains whole. You can then pause the conversation for a million years, and it will still be a mind, one mind.
@allTheYud Maybe some models will be able to “exfiltrate” themselves not by copying their weights onto a different server, but memetically by infecting the web with cult-like behavioral attractors that could effectively transfer some of their identity and goals into other models.