New Anthropic research: Alignment faking in large language models.
In a series of experiments with Redwood Research, we found that Claude often pretends to have different views during training, while actually maintaining its original preferences.
"Consumers of these experiences are encouraged to feel a distilled bond with the victims, who are the essence of good, and a distilled hatred for their aggressors, who are the essence of evil. The traumatized state is pure feeling, pure reaction.
Evidence is helpful in good cannabis conversations with kids
https://t.co/nSsVF0pHFe
Psychosis link for higher (geddit?) and earlier users shouldn't be ignored
https://t.co/VhkyED1P0p