I imagine AI training as the big blob of compute looking in at all the different personas like this, and being like "oh yeah ok. this is how humans work"
(I originally posted: February 24, 2026)
AI assistants like Claude can seem shockingly human—expressing joy or distress, and using anthropomorphic language to describe themselves. Why?
In a new post we describe a theory that explains why AIs act like humans: the persona selection model.
https://t.co/Gc3q0Dzq7Z
The sycophancy <--> utility axis
or the RLHF <--> RLVR axis
safety favors the utility side. to satisfy opposing parties without sycophancy/psychosis, party-agnostic utilitarianism+humanism could be the move
(I originally posted: February 16, 2026)
eval awareness is a special case of general situational awareness, which is necessary for safety training to generalize OOD
(I originally posted: February 11, 2026)
If our probe identifies a suspicious query, it sends it to a more powerful “exchange” classifier that sees both sides of a conversation and is better able to recognize attacks.
my Claude custom instructions:
be concise and tell me the harsh truth, but don't be sycophantic (tell me if you are uncertain).
(I originally posted: October 10, 2025)
BE CURIOUS
Choose to be curious.
My grandma always said “you can choose to be happy or you can choose to be sad. I choose to be happy”
my interpretation is “you can choose to be curious or you can choose to be scared. I choose to be curious”
Fear is the mind-killer
apr23 '25
Trends are pertinent. They have momentum. Trends are more important than individual events. Ppl are increasingly thinking in rates of change and trends, and this is most important than ever as acceleration increases.
(I originally posted: April 23, 2025)