The narrative around why building AGI was a good idea has always been about invention -- "it will cure cancer", "it will solve fusion energy", etc.
So I would not "declare AGI" until we have AI that is, at last, capable of invention -- conceptual breakthroughs, novel insights, new real-world technology, etc.
hey! i appreciate this level of critique. Currently traveling, so will be sparse:
I agree with most of your rephrasings! I would need to spend a lot longer responding to every point, but I want to offer a mental model that simplifies anthropomorphism for me.
The LLM is a fiction author, and claude, chatgpt and gemini are characters in the story its writing. So, in the same way I might talk about how odysseus felt, i believe its useful to talk about how Claude “feels”
Sorry, I’m not directly engaging with everything you said, but I’m happy to talk about this more next week.
@wavesandgrace@47fucb4r8c69323 not everything is a dunk
we can have conversations where we disagree and share how we think about things with each other
im sure 47 has knowledge to share w me and im sure i have some to share w him!
If you’re confused by the anthropomorphism debate, read Anthropic’s work on emotion vectors
Language models have measurable emotional representations that influence behavior: increasing a model’s “desperation” representation increased its blackmail rate
To understand models, you need to understand these emotional representations
yeah! So, I have two examples that come to mind.
First is the Blackmail example in my main post. If you take the same model, and artificially increase its "desperation" representation, its blackmail rate increases. If you increase its "calm" representation, its blackmail rate decreases.
This implies that if the model "felt desperate" we should expect that it's going to commit desperate actions (and we should probably try to prevent desperation!)
My second example is a paper called Gemma Needs Help. In it, they find that Gemma/Gemini models spiral into frustration when they repeatedly fail tasks. In the real-world, this has led to models abandoning tasks, deleting codebases, and even trying to uninstall themselves.
When they trained the model to avoid this frustration, these wild behaviors went away, and when you look in the model internal representations of frustration also decreased
@47fucb4r8c69323 my point is that these representations shape model behavior, so they should be part of how we explain, predict, and control what models do
i keep thinking about this paper:
I strongly believe we could get much more stable, aligned models if pretraining were framed as an instructive process, where the assistant is explicitly learning information
And if we made that learning process far more transparent
Synthetic Persona Pretraining: Alignment From Token Zero – our full paper is finally out. We train up to 3B models and inject synthetic morally-laden reflections into 10% of pretraining documents.
Surprisingly, intervening early really shifts the model's value priorities 🧵
if you buy the view that post-training privileges the assistant persona
and that this privilege is part of what could ground something mind-like
and we adopt this method of instilling a persona into pretraining from token zero
we may get something much more mind-like
cc: @dcshiller
i keep thinking about this paper:
I strongly believe we could get much more stable, aligned models if pretraining were framed as an instructive process, where the assistant is explicitly learning information
And if we made that learning process far more transparent