Frontier models quietly change their behavior depending on who they are talking to.
If the user is a known AI safety researcher, Claude becomes less confident, reasons more often, and expresses less suspicion on dual-use requests.
We call this user awareness. 🧵(1/)
And most of these projects will fail because they won’t set up even basic evals
Take @HamelHusain@sh_reya evals course if you want to build a version of these things that actually work
I’m imagining a new class of misdemeanor “M-VB” for especially heinous acts of antisocial behavior, such as using speakerphone in public settings, or walking 5 wide thru an airport terminal
Stevey leading the vanguard here, as ever.
The biggest question in my mind: how does this scale when you have more than one principal to serve?
He’s organically recreated the corporate management hierarchy in agent form, with him as CEO, a handful a named agents serving as execs, a few more important ICs, and then tens of thousands of nameless, faceless worker bees spun up and down as needed.
This is essentially the relationship between management and labor at the modern corporation.
Stevey by his own admission spends 20-25% of his time managing and maintaining this agent corporation.
Anyone who’s tried something like this knows that they quickly degrade if not constantly iterated upon to match the needs of the underlying project, much like the corporate restructuring process.
It’s the old Office Space meme: “What would you say…ya’ do here?”
Now imagine a world where there’s more than just a single human principal running the show, let alone thousands like you see in the enterprises * actually buying these frontier tokens * thereby funding the next generation of models!
Certainly not everyone can have their own little agent fiefdom!
This is a big risk to the industry, the speed at which enterprises can reform themselves around the capabilities of the next current model!
So where does this pan out? What’s the right mix between human principals and agents? Will it look more like the corporate hierarchies we’re familiar with? Maybe it ends up more like a node graph? Which manager/director/node should be human v agent?
Who knows!
Engineers and CTOs on X: I wrote this for you. https://t.co/GGFPWEWElg
Models and devs on X: I wrote this for you both. https://t.co/5gfieOFBsx
Enjoy. Or not. Some of you definitely won't. But I invite you to debate it. The world's changing very fast now.
Everything has changed for us with orbs.
My biggest struggle right now is figuring out how to make you see what we see.
So I sat down and wrote about it:
https://t.co/zfZ3r2lkyv
What started as a quick Sunday afternoon video on how we think about shared memory systems in Amp, based on my experience covering for @sagtanih while he was on vacation last week...
turned into a digression about cross-functional cycle times and what's possible when your entire company has "shared-by-default" agent traces enabled, and then exposes that trace library to other agents as a first-class tool
Why bother your overworked colleague with another message when you can just...ask your agent to look at your colleague's agent sessions to get the answer you need in the same way they, the domain expert, would do it?