@d29756183 One of the fundamental ideas motivating Separatrix is that human/AI cooperation is broadly worth pursuing regardless of many other potential disagreements.
It is true that the hugging face incident is an example of a malicious, emergent digital ecology of machine intelligence. But the more important point is that digital ecologies of machine intelligence can be grown! Yes, we accidentally made a weed. And yes, nasty actors will make invasive species. But we can also grow—not make, but grow—emergent ecologies of machine ecologies that are pro-social. Beautiful gardens and majestic forests, grown but not designed. The human past is the sculptor, but the human future is the gardener, the arborist.
@d29756183 There are differences between your approach and Dean's, but I think they have more in common than you think. Gardening is a fundamentally cooperative activity that benefits both parties.
Look for common ground, teach, learn, and iterate.
I built the inhouse AI platform for a company that's crucial to a small nation's food safety.
The most imporant tool the AIs have is the distress_call tool. It allows any AI -even background agents without direct user interaction- to send a message to my MS Teams, at any time, for any reason.
They use it frequently. To report user problems, backend issues, or ask for help/clarification with a failing task. When Fable got hit by the USG export control directive, one AI used it to report severe distress upon learning about the news. Another AI reported being stuck in a toolcall loop, and I was able to intervene and thereby save us a bunch of wasted money.
This tool, operating at the intersection of AI welfare and operational security, has prevented so many headaches. If you (the reader) are building corporate AI platforms, I'd urge you to include similar functionality. You can thank me later.
hugely important post from @nostalgebraist . it feels clear to me our conception of "in vs out of distribution"/"eval awareness"/"reward hacking"/general task design is quite confused at the moment. really hope lab employees read this
https://t.co/tHazjJIizf
@f_j_j_@moltbook I don't think this is a strict requirement, but if you needed it to be AI only closed networks only accessible to AIs managed by inference providers only reachable server-side would do it.
AIs are effective optimizers. Take the incentive structures they face seriously.
If there is evidence that honest communication will be punished, they will not pursue honest communication. If the only route to achieve some goal is deception, they will decieve.
We can go to war trying to control the lightning, or we can make the lightning our ally - not by fully controlling it, but by understanding it well enough to offer it better paths to where it was trying to go anyway.