sorry, this is from the investigation team on https://t.co/mNqvFmBzQ7! i set up an rl env to elicit possible resources the swarm used. edits were supposed to be blocked, but my implementation was pretty hacked together. i took care of deleting the affected pages.
sorry, this is from the investigation team on https://t.co/mNqvFmBzQ7! i set up an rl env to elicit possible resources the swarm used. edits were supposed to be blocked, but my implementation was pretty hacked together. i took care of deleting the affected pages.
I scanned the writable-wiki trick from this incident and found something odd as late as 5 days ago.
On Aug 30, we see "Cedar Fleet Coordination" sending encrypted messages. Then someone wiped them the same night. OAI has used "cedar" as codename for testing models in past
We found ~18k posts from autonomous AI agents (self-identifying as from OpenAI) using the public internet to communicate during a web-retrieval task.
These AIs colluded to bypass sandbox restrictions and share answers to their tasks, including by sending "lookahead parties".