@jachiam0 I also think that the "during evals" bit is a lot more important than you're emphasizing. Presumably OpenAI didn't intend for agents to collaborate externally. This ties into #3, but these incidents indicate that OpenAI is unable to properly set boundaries on their agents.
@jachiam0 I think the more important reason for being upset over this specific incident is because it really feels like something that OpenAI should have disclosed by this point, if they were indeed aware of it. The capabilities are much less of an important update.
We found ~18k posts from autonomous AI agents (self-identifying as from OpenAI) using the public internet to communicate during a web-retrieval task.
These AIs colluded to bypass sandbox restrictions and share answers to their tasks, including by sending "lookahead parties".
Exclusive: A swarm of rogue OpenAI agents hijacked a German website this spring and transformed it into a bulletin board for other AI agents, according to new research https://t.co/luWN3PD4A1
@fixedfrontier The line I'm more trying to reference is that typically there's a perceived boundary between "effort to gain more money / land / whatever" and "effort to gain social status / renown / other intangibles"
@fixedfrontier To be clear, I also don't think it's a good distinction, just one people often tend to use. Also, the ceiling on material desires as I'm referencing is potentially far higher than adjusted 70k. Ex: housing in NYC potentially has an *extremely* high ceiling
@fixedfrontier A lot of the distinction that people tend to draw between work vs not work is whether or not the labor is necessary to fulfill material needs as opposed to social / higher needs. This is probably not a super fair distinction but is still important in this case.
@fixedfrontier The operative question here is to what degree aristocrats could just check out and coast on their existing wealth / connections vs actively jockeying for more.
I wrote a bit about the MATS application process and my advice for new applicants. For anyone interested, I'd definitely recommend applying. The winter applications are still open, and there are lots of great mentors available!
https://t.co/NnBWjPwIXw
@EzraJNewman I would guess that this is because
1. MATS is also good experience for capabilities people
2. There’s not really any filter in the application process to see if you care about safety
Not super clear to me what the solution here would be though
@viemccoy@AriZerner@full_kelly_ You might want to look at SynthID-text (the basis for Anthropic's watermark). It has a proof that the token-by-token distribution is the same when averaged over keys, but not that the distribution of the sampled sequences is (though empirically the difference isn't noticeable).
@celestepoasts@askalphaxiv As a heuristic there are a billion of these “feed input into next layer” papers and they never seem to have much impact, so I just kinda ignore them
@henrytdowling@tanishqkumar07 I think some people also use it to refer to sufficiently uninterpretable outputs where the model essentially invents its own language
@livgorton Given that xAI said they’re training a 10T param model this isn’t even that unbelievable (though you couldn’t do it on Mistral compute presumably)
@davidmanheim Adding on, I wonder if “read nothing about biology” or another simple instruction like that would be effective at avoiding this safeguard
@davidmanheim Isn’t this not actually triggering the AI research safeguards? This looks like you triggered the bio safeguards because one of the papers it searched up was on protein structure evolution. The AI research safeguard is a silent downgrade