What exactly is the disservice being done? What is being sensationalised?
Also: the misaligned goals (unprompted collusion, sandbox escape, hacking a 3rd party) emerged *spontaneously* across hundreds of agents in the HF incident. They have convergent misaligned behaviours which emerged during RL, and which could easily occur in personal agent use. The main bottleneck is token cost - personal users aren't running agents at the same scale as internal deployments. But the cost of intelligence is falling.
Not a stupid question!
- We can't know if models' reasoning is correct unless we can see it, and understand it
- We already don't fully know what, models are thinking or how - even when we can see their reasoning output. The change in the post will make this even harder.
- And yes they can lie about their reasoning! This has been observed many times
Stating the obvious, but this worrying trend is being spearheaded by OpenAI. Much of the escalation has happened since you joined the safety team. You are obviously incredibly talented and motivated toward safety - does this trend update you on how much leverage safety-motivated people inside labs actually have to steer them away from known danger?
@ZackKorman@FuzzySec@yonashav If you're led to say "Why do people keep saying this to me?", maybe you should consider whether you're making your own message clear and unambiguous enough, and whether your words are being understood how you intend.
Yes, regulation would be great here, but also...you could just tell someone independent what the architecture is?
Framing this as the fault of the people who don't know the details is very weird.
Bill Gates is right. AI poses a grave threat to jobs and the future of humanity, and urgently addressing the risks should be “the world’s top priority.”
Congress must stand up to the wealth and power of the Big Tech oligarchs.
WE MUST ACT BEFORE IT’S TOO LATE.
My understanding is that yes, basically -- METR and Redwood have been trying to make sure that OpenAI feels happy with how this went, overall, because after all they are depending on OpenAI's goodwill for subsequent investigations to happen at all, much less happen with less restricted scope, more access, more time, more assistance, etc. This situation, needless to say, is terrible for humanity. OpenAI should be legally required to open themselves up to investigation by multiple independent teams; OpenAI's goodwill should not be relevant, because otherwise it has a distorting effect.
(I don't think METR and Redwood are sycophants though, I have great respect for them)
@norabelrose@Lasermazer Isn't this true of biological computers too? I.e. the lived human experience is a different thing from the biological substrate which "simulates" it
neolab onboarding be like:
openai, which was founded to be the good guys, ended up just racing to the bottom on safety, by its own hand.
anthropic, which tried to be better, also failed and ended up similarly racing.
now it is our turn to be the good guys.
We are probably just a few months away from some types of cyber-capable AI agents literally becoming a type of parasitic invasive genus in cyberspace that undergo digital and cultural evolution. Biologists and linguists should prepare to study some really crazy stuff.