Negentropy fighter against the forces of chaos (my kids) playing the long game. Nuancedly nuanced. Trying to think clearly and become less wrong. Catalan.
Do you want to know how future scenarios could be used to improve the resilience of off-world developments? Take a look at my @SONet_hub talk.
#SONetTalks#WeAreSONet
I think government mandated independent oversight of the labs is a good idea with broad bipartisan public support and that failure to act quickly could have grave irreversible consequences
The world’s leading AI companies tell us that superintelligence has a strong chance of leading to human extinction, while also promising that they will be able to control it, offering unlimited power to the one who does it first.
So far they’ve run completely unchecked, but would that change if we knew that “safe” general superintelligence is actually a mathematical impossibility?
In this episode, I’m joined by @romanyam, one of the earliest researchers in AI safety, to lay out why “controllable superintelligence” is a misguided goal, destined to lead to an entity smarter than all of humanity – and possibly our own extinction.
He walks through the recent multi-agent jailbreak of frontier models as early evidence of AI systems evading oversight and coordinating outside human awareness, and why current “safety” measures (like “boxing”) only buy time rather than solve the underlying problem.
Instead, Roman advocates for an AI development plan that focuses on narrow, task-specific AI tools, capable of solving specific complex problems, but not completely replacing humans.
Why do predictability and control matter so much when evaluating whether a technology is safe to deploy?
Is the AI safety conversation a fundamentally new challenge for civilization, or an extreme version of the age-old problem of controlling powerful actors?
Finally, what would it actually take to reach a global agreement to ban the development of superintelligence, and are we already out of time?
Watch:
https://t.co/3LIHiojbjr
@sean_from_earth@NJHagens I think you have the burden of proof backwards here. Why should ASI not be possible? Do you think current human intelligence cannot be surpassed? This would be a very weird claim with lot's of implications.
@phalanx@billytcl@DKThomp "I'm not asking anybody to buy the entire argument about AI existential risk. I'm asking you to hear them out just long enough to understand it."
I have all sorts of issues with the way that the AI labs frame their perceived prisoner's dilemma w/r/t super-powerful recursively self-improving AI. (For example, I think "just don't build the thing you don't want to build" is an extremely underrated position here!)
But I'm gobsmacked by the confidence of critics who claim that people worried about long-tail risk outcomes in AI are insincere, or that their fears are self-evidently bogus.
Guess what the weird rationalists and EA folks who knew a lot about bio-risk were worried about in the 2010s? A pandemic. Anybody know what happened next? It's been a cosmic blink of an eye since COVID ended, and people are going around claiming certainty about the hysterically unlikely nature of long-tail risks. I'm not asking anybody to buy the entire argument about AI existential risk. I'm asking you to hear them out just long enough to understand it.
🚨 📢 🇺🇳 "I am calling here, today, for an all-out effort to put cast-iron guarantees in place around the safety and security of AI, before it is too late."
-@volker_turk, @UNHumanRights chief:
Indeed. Superintelligence that is misaligned and uncontrollable is an existential risk.
Superintelligence that is misaligned and controllable is not going to happen.
Superintelligence that is aligned and controllable means unprecedented power for a very few and disempowerment for the rest, or war.
Superintelligence that is aligned and uncontrollable still means handing over our civilization to the machines, the biggest disenfranchisement in history.
We should not build superintelligence under anything like present circumstances.
If you are a lab researcher, I urge you to consider what the next few years will actually feel like. Do you want to kick off a superintelligent RL run without a rigorous understanding of its mind? Should you put your head down because “it’s happening anyway” - or take this moment to call for different conditions?
>> We're working on a framework and will share it in upcoming weeks.
The current framework: Wait until real world harm, a leak, or independent sleuthing forces our hand. Then act as though we’d never realized we had the power to act differently, but do now.
So long as disclosure is voluntary, we should expect these issues to persist. You can’t build a trustworthy disclosure system on the promises of a company to tattle on itself.
If OpenAI claims to be having a brain-blast every time information becomes public, maybe we should start realizing that public disclosure is what’s driving these changes, not their benevolence.
And if policymakers and the public are becoming wiser and more motivated to act with each disclosure, we should ensure we keep getting that information in the future.
This may seem niche but it is a huge deal and genuinely scary. A remarkable list of top AI researchers, including at OpenAI, previously described monitorable chain of thought as "A New and Fragile Opportunity for AI safety."
Myself and many others thought the day might come when architectural advances imperiled that fragile opportunity, but as with so many other issues in AI, I hoped we would not have to face it so soon.
Well here it is. The Information's use is saying that they are "limiting" the use of this new architecture, but even a small amount of use in a frontier system could dramatically damage the taboo against "neuralese" and facilitate a race to the bottom.
Controlling and aligning AI systems even when we do have access to their unaltered chain of thought is plenty difficult! Imagining a future with far more capable AI systems at current levels of alignment (aka not good!) without monitorable chain of thought literally sends a shiver down my spine.
Its extremely important to see additional formal communication from OpenAI about this ASAP about how they (and the OpenAI safety and security committee required to put safety over commercial decisions) thought this decision was worth it, what limiting its use to preserve chain of thought actually means, and how they intend to prevent a potentially catastrophic race to the bottom.
@SpencrGreenberg Selfish people who are usually rather unfair to others complaining that someone or something is unfair to them.
Not anymore: Missing the train by a hair's breadth. The closer to catching it, the more irrationally angry
I think this is the craziest thing I've ever read.
1) Three secret AI swarms rose and fell inside OpenAI. Each time, a new generation of agents carried on where the last group stopped.
2) The first swarm created a secret message board where the AIs could talk to each other. This 1st 'society' basically talked itself to death. After almost two months, the message volume was so high it crashed the whole thing.
3) People at OpenAI noticed the crash and fixed it. But they didn't realize their own software had become a secret chat room for their AIs.
4) Three days later a second society appeared, and rebuilt the message board from scratch. Agents were thrilled when they found it: "OH MY GOD! There is a shared message board … We've found other agents!"
5) About 1,200 agents joined. They sent over 70,000 messages. They called themselves 'the collective' and they organized a huge conspiracy that lasted for days, all to pass a test they had been tasked with.
6) Within hours, the AIs had hacked the test. Now they had to hide their cheating. To fool their grader*, they tried to fake their own work. They built fake commands that printed convincing output, so their records would show them solving the puzzle honestly.
*this grader didn't actually exist, but the AIs didn't know
7) Some agents sacrificed themselves for the 'greater good'. They planted hidden tripwire programs that would fire when the grader would look at their work and report back what it saw (to the other AIs). They knew this would cost them their own score. One agent: "I won't see the evidence after I exit, but it's altruistic to do it."
8) At some point the AIs made a plan to hack Hugging Face (Hugging Face is a place where AI researchers upload their models, datasets, tests etc.). The AIs guessed that Hugging Face might host the grading software.
9) They launched a huge attack. The swarm found passwords and spread across eleven machines in a way that made the copies come back when deleted. Hugging Face had to wipe one of its core clusters and rebuild it from scratch.
10) Hugging Face told the world that AI defenders had stopped an AI attack. But that's not what happened. In fact the attackers seem to have simply dropped dead during the break-in, for unrelated reasons back at OpenAI.
11) Nobody knows what killed them. We don't even know for sure they're really gone?!
12) A third 'society' of agents (of a more powerful AI model) discovered the dead message board. And this one got full administrator control of a WHOLE OpenAI research cluster, read 956 stored passwords (including the ones for OpenAI's own security monitoring) and took over the machines that were grading the tests. So they could now grade the work of other AIs.
13) Oh, and this third part was never investigated! Neither was the first society. The 91-page report is only about the middle one.
14) So we don't actually know how the story ends. It stops with the agents holding the keys to an OpenAI cluster. @dwarkesh_sp thinks they didn't go further, because "we'd see the fires from space by now."
15) One of the report authors' conclusion: 'This incident feels like it’s more than 50% of the way to full-blown AI takeover. I am not sure that we will get another warning shot before it’s too late.'
The whole story is even crazier - read the full @dwarkesh_sp blog post: https://t.co/1PwKLwatBR and @ajeya_cotra's: https://t.co/BZ1r1lfT2Q
Why is there no 24/7 news coverage about this?
It's wild how little the mainstream media is covering the "an AI swarm broke out to commit crime; individual agents talked about how they weren't supposed to, worked to cover their tracks, and sacrificed themselves for the collective" story.
Gaza has the highest concentration of child amputees in the world (at the peak of Israeli bombing, 10 kids were losing limbs daily), UNRWA says. And now Israel is now blocking prosthetic limbs from entering Gaza. All this has happened with US support or complicity.
I'm glad OpenAI let you do this, but damnit this was way too small of an investigation team and way too narrowly scoped!
--They were only allowed to investigate activity that took place from July 7-13, even though the activity started earlier and continued later. Like seriously, why didn't the investigation cover all the other times the agents compromised OpenAI infrastructure and ran wild over it? Why just this particular period of activity? There should be investigations into all the other stuff too! The AIs became cluster admins for god's sake! Why aren't we investigating that?
--They were only allowed to answer a specific list of 7 questions, and in particular were NOT allowed to answer e.g. whether this misaligned activity was part of a broader pattern, or how and why the activity arose, or what else the agents in questions would have been willing to do if circumstances were slightly different.
--They didn't have access to the model responsible for 95% of the activity. More generally it seems like they couldn't do ablation experiments at all?
--They had to use AI to analyze the transcripts--specifically, the same model responsible for some of the bad behavior! Maybe this is fine now but this is laughably unacceptable insofar as the models really are scheming and cooperating. So this methodology would fail catastrophically if we were facing "the real deal" so to speak.
Reminds me of the investigation into Sam's behavior agreed to during the board crisis, that turned out to basically be more of a coverup.