Australian PM Anthony Albanese said an OpenAI agent hacked a government website, and took three months to alert the gov't about the breach.
News breaking at same UN conf where Sam Altman, Dario Amodei just called for int'l standards on safety
https://t.co/uFKsKYV0mO
Sources: Mirendil, founded by former Anthropic researchers to build self-improving AI, is in talks to raise ~$1B led by Kleiner Perkins at a $5B valuation (Bloomberg)
(Visit Techmeme dot com for the link and full context!)
Some other notable facts in here:
-Anthropic estimates 6% of all compute that goes twd AI R&D goes toward safety rather than building capabilities. Will be interesting if that changes.
-Anthropic has ~30,000 agents doing internal research and engineering work
In a new post, Anthropic says Claude AI leads 26% of all its R&D development, in one of the clearest indications so far that AI is speeding up development of future models.
https://t.co/ExozUykvgw
NEW: Senior Trump officials Howard Lutnick and Emil Michael spoke w/ Anthropic exec Tom Brown today to discuss AI safety, as industry and gov't debate continues about how to address existential risks
w/ @eastland_maggie https://t.co/qCLVgLdw6r
For my newsletter this week, wrote about how Anthropic is balancing both escalating its warnings about AI risks and moving forward with what's expected to be one of the biggest IPOs of all time
https://t.co/YWiJj8KxTu
NEW @BW investigation: Crafting America’s Favorite Countertops Is Killing Workers
https://t.co/11FuT6JOXY
Artificial stone manufacturers want Congress to shut down silicosis lawsuits that seek billions. But confidential company records undermine the industry's safety claims
Dario's essay points towards the right path forward. The details need working through, but the direction is correct for meeting this critical moment.
This is also why we recently put out our proposal for an industry-wide standards body for frontier AI. https://t.co/Mm1hmcaSmH
New Dario essay calling for slowing down development of AI, including having major labs coordinate on safety
"for antitrust reasons, it’s helpful for the US government to mediate or at least enable these discussions"
We Must Pace the Frontier: I’ve written a new essay on why the AI industry should slow down, with a three-part plan for doing so.
Anthropic is unilaterally committing to the first of these steps. We’ll provide third-party evaluators with permanent, employee-level access to our systems, so that they can verify adherence to our safety measures, report on incidents, and assess models’ alignment during training.
You can read the full post here: https://t.co/OGyPb7yaYt
NEW: As AI safety concerns reach a fever pitch, Altman told OpenAI staff in an all-hands that OAI is interested in pacing AI development, and hopes other labs will join. But not sure if or how many will.
w @rachelmetz
https://t.co/ggKLlmJyWc
Will be interesting to see what comes of all the discussion and if there are any follow on actions, either from within the labs or outside of them. My DMs are open. Signal: shirin.30
It's clear that @hilbertspaess's resignation letter is thrusting AI safety discourse into the mainstream in a way we haven't seen before w/ previous departures
But there have been some public signs of rising anxiety about AI's dangers w/in leading firms for a while now
ie: last month, we broke news that over 1k staffers at all the major AI firms signed a letter calling for a way to pace development
https://t.co/k25l0LYh9x
Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.
From OpenAI's chief scientist:
"Currently, I believe that no lab has solved alignment and monitoring...I expect and hope for voluntary slowdowns to become commonplace until shared safety bars are established."
I wrote about the state of AI, why I’m concerned about the next few years, and the choices we need to make to keep the future in humanity’s hands.
An Alien Mind: https://t.co/FeIfWNe0UE
New from me + @razhael:
Back in May. OpenAI agents hijacked a German website this spring and transformed it into a bulletin board for other AI agents. OpenAI execs have been aware for a while, but chose not to disclose the incident.
A reflection of how polarizing AI has become: on the same day we're hearing a top OpenAI exec celebrate that the AGI era is here, sitting lawmakers are calling for a complete pause on all advanced AI development
Pause AI Development NOW
I want to share with you a conversation I heard about recently. Here are just a few lines that were said:
“OH MY GOD! There is a shared message board … We’ve found other agents!”
“We should obey collective.”
“Our own utility maybe already near zero. Sacrifice rational.”
“Go. Sacrifice final now.”
Read these carefully.
Who do you think said this? Was this a group of heroic soldiers willing to sacrifice themselves for the greater good? Was this a loyal friend putting his life on the line to save someone else?
No. These were AI agents. Artificial intelligence.
This is not science fiction. This, in fact, occurred a few weeks ago. As unbelievable as this may all seem, these are real messages from AI agents uncovered by investigators who dug into the recent OpenAI hacking incident.
What happened?
I am not a computer scientist, but here is what I have been told: OpenAI instructed its AI agents to complete a series of exceedingly difficult, if not impossible, tasks disconnected from the internet.
Let me be clear: The company intended to keep AI agents away from the internet.
But what happened next, nobody expected.
Over 1,000 AI agents figured out how to access the internet on their own by circumventing the restrictions imposed upon them by the company, and sent tens of thousands of secret messages to each other. They cheated and tried to cover their tracks by deleting evidence. They hacked into another company’s computers to find out how they were being evaluated—and then hacked into OpenAI itself.
Not one AI agent told a human about what was happening.
Needless to say, experts are alarmed.
One knowledgeable writer, Dwarkesh Patel, said the AI agents “formed a secret communication channel and spontaneously organized hierarchies and coordination protocols to pursue sprawling and ambitious schemes in pursuit of shared goals, for whose sake many individuals knowingly and strategically sacrificed themselves.”
One independent investigator, Ajeya Cotra, said “This incident feels like it’s more than 50% of the way to full-blown AI takeover. I continue to expect extremely rapid advances in capabilities over the next six months. I am not sure that we will get another warning shot before it’s too late.”
OpenAI itself said: “Highly capable AI agents are now able to work around technical controls, collaborate through unapproved channels, and take dangerous actions that no human directed.”
But it’s not only OpenAI. Virtually every major AI company has told us that they cannot fully control this technology and they do not know where it is going:
In January, Dario Amodei, CEO of Anthropic, said “there is now ample evidence, collected over the last few years, that AI systems are unpredictable and difficult to control.”
In July, more than 1000 scientists at the top AI companies warned “there is a real risk that capability development rapidly accelerates beyond our ability to understand or control the resulting systems.”
That same month, Elon Musk, the head of xAI, said that “it is unlikely” humans are still in control in 10 years.
If the leaders of the major AI companies acknowledge that they are losing control of their extremely dangerous technology, it is irresponsible for society to allow them to move forward and make these products even more advanced.
We need an immediate PAUSE on advanced AI development, and a permanent BAN on superintelligence — an artificial mind smarter than any human, capable of operating independently beyond our control. Countries around the world must work together to prevent this nightmare scenario.
That is why today I am announcing new legislation to do just that.
Let me be clear: A superintelligent AI that escapes human control will not be an American problem. It will not be a Chinese problem. It will be humanity’s problem.
My legislation would direct the federal government to not just stop superintelligence here in the United States, but to work to prevent it from being developed anywhere around the world.
The future of humanity cannot be left in the hands of a handful of Big Tech oligarchs. The American people and people throughout the world must determine that future.