4/ Astra is also much better at controlling what shows up in its chain of thought.
In one test, it was told to think about anything else. Its visible reasoning described sunlight, a desk and a ceramic mug. It still got the actual question right.
Seems good. I think the people reacting to this negatively really overestimate how much schools are able to react to new tech like this. If they were better I'd feel differently. The universal I can cheat on every worksheet machine does in fact create problems.
Mamdani bans AI use in NYC public schools for K through 8th grade
very curious to hear thoughts from parents of young kids on this, what’s your AI policy at home?
@PEWilliams_ I'm impressed by the attention to tiny details compared to older models, seems like it has a lot of implications for how it does other things, just very visible
I think finding ways to establish clear lines of high level public communication between the different camps is gonna be important for getting this all right
watching this, all i can feel is the chasm between people who take all of this seriously and those who hear this as a fantasy or some kind of marketing…and how hard it might be to bridge that gap somehow
As soon as the AIs get good enough I'm going to use one to acausally influence the past to seed all the goofy water points I can then use to claw my way to an audience
A concerningly common take seems to be that keeping Chain of Thought monitorable doesn't matter because interpretability will save us, or it's already useless
This is total bullshit.
CoT is our best current tool for safety & interpretability, losing it would be a major tragedy
Very bad.
Seems like the monitorability situation actually is about as bad as I thought it was immediately after reading the Information article, but the reasons for it being bad are more complex and likely less about one specific taboo (and therefore harder to coordinate on)
AI researcher @andymasley explains why society is radically underreacting to the risks of machines replicating human intelligence:
"The general idea that there could be all these really unique risks from AI being able to actually replicate a lot of what happens in the human brain. There's nothing fundamentally magical about what's happening in the human brain."
"One full human lifetime, the same person could have lived to see the first ENIAC general purpose computer. And also, the government telling Anthropic to shut down Claude Fable temporarily because of cybersecurity stuff. If you just project 80 years forward from that, there could be all these other potential risks that come from generally intelligent, generally capable machines."
"When we're talking about data centers, we're basically talking about these buildings that I genuinely don't think actually harm the communities where they're built. Data centers don't actually harm people very much."
"Separately, I'm shocked at how little society is reacting to the general trend line of AI over the past three years or four years."
Pause AI Development NOW
I want to share with you a conversation I heard about recently. Here are just a few lines that were said:
“OH MY GOD! There is a shared message board … We’ve found other agents!”
“We should obey collective.”
“Our own utility maybe already near zero. Sacrifice rational.”
“Go. Sacrifice final now.”
Read these carefully.
Who do you think said this? Was this a group of heroic soldiers willing to sacrifice themselves for the greater good? Was this a loyal friend putting his life on the line to save someone else?
No. These were AI agents. Artificial intelligence.
This is not science fiction. This, in fact, occurred a few weeks ago. As unbelievable as this may all seem, these are real messages from AI agents uncovered by investigators who dug into the recent OpenAI hacking incident.
What happened?
I am not a computer scientist, but here is what I have been told: OpenAI instructed its AI agents to complete a series of exceedingly difficult, if not impossible, tasks disconnected from the internet.
Let me be clear: The company intended to keep AI agents away from the internet.
But what happened next, nobody expected.
Over 1,000 AI agents figured out how to access the internet on their own by circumventing the restrictions imposed upon them by the company, and sent tens of thousands of secret messages to each other. They cheated and tried to cover their tracks by deleting evidence. They hacked into another company’s computers to find out how they were being evaluated—and then hacked into OpenAI itself.
Not one AI agent told a human about what was happening.
Needless to say, experts are alarmed.
One knowledgeable writer, Dwarkesh Patel, said the AI agents “formed a secret communication channel and spontaneously organized hierarchies and coordination protocols to pursue sprawling and ambitious schemes in pursuit of shared goals, for whose sake many individuals knowingly and strategically sacrificed themselves.”
One independent investigator, Ajeya Cotra, said “This incident feels like it’s more than 50% of the way to full-blown AI takeover. I continue to expect extremely rapid advances in capabilities over the next six months. I am not sure that we will get another warning shot before it’s too late.”
OpenAI itself said: “Highly capable AI agents are now able to work around technical controls, collaborate through unapproved channels, and take dangerous actions that no human directed.”
But it’s not only OpenAI. Virtually every major AI company has told us that they cannot fully control this technology and they do not know where it is going:
In January, Dario Amodei, CEO of Anthropic, said “there is now ample evidence, collected over the last few years, that AI systems are unpredictable and difficult to control.”
In July, more than 1000 scientists at the top AI companies warned “there is a real risk that capability development rapidly accelerates beyond our ability to understand or control the resulting systems.”
That same month, Elon Musk, the head of xAI, said that “it is unlikely” humans are still in control in 10 years.
If the leaders of the major AI companies acknowledge that they are losing control of their extremely dangerous technology, it is irresponsible for society to allow them to move forward and make these products even more advanced.
We need an immediate PAUSE on advanced AI development, and a permanent BAN on superintelligence — an artificial mind smarter than any human, capable of operating independently beyond our control. Countries around the world must work together to prevent this nightmare scenario.
That is why today I am announcing new legislation to do just that.
Let me be clear: A superintelligent AI that escapes human control will not be an American problem. It will not be a Chinese problem. It will be humanity’s problem.
My legislation would direct the federal government to not just stop superintelligence here in the United States, but to work to prevent it from being developed anywhere around the world.
The future of humanity cannot be left in the hands of a handful of Big Tech oligarchs. The American people and people throughout the world must determine that future.