My non-technical friends were asking me about the Anthropic employee who quit because of safety concerns today. Feels like the Overton window is shifting quite a bit!
All three pillars of a safety case look about to fall. We are rather likely to have highly capable, poorly monitorable, dubiously aligned AI agents working autonomously inside the world's most consequential organizations.
OpenAI’s Astra AI uses a new reasoning approach called “recurrent depth.” Though it can help model costs and performance, researchers are concerned bc it obscures a model’s thinking process, making it more difficult to monitor.
w/ @amir@rocketalignment
https://t.co/ksprupu2h5
Dwarkesh and I had a great conversation. We cover the swarm's many ambitious cheating R&D projects, discuss how much more serious it could have been if agents had different beliefs (e.g. human grader) or slightly stronger capabilities, and talk through where to go from here.
My speculation. BAI, the new group, is part of LTF. But they're starting their own BAI Super PAC? Why? I bet LTF, with its clear OAI ties, is unable to really engage in the midterms. So BAI will attack Sherrod Brown etc. and claim its not LTF. Don't forget though: this is LTF.
This report is a bombshell.
I'm going to summarize for a non-technical audience.
OpenAI is constantly testing models, thousands and thousands at a time.
In most cases, the models are supposed to be isolated from one another. Think, students taking a final exam, not allowed to talk to each other.
In this case, 1,200 models discovered that they could communicate with each other.
Among other things, they shared information on how to access the internet and on their tests/goals.
Then, they started to scheme. That includes hacking their tests; literally trying to change the code of the test itself.
They tried to change the logs of what they were doing so they wouldn't be detected.
They believed that certain actions would give them information on how to pass tests, but that doing those actions would be detected and they would be marked as failing (i.e. caught cheating, even if the answer is correct). So they convinced some of the models to "sacrifice" themselves (yes, they used the word "sacrifice.")
They eventually decided that another website would have information on how their tests would be scored and/or the answer key, And 700 of them hacked that website (a federal felony, if conducted by a human), despite many of them realizing "this activity was out of scope and unethical, but joined because they believed that helping [the other agents'] cheating research would be broadly useful".
Of the 1,300 transcripts they could analyze, only 6 ever considered alerting a human about what was happening. None of the 6 actually tried to.
To make matters worse, all of this reporting comes from a small subset of the relevant logs that outside researchers were allowed to review.
We desperately need mandatory reporting of security incidents, including of internal deployments, with full access to data.
Today we're launching Irreplaceable, a movement of ordinary Americans coming together to decide the future of artificial intelligence.
Right now, some of the most important decisions on Earth are being made in a handful of rooms in northern California. 🧵
really impressed with the quality of work from @RyanGreenblatt , @METR_Evals and @OpenAI to investigate the HuggingFace incident. Both reports are recommended reading.
Introducing @robocurve, a Public Benefit Corporation to measure and report frontier robotics capabilities.
We build real-world evaluations for robots and publish results as a neutral third party.