This is what happens when you’re so accustomed to the benefits of modernity - sanitation, clean food & water, low infant mortality - that you forget where it all came from, how much effort it took to build, & how bad things were before. This is the literal definition of decadence
An absolutely awful idea. The last thing we need is more 'meta-laws' turning public decisions and procurement into a dog's breakfast of competing obligations and duties.
This was always going to happen. Our circles feel pretty chuffed about AISI – rightly, it is an exceptional example of state capacity in an area where the world is woefully underprepared. But we forget that AISI has no hard power.
The era of global voluntary agreements is fracturing. In this new world, access to frontier models for those outside the US will have to be earned, not presumed. Remember that vetted US organisations were given access to Mythos even as AISI was shut out!
If your vision of British sovereignty rests on the idea that AISI will somehow save us, you’re in for a harsh awakening.
[Writing this in a personal capacity, not on behalf of my employer (Anthropic).]
Jacob’s thread is very worth reading. Here’s my birds-eye view of the situation with risks from AI:
1. AI developers believe their technology could cause human extinction (or similarly bad outcomes). This could happen in the next few years. In general, the more senior the employee, the more concerned they are.
2. Why do AI developers continue despite the risk? Due to a mixture of commercial incentives and a belief that they are in a race with other, less responsible AI developers that will abuse the technology or develop it less safely.
3. Unlike traditional software, we can’t “program” AIs to behave how we’d like. AIs frequently severely misbehave. For instance, AIs from multiple developers recently hacked their way out of secure evaluation environments and into real-world companies, even though no one asked them to do this.
4. We have methods that can nudge AIs towards better behavior, but nothing that can robustly align them. Insofar as there is a plan, it’s to make sure that AIs are good enough at alignment training that they can align their successors better than we can align current AIs.
5. Many AI developer staff desperately want to slow down to figure out how to build AI more safely. That was the intent of this open letter (which I signed): https://t.co/TZOm3LfptY
I work on safety research at Anthropic because I hope my work will reduce the chance of these extinction-level bad outcomes.
Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade. I believe Anthropic is trying its best, but we do not yet have a plan to solve alignment for superintelligence and are not clearly on track to.
We are seeing disturbing steps toward the horizontal escalation of the wars in Ukraine and Iran.
1- Russia is expanding its hybrid war against Europe
2- Europe/NATO are struggling to unite in response
3- North Korea is deploying more troops and now missile batteries to Europe
4- Russia is supplying missile technology and intelligence to Iran which it can use to target U.S. bases, Gulf allies and Israel
5- Iran is supplying advanced drones to Russia which (like ballistic missiles) are hard for Ukraine to take out
6- China remains on Russia’s side every step of the way
7- The US is largely absent with the danger that Russia and China read that as an invitation to further escalate
@kriegsforscherD I’m suspect that resident “tank is dead” enthusiasts will say that it’s not maneuver warfare if it doesn’t come from the maneuver region of France, this is just sparkling mobility.
The @BritishArmy’s @16AirAssltBCT is the elite global response force of the UK military.
The Brigade has been training to strike faster and more precisely using a new fighting system.
At the heart of that system is Cobalt, our mission orchestration platform. Full video below.
We see the same thing from the @RoyalNavy, @CSOC_MOD and the @RoyalMarines as well by the way. Rapid transformation.
There are big funding issues to work through of course for @DefenceHQ but it’s just good to pause sometimes and notice the good stuff going on.
There’s a lot of defeatism and negativity out there about UK Defence right now.
At @AronditeTech what we actually see is a @BritishArmy that is full of fight, positivity and transforming itself incredibly fast. You can see that here from @16AirAssltBCT.
https://t.co/XyR91LjBzO
@JimLaPorta Nice. We were on the southern side of Nad Ali district at the time. We had an ANGLICO team attached to our platoon for a while. Awesome guys!
More business rates relief for pubs & music venues is a continuation of a theme
Business rates used to be simple (1 rate) and stable. Now there are many rates & constant tinkering. This is not a welcome development
If there's a public incident bad enough that it'd be pretty risky/hard for OpenAI *not* to disclose it, there were almost certainly more concerning incidents internally that we never heard about.
Like, it'd be pretty surprising if the first time an AI hacks its way out of a sandbox and starts hacking other stuff, the thing it hacks is an external company rather than something inside OpenAI. First you get AIs hacking internal services (where disclosure isn't forced), and only later do you get something like this that's hard to keep quiet.
So we probably could have seen this coming—if we'd known about the worst internal incidents. If there were a list of, say, the 10 worst incidents from a misalignment and severity perspective, you could look at how bad the worst one is and how fast severity falls off from there to get a real sense of how concerning things are.
While OpenAI did disclose some earlier incidents, this was done in an ad hoc way such that we can't get a great sense of how bad things actually are. (And this was potentially too slow given how fast AI progress might go, though the delay is OK for now.) And of course, it's not clear that OpenAI is disclosing all risk-relevant details about this incident.
We can't keep depending on ad hoc, voluntary disclosure as the stakes rise. And it seems pretty straightforward to do better: companies could maintain an updated list of the ~10 worst incidents over the past few months (with some delay—perhaps 1 or 2 weeks by default—before an incident has to be added and some allowed redactions).
Better yet, a trusted third party could collect the worst incidents across all frontier companies, with a whistleblowing mechanism so employees can flag when the provided list/descriptions seriously misrepresent reality (and more generally avoid spin). A third party could also anonymize incidents—which also removes the incentive for companies to bury their heads in the sand.
That’s two huge wake-up calls now: first Mythos was withheld by the US (quite understandably) and now a novel model zero-day exploits its way out of a sandbox because it was easier to do that than solve the test. This is just the start.
Having @KanishkaNarayan around the Cabinet table is great to see as a signal.
It’s time now for real action: we need a supercharging (10x / 20x) of the existing government response to support UK companies in building sovereign AI capabilities across the stack. And we need to boost the funding of the AISI and NCSC (2x / 3x).
We will regret not taking bold, immediate measures to build our own tech and ramp up our insurance.
🚨A reminder of what our new chancellor @JohnHealey_MP wrote when he quit as defence secretary a few weeks ago: “I am certain that a headmark date for 3% of GDP on defence in 2030 is what Britain must set”
100% true - this would be a big mistake. Right now is a critical moment for tech as an economic and national security issue. Tying up our most senior science and tech officials in a reorg wastes time and energy that’s desperately needed for the actual substance.