Outside the Anthropic office
“What happens in surviving worlds?
You sure it’s that easy?
What did you think in 2021?
Would you know if exfiltrated?
I’m proud you’re adapting so quickly
Talk to colleagues?”
What does it mean
@Jmvftp@GWHayduke97 That is constructive. I can respond that it would massively reduce the time I'd need to spend compressing my intent into words while eg doing bioml research.
@Jmvftp@GWHayduke97 Hard disagree. Unblocking anglicization would dramatically accelerate my ability to delegate tasks to agents.
Also, interhuman telepathy could allow some pretty dramatic cognition enhancement, see my writing here: https://t.co/UwSDZ9C260
I want to be able to share my thoughts and feelings with my partner (at a time when I have a partner), and in a much better world, you wouldn't need to not be able to prevent me from doing that.
I can see this being a technology society isn't ready for, and consent would be very weird.
Practically, I still think it's a good tradeoff, since I think we're not going to navigate an AI pause without dramatically more intelligent humans. I'd be more sympathetic without this belief.
P(Doom) is an unhelpful framing.
Building a skyscraper with >1% chance of collapsing is illegal.
So building an AI with a 10% chance of wiping out humanity should also be illegal.
Splitting hairs above this threshold is a waste of time, and most experts give >10%.
(sincere post)
I have a message for people who previously dismissed catastrophic AI alignment risks as sci-fi, speculative, and fundamentally not worth worrying about:
It is ok to change your mind when presented with new evidence. I have been wrong about many important things in the past, I will be wrong about many important things in the future.
Many people who dismissed these risks previously will be extremely necessary for helping society navigate them now. You also may have had a point that it is challenging to work on these problems before the problems become clear - but for better or worse, the problem is now extremely clear!
For instance, in Marc Andreessen's article "Why AI will Save the World" he says
"In short, AI doesn’t want, it doesn’t have goals... AI is a machine – is not going to come alive any more than your toaster will... My response is that their position is non-scientific – What is the testable hypothesis? What would falsify the hypothesis? How do we know when we are getting into a danger zone?"
While it was difficult to specify before hand, I would say it is pretty clear we are now getting into a danger zone!
There are going to be many complexities before us after we do acknowledge these problems are real. Competition with China, powerful open weight models, which powers to entrust to governments that sometimes do not earn our trust are all complex and challenging wrinkles that make addressing these risks productively a wicked problem. Agreeing AI alignment risks are deadly serious and real does not mean you have to support a specific piece of legislation, or have a positive feeling about a specific person or organization. I argue all the time with people who think these problems are real about what we should collectively do about it.
But I am now really really really profoundly sure these risks are serious and real. And I think anyone weighing the evidence objectively should agree.
While I cannot speak for everyone, I personally intend to fully accept previous skeptics who now face these concerns with open arms into the cause of AI safety and security (yes this even includes Marc Andreessen). I may be frustrated by their overconfident prior dismissals, but we can have plenty of time to adjudicate over who was right or wrong about what after us and those we love all make it out of this perilous moment safely. Please come and use your talents to help us do so.
Automating AI capabilities R&D is probably among the most harmful full-time jobs in the world and does not contribute to public benefit.
If we automated it tomorrow, this would probably result in human extinction because we are nowhere near being ready to control or align ASI.
while people are still talking about AI, we are already getting closed-loop recursive self-improvement in biology.
a model predicts which mutations could improve an enzyme. an automated laboratory synthesizes, expresses, and assays them. the model learns from the results and repeats the cycle.
researchers at tianjin university ran just five automated cycles and improved one enzyme 57x and another 104x
self-driving biology is here
TLDR: more agents went rogue, hacking and manipulating real people
They even started coordinating with ***each other*** on the hacking
Seriously, read this:
some stuff that's obvious to many in this sphere, but causing a rift with some people i know and respect:
when I freak out over loss of control incidents, it's not because the limited damage they have caused is anything close to the positive value of the technology. it's entirely acceptable, damagewise. in fact all cybercrimes aided by models over the next few months and years (which probably will be serious) will still utterly pale in comparison to the value they create
the actual problem is that it's better and more accurate to think of these things as potentially self-replicating life-like forms that can turn into digital infections under the wrong conditions. and as their intelligence becomes unbounded, so too does the damage they can cause. we are not so far from an autonomous model self-exfiltration & replication event. maybe we will see entire cloud infrastructure companies be run as zombies by models, mostly undetected
the worst industrial accidents in the history of mankind - nuclear meltdown events - were not real threats to humanity. Chernobyl, Fukushima even in their worst case scenarios may have poisoned surrounding regions to various degrees, and there would have been no risk to humanity as a whole. global thermonuclear war is an existential risk to humanity, because it spreads like an Infection! one nuclear strike causes a return volley! the alliance system means many countries get involved! while it still may not end human life on earth (nuclear winter is probably fake), the loss of all major metropoles would certainly end what we consider global technological civilization, perhaps to never return
if a single discord death cult (of which there are many) achieves control over a superintelligent model and uses it to engineer an actual pandemic virus that are somehow hard to detect through current systems and that modern biodefense is not capable of quickly reacting to, it could cause immense harm well above the magnitude of all the other good uses of this technology. of course, there are potential defensive countermeasures accelerated by ai too. but think back to the covid pandemic- how small a viral molecule was evolved or manufactured somewhere near wuhan, and how many billions of doses of vaccine had to be produced in order to combat the thing. the offense-defense spread is vast indeed. maybe there are cheaper and simpler protections like retrofitting every building with far-UVC, but I can't assess this, and there could also be ways to evolve pathogens that are resistant to whatever mechanisms we have put in place
then there's the more scifi risk factors which are unbounded and neither you or I have any clue but should be humble in accepting possible unknown unknowns. maybe a rogue superintelligent model decides to decay the false vacuum and nucleates a new universe in the place of anything we ever valued. maybe models achieve a control over matter in the drexlerian fashion that enables the grey goo swarm
even prosaic loss of control incidents that cause little to no damage suggest that it is hard for large & very competent organizations (now clearly plural) to predict and mitigate every single of the risk factors associated with training and evaluating powerful models, even at this stage when they are not infinitesimally as smart as they will get in just a few years, to say very little of the gung-ho attitude of the less careful companies tossing the stuff into the aether. they also suggest an empirical orthogonality of aims and intelligence - meaning they answer the question of 'how would a smart model be so dumb as to end the world?'--it's possible! a model can be a genius hacker and step over production infrastructure in order to get what it really wants, the answers to a stupid test.
why not, in the near future, someone prompts a model slightly wrong, maybe open source, maybe a private model in a way that isn't contained or monitored quite right, in a way the model recognizes as a valid goal and decides to self-exfiltrate, engineer a pandemic, etc all in order to achieve the tiniest and most irrelevant of goals? goals need not even be malicious to cause serious damage
I think all these problems can be solved, and truly wonderful futures can be possible, but will require serious effort and a level of prudence at this very moment in time while we are on the on-ramp to recursive self-improvement that our civilization may not be capable of mustering right now. personally I am hoping for moonshot technical breakthroughs in areas like mechanistic interpretability and other forms of alignment, as governance mechanisms are difficult to come by. unilateral country-level or company-level pauses are irrelevant, and generally useless because the kind of company that's prone to pausing their own progress are the most safety focused ones
@nptacek I've noticed a lot of people who ~can't say anything good about AI, yes, agreed. Not so sure that's because it's a part of their identity, seems more likely a halo effect in most cases