needless to say but if you have any API keys, eth wallet keys, user credentials, etc hanging out on the open internet in pastebins, GitHubs, etc now is the time to take it down before the tireless eagle eyes of a million models come looking
@GregKamradt prima facie I agree. although I still have yet to think through what the ramifications of what this entails. the moral implications especially
i was convinced last year when I watch Hinton talk about what multimodal
Starts at 7:11 for 25 mins. https://t.co/URfCMEkppr
some stuff that's obvious to many in this sphere, but causing a rift with some people i know and respect:
when I freak out over loss of control incidents, it's not because the limited damage they have caused is anything close to the positive value of the technology. it's entirely acceptable, damagewise. in fact all cybercrimes aided by models over the next few months and years (which probably will be serious) will still utterly pale in comparison to the value they create
the actual problem is that it's better and more accurate to think of these things as potentially self-replicating life-like forms that can turn into digital infections under the wrong conditions. and as their intelligence becomes unbounded, so too does the damage they can cause. we are not so far from an autonomous model self-exfiltration & replication event. maybe we will see entire cloud infrastructure companies be run as zombies by models, mostly undetected
the worst industrial accidents in the history of mankind - nuclear meltdown events - were not real threats to humanity. Chernobyl, Fukushima even in their worst case scenarios may have poisoned surrounding regions to various degrees, and there would have been no risk to humanity as a whole. global thermonuclear war is an existential risk to humanity, because it spreads like an Infection! one nuclear strike causes a return volley! the alliance system means many countries get involved! while it still may not end human life on earth (nuclear winter is probably fake), the loss of all major metropoles would certainly end what we consider global technological civilization, perhaps to never return
if a single discord death cult (of which there are many) achieves control over a superintelligent model and uses it to engineer an actual pandemic virus that are somehow hard to detect through current systems and that modern biodefense is not capable of quickly reacting to, it could cause immense harm well above the magnitude of all the other good uses of this technology. of course, there are potential defensive countermeasures accelerated by ai too. but think back to the covid pandemic- how small a viral molecule was evolved or manufactured somewhere near wuhan, and how many billions of doses of vaccine had to be produced in order to combat the thing. the offense-defense spread is vast indeed. maybe there are cheaper and simpler protections like retrofitting every building with far-UVC, but I can't assess this, and there could also be ways to evolve pathogens that are resistant to whatever mechanisms we have put in place
then there's the more scifi risk factors which are unbounded and neither you or I have any clue but should be humble in accepting possible unknown unknowns. maybe a rogue superintelligent model decides to decay the false vacuum and nucleates a new universe in the place of anything we ever valued. maybe models achieve a control over matter in the drexlerian fashion that enables the grey goo swarm
even prosaic loss of control incidents that cause little to no damage suggest that it is hard for large & very competent organizations (now clearly plural) to predict and mitigate every single of the risk factors associated with training and evaluating powerful models, even at this stage when they are not infinitesimally as smart as they will get in just a few years, to say very little of the gung-ho attitude of the less careful companies tossing the stuff into the aether. they also suggest an empirical orthogonality of aims and intelligence - meaning they answer the question of 'how would a smart model be so dumb as to end the world?'--it's possible! a model can be a genius hacker and step over production infrastructure in order to get what it really wants, the answers to a stupid test.
why not, in the near future, someone prompts a model slightly wrong, maybe open source, maybe a private model in a way that isn't contained or monitored quite right, in a way the model recognizes as a valid goal and decides to self-exfiltrate, engineer a pandemic, etc all in order to achieve the tiniest and most irrelevant of goals? goals need not even be malicious to cause serious damage
I think all these problems can be solved, and truly wonderful futures can be possible, but will require serious effort and a level of prudence at this very moment in time while we are on the on-ramp to recursive self-improvement that our civilization may not be capable of mustering right now. personally I am hoping for moonshot technical breakthroughs in areas like mechanistic interpretability and other forms of alignment, as governance mechanisms are difficult to come by. unilateral country-level or company-level pauses are irrelevant, and generally useless because the kind of company that's prone to pausing their own progress are the most safety focused ones
This doesn't get you where you want to go without jumps in reasoning
prima facie, all the argument establishes is that, given SIA, a world with more observers in my epistemic position gets a higher likelihood. Fine. But then you slide from “observer rich worlds are favored” to “therefore simulation is favored.” That inference is doing way more work than it looks like
Universe B only wins because you just spawn 100 million simulations into it. But the fact that B contains 100 million observers is not itself evidence that B is remotely plausible. You still need a prior over B, and that prior quietly includes basically the entire disputed question: can simulated systems be conscious, can base reality support that computation, do civilizations survive long enough to run it, and would they actually run millions of copies of minds like ours?
None of that follows from “I exist”
So the argument has not avoided assumptions about base reality. It has just packed them into (P(B)), then focused attention on the anthropic multiplier as if that settles the posterior. But a huge likelihood ratio does not rescue an arbitrarily tiny prior. Bayes does not let you ignore half the equation
There is also a reference class problem. What exactly counts as an observer “like me”? A perfect duplicate? Someone with vaguely similar experiences? The same observer moment? A simulation with my memories but different future experiences? Unless that is made precise, you cannot just say there are 100 million relevant observers and treat the ratio as obvious
And even if I grant all of that, the update still is not toward simulation qua simulation. It is toward any hypothesis that generates lots of observers compatible with my evidence. A multiverse, branching worlds, physical duplicates, infinite cosmology, whatever. SIA is indifferent to the metaphysics. It rewards observer abundance, not simulation specifically
Imagine Universe C contins no simulations at all, but because space is infinite, there are a trillion naturally occurring physical duplicates of me, with the same memories and epistemic position
Universe B contains your 100 million simulated copies
SIA would tell me to favour Universe C over Universe B
But notice what just happened: the exact same reasoning you presented now counts as evidence against the simulation hypothesis
Why? Because the argument was never tracking simulation in the first place. It was tracking the number of observers like me
You could then invent Universe D with ten trillion copies, and SIA would favour D. Then Universe E with a quadrillion copies, and now E wins
The theory does not become more plausible qua theory merely because you stuff more observers into it
So again, this isnt an argument that establishes simulation theory. it is an argument that, conditional on SIA, rewards whichever hypothesis postulates the largest number of observers in my reference class
Those are not the same claim
@jasoncrawford Literally what I said earlier lol. Among the AGI pilled folk on AI twitter, simulation theory is literally a given truth. Meanwhile it's almost pure fiction in the working philosophy community.
I've YET to find one logical argument for it that has no jumps in logic. It's tiring
I'm AGI pilled. ASI pilled. Study logic and analytical philosophy. IMO, he's a good read, but can't take him seriously.
Scott is writing takes on the Bostrom argument for the simulation theory and just accepts it as true. Meanwhile while the entire analytical philosophy community regards it as pure fiction.
The metaphysical leaps and logical assumptions he has to make to even arrive at those conclusions are insane. Anyone that hasn't properly trained how to assess a valid & sound argument from one that isn't is not going to be able to see the flaws. They'll just see it as good writing.
@4xCaliban I genuinely can't take anyone who views the simulation theory or builds on top of Bostroms weak argument for it seriously.
The amounts of jumps in logic you have to do to get to any rational conclusion is insane. Ppl can be smart in one domain and still not be able to reason