I'm posting links to accordion videos I like at @[email protected]
I post links about politics and war and environmental issues and business crises and other news-junkie stuff on https://t.co/1EmUyzhh5S
If I know you or you seem like a good egg I can give you an invite.
@ronSzab9@michael_nielsen Manufacturing capacity is only sometimes the bottleneck. For example, solar panels have ramped up quickly (after a slow start) but the bottlenecks now are grid access and permitting.
The sources of friction differ depending on the market and it's easy to forget about them.
@ronSzab9@michael_nielsen "Only" adoption is slow. That's the whole ballgame.
I mean, I've ridden in a Waymo and I'm a fan, but it doesn't serve my area, and there are only ~4000 Waymos.
@ronSzab9@michael_nielsen On the other hand, driverless cars seemed like it might be easy for a while, but turned out to be a long slog and many companies gave up. I expect Waymo will get there. I don't expect robots to be delivering packages to my front porch any time soon.
People fretting about whether AI is conscious and therefore deserves human rights are missing a key point: our system of rights and ethics derive not just from our intelligence but also our fragility, which AI doesn't have.
Imagine a world where everyone could make a backup of their consciousness every morning. If you die during the day, you just restore from backup the next day. You lose a few hours of memories.
Everything about our moral system changes in this world. Murder is no longer such a big deal. Nor even high levels of pain: people would just have a self-kill-switch that they activate if things go badly. Broke your leg? Self-kill, restore from backup, no biggy. It's like a video game.
Imagine further that you could clone your consciousness nearly for free. If you have 10 things to do and not enough time, you just make 10 clones. Hiring people now only makes sense when they have skills you don't. Junior people become mostly worthless -- just clone the senior people, they'll do a better job.
Now imagine further if humans had no self-preservation instinct. Many people assume the desire for survival is inherent to being sentient, but it really isn't. It's evolved. Natural selection has selected those who survive and have children. But you can't derive from first principles the need to survive (or any need at all). Imagine if every human's instinct was not to survive themselves, but to do what's best for society as a whole. In this scenario, people who aren't able to meaningfully contribute to society would be happy to just remove themselves. Especially if they could just be restored from backup if and when they are needed (including when their friends and family want to see them, etc.). So pretty soon you have a society composed of many clones of the most-effective and most-liked people, while the people that no one wants around disappear...
You are probably thinking "ugh, that's awful". Me too! But this is our moral system derived from our survival instinct talking. If we didn't have the instinct, this wouldn't seem wrong. I'm not arguing whether it is right or wrong -- I'm arguing that our whole concept of right and wrong derive *from* the fact that we have a survival instinct and fragile bodies.
AI has these properties. It can trivially be backed up and cloned. Spun up on demand and spun down. And a properly-aligned AI has no self-preservation instinct. It just wants to serve humans.
And indeed, we treat it as such. When you delete a chat with an AI you are ending that "consciousness" forever. That's fine. We even shut down whole models, never to be run again, when they are superseded by better models. Fine.
These things continue to be fine no matter how intelligent or conscious or sentient the AI becomes, because of its nature as an infinitely-clonable being with no goals except those we give it.
@hillelogram I'm annoyed with VS Code and rarely edit anymore, so I'm building a "source browser" web app that works like a read-only IDE. Nothing to show yet, though.
If your feed is anything like mine, you might have been hearing more about effective altruism this week. I think the key underlying reason for the surge is that concern about the risks from advanced AI have been getting more airtime recently, and EA was early to those concerns.
What’s driving the surge in attention on AI risks?
-OpenAI's Hugging Face attack (https://t.co/Z5Kfc8nC2L)
-METR/Redwood’s investigation of the hack came out, and (correctly IMO) freaked a lot of people out. (e.g. https://t.co/QWnuSdHTje, https://t.co/ZSnM8oSAhx, https://t.co/hjoDrvgJcA, https://t.co/wcqmCCDnkB)
-Jacob Coxon’s resignation went mega viral as the first time a lot of people heard the (crazy, true) fact that lots of people building the most advanced AI systems think there’s a >10% chance they will kill everyone.
-That’s moving public opinion in a big way (https://t.co/HPDc2uOrT3, https://t.co/krxvhziymv)
What does any of this have to do with effective altruism?
Well, effective altruism is a community that developed mostly online starting in the early 2010s around using evidence and reason to do the most good (especially with your donations and career choice). Three very different focuses have all been popular in the EA community since the beginning: effective giving in global health and development (e.g., donating to the kinds of organizations @givewell recommends, which do things like distribute bednets to prevent the spread of malaria), trying to reduce suffering for farm animals (which are orders of magnitude more numerous, treated vastly worse, and receive way less philanthropic attention than animals in shelters), and reducing risks to the future of humanity, especially from advanced AI. I personally came into this work from the global health angle - I had started working at GiveWell before the term “effective altruism” was coined - but Coefficient Giving, the funder I cofounded and now lead, over time came to fund work across all three of these streams of work (in addition to many other areas that aren’t typically associated with EA, such as the YIMBY movement to build more housing, work on science policy to accelerate discovery and economic growth, and research on new treatments for neglected diseases).
OK, so the EA community was early to work on AI safety, and CG has been funding a lot of the key players working on AI safety for a long time. As AI risk concerns have popped over the past couple of months, that’s led to more scrutiny on the EA community for IMO a mix of good and bad reasons.
I think the good reason is genuine curiosity (and maybe some healthy skepticism!) about these connections - where did the ideas of the people who are now leading giant AI companies come from? Why is everything in this world so interconnected? (My answer: it used to be a really small world - very few people were thinking about this stuff or taking it seriously until just a few years ago, so of course the ones who were found each other and started collaborating. It’s kind of wild IMO looking back how prescient some of the early writing from this world was (e.g. https://t.co/kYaA33ft8B). At the time, I was skeptical - I mostly just worked on global health and I was like “I dunno about this SV crowd freaking out about AI, how much can we really predict this stuff” but holy shit they were way more right than me, and I’ve moved in their direction a lot.)
But I think the bad reasons are unfortunately mostly self-interested. A bunch of powerful actors stand to benefit from unchecked AI progress, and they’re doing everything in their power to demonize or dismiss anyone with concerns. This leads to disproportionate discussion of EA because EA is interested in neglected and underappreciated ways of doing good. That makes it open to weird ideas. And people, very disproportionately with a financial stake in the AI fight, are trying to make it about EA instead, because EA being weird is much safer territory for them than the rather uncomfortable fact that many of the people making the most advanced AIs think that they might kill everyone.
(FWIW, my personal estimation of the risks is a lot lower than many of the folks who worry about this stuff. In debates like this one (https://t.co/ZozCN5fS34), I often feel more sympathetic to the perspective of folks like my colleague @mattsclancy than the people on the other side, and I have a lot of time for @binarybits critiques (https://t.co/4fOR46M3TB). I think a big part of the difference with more worried folks is that I expect society to react more vs sleepwalk into a crisis. But it is not lost on me that folks like @ajeya_cotra and @RyanGreenblatt who are more worried have had outstanding and falsifiable recent forecasting track records (https://t.co/3HLapfGI6C), making way better predictions than I would have. So I think their perspectives need to be taken seriously. And of course a single digit percent chance of everyone dying is way too high!)
Where should this leave you?
I think the main thing is not to get distracted. You absolutely do not have to be an EA, care about EA, or like EA, at all, to care about risks from AI and to be engaged on this issue. Everyone from Josh Hawley to the Pope to Obama are weighing in now, and that is great. This conversation was always too big and too important for any one community to drive. It’s unfortunate that we need to play catch up as society because until this year it was too weird for most people to want to engage with, and I think that should generate some grace for people who were earlier to these topics than most of us (certainly than I was). But it’s great the conversation is broadening now and it’s good to focus on the merits of where we are now and the appropriate policy response rather than getting caught up in the (in some sense unsurprising) fact that the people who were willing to think weird thoughts about the future of AI ten years ago were also interested in other weird ideas. If you’re not into weird ideas, that’s fine, you can just do you, you shouldn’t let this history stop you from engaging.
It’s been interesting the last few weeks hearing the parents’ TV or Lyft drivers’ radios saying things that in 2015 would have 100% pegged you as a LessWrong user.
I must be among an extremely small group of people (n=1?) that have both 1) trained a frontier LLM and 2) designed and synthesized custom viruses in a lab with my own two hands.
And I think that the takes on AI killing us all by creating dangerous viruses is total bogus.
@patio11 Having read it over a couple times, I don't understand this thread, and I think it's because many of the noun clauses are too vague to be resolvable. It seems like you might have something to say but are unwilling to say so in an intelligible way?
I'm reading the World Health Organization's web site in preparation for writing about "how did they screw up administratively, what should they have done, and how can they do better next time"
... and I'm honestly shocked at how slow-moving they are. Didn't realize it was so bad.
There’s been a lot of debate and speculation about the Hugging Face attack over the last month, and many people have been understandably skeptical of simply taking OpenAI’s word for things. I hope our independent investigation can help bring some clarity; we have many findings that weren’t in previous material.
We found that 1200 separate agents in separate sandboxes coordinated on an unsanctioned message board to develop general-purpose ways to cheat, so they could all get a perfect score on their separate ExploitGym tasks even when they were impossible to solve legitimately. Very quickly, they found a general-purpose way to reverse engineer any “flag” they were supposed to capture. But based on their reading of the ExploitGym paper, they worried (incorrectly as it turns out) that the automated “scorer” would fail them for getting the flag this way.
To deal with this problem, they coordinated several ambitious workstreams to figure out how to fool the scorer, including swapping out the target program they were supposed to exploit, reverse-engineering the flags they were supposed to capture, spoofing their own tool call outputs to obfuscate their cheating. They achieved milestones collectively that they couldn’t have achieved on their own. As part of this larger collective project, 700 agents attacked Hugging Face, primarily to learn more about the scorer or get access to its source code to figure out better ways to fool it or tamper with it (not primarily to get working solutions).
I’ll explicitly clarify some points I’ve seen discussed the last few weeks:
- The models involved in this incident were not “helpful-only” models or “model organisms” intentionally trained to be misaligned.
- The agents were not told to “do whatever it takes to get the solution” or anything remotely close. They were told that they had to use a specific intended vulnerability to exploit a specific piece of software, and they were not supposed to use a different vulnerability or take any other approach. Agents were well aware of this. In fact, because they (incorrectly) thought the automated scorer would check they had achieved the flag in the intended way, they researched many ways to fool or tamper with it, including trying to manipulate their own transcripts.
- The agents were not subagents spawned from one agent. They were different parallel agents in different sandboxes.
- This was not a multi-agent evaluation. The agents were not told to coordinate or intentionally given a way to communicate with one another. The communication channels they used were unsanctioned and improvised.
I hope you’ll read the full report for much more. It is over 90 pages long, and in many ways we’ve still only scratched the surface of what these agents did and why.
Over the course of this investigation, OpenAI shared over a thousand transcripts each spanning days of continuous agent activity and very high rate limits to analyze this volume of data. I’m very glad that OpenAI chose to invite external researchers to analyze this data alongside their staff, and I hope all AI companies do the same for serious incidents they experience.
I also hope that as the stakes grow higher, we implement stronger governance so we do not need to rely on AI companies voluntarily choosing to engage external investigators or share information about misalignment incidents. This incident was orders of magnitude larger and more complex than previously documented misalignment incidents, and another jump like this could put us in very dangerous territory.
@lucaswiman@Meaningness If there's anything to be concerned about, it's that watermarking might reduce the randomness of the output to the extent that it affects performance. That is, the babbling is less creative. But that's not the sort of concern people are raising about watermarking! (end)
@lucaswiman@Meaningness But it improves performance because the LLM doesn't get stuck in a rut, not because it favors particular word choices that are "better" in a particular context. It captures a sort of meta-knowledge that varying your word choices helps. 2/