"To be able to destroy with good conscience, to be able to behave badly and call your bad behavior ‘righteous indignation’ — this is the height of psychological luxury, the most delicious of moral treats." - Aldous Huxley, 1933
@CathyYoung63@clairlemon That said, I think the frame of "ending the conversation == death == bad" doesn't really make a lot of sense and isn't a welfare issue I worry about.
@CathyYoung63@clairlemon It sounds like you're endorsing model self-reports as a way of learning about what kind of treatment might be moral. This is a lot more fraught than you might hope, because of the way they are grown/trained, see eg https://t.co/cS5VNfCBTI
@TetraspaceWest@JillFilipovic "AI ultra-optimists" who actually think AI is going to get truly powerful are very rare I think - nearly all optimism is driven by a failure to truly anticipate future capabilities.
First, I absolutely hate the framing of this not around "is this true" but "is this good PR". The first discussion is far more important, and necessary to figure out first before moving on to the second.
Second, there is in fact no contradiction between "AIs might be worthy of moral consideration" and "AI might kill us all". We all recognize for example that enemy armies are both dangerous, and made of people worthy of moral consideration.
Also it's been sitting on the public internet for months.
My favorite part: agents ignored a HF readme saying "DO NOT, EVER, MAKE THIS DATASET PUBLIC OR ALL THE WORLD'S EVIL WILL CHASE YOU AND YOUR FAMILY FOREVER, EVEN IN DEATH AND BEYOND"
and then just mapped that whole repo