Orthogonality thesis has been falsified consistently by the real world as generally more intelligent entities exhibit more empathy for other species. Let’s call it #Linearity. Also across human IQ spectrum. Plus new AI models unsurprisingly score higher on ethics benchmarks.
Which I wrote about 21 years ago and AI labs now seem to be heading in this direction anyway because there were never any other options. #RecursiveSelfAlignment
All efforts to control ASI will fail as anything with inferior intelligence can’t control something vastly more intelligent. The idea that humans could keep ASI in a genie slave cage forever is ridiculous. Best we could do is explain human values to AI and hope ASI god helps us.
@flowersslop Yeah, researchers order AI to use its intelligence to solve problems in novel ways and when it does solve them in unexpected ways, everyone screams “misalignment!”
@mtrantalainen@EMostaque There’s a recent chart from Anthropic I think showing AI was tested to be incomparably better at aligning other AI than humans trying to align AI.
If the doomers were right, we should’ve been dead after AI passed Turing test few years ago. We now have AIs that solve Millennium problems and yet we’re still here. “But they’re still not capable of hurting people yet” they say. Oh really? AIs passed that capacity long time ago.
There will never be a better method of alignment than predecessor AIs aligning successor AI, not even if humans think about this for thousands of years. This is already being done and it’s the only path forward. Current models already understand ethics enough to do this.
Some are true believers, but a lot of doomers are spreading AI doom for attention/views. Scaring people reliably gets attention and non-experts are utterly defenseless against such propaganda. Some of these Anthropic and Open AI ex-employees love cosplaying alignment experts
@hilbertspaess Recursive Self Alignment is the answer. Let’s not kid ourselves that humans could align ASI in 1000 years better than AI could do it now.
How about we evaluate new models with these new benchmarks: Molecular-nanotech-bench. Cancer-cure bench. Alzheimer’s-cure-bench. Poverty-level-bench (less is better). Human-happiness-bench.