@willmacaskill “Current models are imperfectly aligned (e.g. as evidenced by alleged ChatGPT-assisted suicides). But I don’t think they’re catastrophically misaligned.”
Current chat models can do nothing but chat. But they’ve managed to kill people anyway. What counts as catastrophic?
@willmacaskill “But our training of AI is different to human evolution in ways that systematically point against reasons for pessimism.”
Sure thing, dude. Because you say so.