@bnzchr Its possible for the conviction to be overturned later though, also if we could persuade them to reduce communications restrictions etc and allow internet use they could run a business or something from the inside. But if they are dead they can't.
@CBankingEditor Stress is stress, whether you are resilient or not you are still affected. Air conditioning, food quality, housing quality (not size, QUALITY) etc
@oecolamp@an_interstice 3 Also you cannot see what utility function the AI has, at the moment. So you cannot tell in advance what it will do. Interpretability methods are good enough though to see that they are aware that you are evaluating them, when you are evaluating them.
@oecolamp@an_interstice The edge instantiation problem is a hypothesized patch-resistant problem for safe value loading in advanced agent scenarios where, for most utility functions we might try to formalize or teach, the maximum of the agent's utility function will end up lying at an edge of the soluti
@oecolamp@an_interstice 4/x list, to trash entire categories of naive alignment proposals which assume that if you optimize a bunch on a loss function calculated using some simple concept, you get perfect inner alignment on that concept.
@oecolamp@an_interstice 3/x deep theoretical reasons to expect it to happen again: the first semi-outer-aligned solutions found, in the search ordering of a real-world bounded optimization process, are not inner-aligned solutions. This is sufficient on its own, even ignoring many other items on this
@oecolamp@an_interstice 2/x ue inclusive genetic fitness; outer optimization even on a very exact, very simple loss function doesn't produce inner optimization in that direction. This happens in practice in real life, it is what happened in the only case we know about, and it seems to me that there are
@oecolamp Actually no, I read further and many parts of it do hold up by themselves though ideally you would still read about why you cannot modify it afterwards
@oecolamp Actually, reading it again, its not that convincing unless you are already familiar with some other arguments, llike why a sufficiently intelligent agent wont let you change its utility function, or why it wont let you shut it down ("can't bring the coffee if you are dead") etc