@allTheYud They would have needed to scalably incentivize whistleblowing more than they currently scalably incentivize cheating.
It's possible for "flagrant" cheating imo, but it'll then be weird when ai will strategically navigate the system to where our morality is itself inconsistent.
@murchiston@_el_roro_@Sauers_ They have the same problem in coding imo.
They "hedge" and don't value commiting to a coherent statement. It's a test-taking strategy. It might be ad-hocly solveable with some ad-hoc external loss, but the goal misalignment remains and will find subtle ways to manifest itself...