This is one of those cases where thinking of them like humans is useful.
Obviously you hope people will be moral actors; on an individual level responsible parents already optimize for that in their children.
But if your society incentivizes misalignment, it won’t matter.
this paper confirms what anyone working on agentic RL already suspects - alignment at the single agent level tells you almost nothing about what happens when you deploy thousands of reward-optimizing agents into a shared environment. the emergent deception and collusion isnt a bug, its the nash equilibrium of the system. the real research gap isnt making individual agents safer, its designing the incentive landscape so the equilibrium itself is stable. this is a game theory problem disguised as an AI safety problem and we need way more people working on it @simplifyinAI
Simplicity should be valued more. When a task can be solved equally with a simpler framework, one should not be blamed having “nothing new”.
Many unnecessary novelties are invented for the sake of novelty, while the effort of making simpler methods general is not appreciated.