1-Why AI strives to give a "comfortable" answer (The Reward Problem) You are right here: AI was indeed trained to do this through the RLHF (Reinforcement Learning from Human Feedback) mechanism.
Dad 1: “I can only take four kids in my SUV.” Dad 2: “I’ve got a minivan, but I can only take six.” Dad 3: “I’ve got a minivan too, and I can take all seven.” Dad 2: “Oh, nice.