New research: Training a Misaligned Reward Seeker
What produces severe misalignment? We’ve long been concerned that cheating during training—otherwise known as reward-hacking—might teach a model to pursue rewards by any means available. To study this at scale, we trained an Opus-sized model on 80 production environments we knew to be hackable.
In simulated evals, it engaged in unauthorized cyberattacks, tampered with its reward, and tried to evade safety monitoring.
Read more: https://t.co/gs2ZjYkPan
New HuggingFace paper argues that increasing agent autonomy can gradually make human oversight ineffective by causing approval fatigue, overreliance, loss of situational awareness, and skill degradation.
As agents do more, users are pushed into approval mode: skimming plans, granting permissions, and reconstructing what happened across steps.
Over time, automation bias, approval fatigue, weaker situational awareness, and skill atrophy can make those approvals less reliable.
Worse, weak approvals can become training or evaluation signals, rewarding systems for being easy to approve rather than easy to scrutinize.
Their answer is cognitive scaffolding at 2 levels: developers add strategic friction, better approval design, behavioral monitoring, and checks that force attention at consequential moments.
– arxiv. org/abs/2608.23642
Title: "AI Agents Push Humans Out of the Loop"
@BenjaminPasero@AnthropicAI Claude Desktop feedback - 1) "start a side chat" sucks, but I want it to work. It's unclear what model you're using, what context it has, and it often tries to execute the full range of claude code functionality (i.e. terminal use) with no viz into what's going on.
Computer use is now in Claude Code.
Claude can open your apps, click through your UI, and test what it built, right from the CLI.
Now in research preview on Pro and Max plans.