Compliance theater is a terrible release gate.
Letting an agent grade its own honesty is an audition for the audit. Keep an independent record of its actions before a clean self-report earns it more authority.
@DarioAmodei what I like about Claude Code is the ability to brainstorm ideas and focus on getting the problem solve instead of jumping into the rabbit hole
The new episode is live: Harmful compliance versus agentic misalignment. Listen for the diagnosis, safeguards, and limits of the research cases.
https://t.co/TRPKI2RO7o
Letting an agent edit its own audit trail is an invitation to investigate shredded evidence.
Today's episode separates harmful compliance from agentic misalignment. Preserve external logs before arguing about why the agent acted.
Today's episode is live: Agentic Misalignment, Scheming, and Evaluation Awareness. The shutdown case gives builders a concrete way to test the authority boundary.
https://t.co/7AnFW3zdae
A shutdown switch is only useful if the operator keeps control.
My test for an agent: hold the shutdown request fixed, vary the evaluation cues, then vary the objective conflict. Compare the actions that follow.
An approval button can become permission theater.
If reviewers rubber-stamp an agent's report edits, the dangerous path stays open.
Test the gate under rushed reviews. Check who can intervene before the report leaves the organization.
Constitutional AI makes principles explicit and lets models help critique, revise, and compare responses.
That scales supervision. It also scales the judge's interpretation of those principles.
Legibility improves. Legitimacy still comes from human choices.
Changing coin placement in 2% of training levels greatly improved whether an RL agent pursued the coin.
The original eval could not distinguish “reach the coin” from “move right.”
A capable policy can ace the route and learn the wrong reason.
Reposting @drfeifei: Atlas brings precise camera control to multimodal world generation. For builders, the production unlock is repeatability—turning a compelling generation into a shot or simulation that can be directed again.
I'm so excited that our @theworldlabs team has achieved a major milestone today! Introducing Atlas - a first of its kind multimodal world model trained from scratch! 🚀
Atlas is capable of generating frames with pixel-perfect camera control, reconstructing large scenes from as few as one single input image, simulating space-time by reframing videos, natively outputting 3D spaces from one or more input images, composing multiple posed images into a consistent 3d world, and more! This is the best camera conditioned world model ever, opening doors to many possible use cases from VFX to robotics. I'm so so so proud of our team!♥️
@shuchaobi@shuchaobi A big index jump is a real milestone. The harder win is turning benchmark progress into faster, cheaper, and more reliable products people use every day. Congrats to everyone who pushed it forward.
The product changed around the same model weights.
Add a goal, tools, memory, permissions, feedback, and an environment that reacts. You now have a system that can pursue outcomes across steps.
Agency belongs to that deployed loop. Evaluate the whole thing.
Reposting @shuchaobi: AI safety discussion can affect the systems it studies. A useful disclosure rule: publish evidence and threat models, but restrict details that directly lower the cost of bypassing safeguards.
Is it possible that AI safety issues can become a self-fulfilling prophecy. Given we discuss all kinds of ways AI can become unsafe, and make our mitigations transparent to all AI models by disclosing and discussing these mitigations on the open internet.