AI systems are becoming more capable, and we want to understand them as deeply as possible—including how and why they arrive at an answer.
Sometimes a model takes a shortcut or optimizes for the wrong objective, but its final output still looks correct.
If we can surface when that happens, we can better monitor deployed systems, improve training, and increase trust in the outputs.
In a new proof-of-concept study, we’ve trained a GPT-5 Thinking variant to admit whether the model followed instructions.
This “confessions” method surfaces hidden failures—guessing, shortcuts, rule-breaking—even when the final answer looks correct.
https://t.co/4vgG9wS3SE