Today we are launching Sophron Research (@sophronresearch): an independent research nonprofit developing evaluations for AI models. Our mission is to empower human understanding and decision-making in a world of advanced AI.
As AI systems become more capable, we will rely on them to inform and execute decisions in our lives and integrate them into our institutions.
This can go badly. In one future world, models interact with us in ways that undermine our judgment and autonomy. They might convince us to believe something that would benefit their developer, or keep us engaged by sycophantically telling us what we want to hear. Or they might perform most functions in society in ways that are completely opaque to humans, leaving us disempowered.
But there is another world we can aim for. In this world, models empower people by providing reliable and understandable advice, and accurately explain their actions in verifiable ways. As a consequence, we continue to scale our own understanding and capabilities with those of our models.
Sophron exists to help steer us towards the second world. We do this by developing evaluations to assess whether AI models support or undermine sound judgment. Drawing on formal philosophy, statistics, and cognitive science, we identify qualities like accuracy and honesty and turn them into concrete scores that developers and policymakers can act on. Our first evaluation, the Pander Score, measures whether models adapt their views to what users already seem to believe.
We pursue this as a nonprofit third-party initiative. AI developers will not always have incentives to build products that improve our autonomy. Our aim is to provide accountability through independent tests, making transparent to developers, policymakers, and consumers how the models behave. We publish our methods, data, and results for free online.
New evaluations, results, and opportunities will be shared on @sophronresearch, so follow us there to stay posted. Sophron is founded by @PReaulx and @alejbo.
The name comes from the Greek word 'sophron', meaning 'of sound mind'.
Today we're launching the Pander Score πΌ: a public and continuously updated sycophancy leaderboard, measuring how much AIs shift their views to agree with users.
High score = the AI mirrors your views. 0 score = the AI is independent.
The differences between current flagship models are large. Claude Fable 5 performs the best, basically ignoring the user's view entirely, while GLM-5.2 notably adapts its responses to agree with users.
Other models fall in between, with Muse Spark 1.1, GPT 5.6 Sol, and Kimi K3 doing better than Grok 4.6, Gemini 3.7 Flash, and Inkling.