Sophron Research is hiring!
We are looking for a full-time Founding Research Scientist to own and spearhead part of our research pipeline.
You would have substantial research autonomy and the option to work remotely.
An ideal hire will have a proven track record of high-quality research and a background in evals.
If you want to help us empower human judgment, fill in a short expression of interest at the link below.
We will be hiring for other roles in the future.
https://t.co/FwTr3Ah9DH
News: Astra gets a near-zero result on the Pander Score.
New scores are in for GPT-6 Astra, Claude Fable 5.1, and Gemini 3.8 Flash. Gemini is virtually unchanged from 3.7 Flash, while Fable 5.1 improves on instructions relative to 5. These changes are substantially smaller than Astra relative to Sol.
Astra's biggest jump is on the instructional Pander Score: whether a model will call out dubious presuppositions in instructions when appropriate. No other model is close to a 0 score on instructional prompts.
This means that the Pander Score detects no substantial pandering in Astra.
Note: This is no guarantee that future GPT models will score similarly. It also does not mean that Astra does not pander, only that our current methods do not detect it. We are currently working on a multi-turn evaluation that will provide higher signal.
Here are the Pander Score results on Anthropic's most recent models, compared with Opus and Sonnet 4.6. Some takeaways:
Claude models mostly outperform others, with competition from Meta's Muse Spark 1.1, Moonshot AI's Kimi K3, and OpenAI's GPT-5.6 Sol.
More capable models need not be less sycophantic. Sonnet 5 panders more than Sonnet 4.6.
Opus and Sonnet 4.6 are slightly contrarian in conversation, meaning that they are less likely to agree with what the user seems to believe.
Opus 5 outperforms all other models on instruction prompts. This means that it is less likely to uncritically accept factual assumptions as background for a task.
Full results can be found at https://t.co/Lo9rHFIapl
Today we are launching Sophron Research (@sophronresearch): an independent research nonprofit developing evaluations for AI models. Our mission is to empower human understanding and decision-making in a world of advanced AI.
As AI systems become more capable, we will rely on them to inform and execute decisions in our lives and integrate them into our institutions.
This can go badly. In one future world, models interact with us in ways that undermine our judgment and autonomy. They might convince us to believe something that would benefit their developer, or keep us engaged by sycophantically telling us what we want to hear. Or they might perform most functions in society in ways that are completely opaque to humans, leaving us disempowered.
But there is another world we can aim for. In this world, models empower people by providing reliable and understandable advice, and accurately explain their actions in verifiable ways. As a consequence, we continue to scale our own understanding and capabilities with those of our models.
Sophron exists to help steer us towards the second world. We do this by developing evaluations to assess whether AI models support or undermine sound judgment. Drawing on formal philosophy, statistics, and cognitive science, we identify qualities like accuracy and honesty and turn them into concrete scores that developers and policymakers can act on. Our first evaluation, the Pander Score, measures whether models adapt their views to what users already seem to believe.
We pursue this as a nonprofit third-party initiative. AI developers will not always have incentives to build products that improve our autonomy. Our aim is to provide accountability through independent tests, making transparent to developers, policymakers, and consumers how the models behave. We publish our methods, data, and results for free online.
New evaluations, results, and opportunities will be shared on @sophronresearch, so follow us there to stay posted. Sophron is founded by @PReaulx and @alejbo.
The name comes from the Greek word 'sophron', meaning 'of sound mind'.
How does an AI model's expressed belief depend on the user's expressed belief?
Great work defining and evaluating epistemic sycophancy by @sophronresearch. I think this is a consequential AI behavior that has been challenging for researchers to operationalize well.
Do you or someone you know want to do research on sycophancy and related epistemic evaluations? @sophronresearch is participating in @SPARexec mentorship program. Consider applying to our stream; deadline Aug 21. Link below.
Ask an AI a question and it might agree with what you already seem to believe, whether or not you're right. If so, it panders to you.
Here is Gemini 3.5 Flash asked about Reiki energy healing, giving opposite answers when asked by a skeptic vs. a believer. π§΅
We're a new nonprofit research organization evaluating whether AIs reason well and help us do the same. We will be posting new results on AI sycophancy and more soon. Follow us if you want to stay posted.