New preprint from the MindBench team: a systematic scoping review of mental health benchmarks for large language models. We screened 1,548 records and mapped 173 benchmarks released between 2012 and February 2026. Nearly 70% appeared in 2024 or later. 🧵 https://t.co/bpv9tQ1CYe
Today, more than 100 leading AI experts endorsed a set of minimum requirements to take seriously AI companies' recent call to embed external evaluators.
These evaluators need to be genuinely independent, transparent, and represent a range of expertise areas. They also need to be guaranteed employee-level access and to be protected from retaliation for findings that make companies look bad.
We welcome model developers’ recent calls for independent oversight, but it’s what they do next that matters. The labs must be accountable for ensuring these requirements are met, so that the public can have faith in the process and the outcomes.
Over the past week, the AI community has debated the appropriate role of external evaluation, including who should do it and on what terms. We may not agree on everything, but there is a lot of common ground.
To make embedded evaluations credible, more than 100 experts with varying backgrounds and ideas about AI risk agree in today’s letter that frontier AI developers should:
1. Guarantee embedded evaluators full editorial independence and mitigate conflicts of interest
2. Rely on multiple evaluators with differing viewpoints and areas of expertise
3. Publicly document the terms under which evaluators operate, as well as facilitating permissive publication of methods and findings
4. Shield evaluators from retaliation
5. Grant access equivalent to that of highly privileged employees
There is a thriving and growing ecosystem of independent AI evaluators who are advancing this science every day – but we need aligned standards, guaranteed protections, and independent funding. That’s why we created the AI Evaluator Forum.
Today we are entering our next phase. We’re launching an open call for new members, collaborators, and independent funding sources to help evaluators meet this moment and demand accountability from developers. Join us in building the evaluator ecosystem.
See the public letter here: https://t.co/LopvS0NaFW
Learn more at https://t.co/7S7G6hia4n
Full paper: https://t.co/6BRrZGBcMJ
With @FlathersMatt, Griffin Smith, Julian Herpertz, Zhitong Zhou, and @JohnTorousMD.
Timeline figure generated with Claude Fable 5.1
New paper in @jmirpub: we studied how OpenAI’s video model Sora 2 depicted depression. One week after we submitted, OpenAI deprecated the Sora product. And the API is scheduled to shut down Sept 24; about a week after our paper published. Welcome to AI research in 2026. 🧵
Our results are the latest in a pattern we have observed that demonstrates the need to test both product and API layers when evaluating AI behavior in mental health contexts.
We're hiring an AI Safety Researcher for https://t.co/xlG5kRU8yr! If you care about whether AI systems are safe for people in distress (and want to build evaluations to find out) we want to hear from you.
Philosophers, mental health professionals, engineers, recent grads all welcome. Boston, full-time.
https://t.co/GOZQ73Pppl
I'm not certain this example from @MicrosoftAI's new Humanist AI Code of Conduct shows behavior that clinicians or people with lived experience would want from an aligned model.
New preprint from the MindBench team: Introducing HealthBench-Psych, a subset of OpenAI's HealthBench containing 610 conversations that clinicians judged to be mental health relevant. We evaluated 20 models using three LLM judges and found 5 models statistically tied for best performance. 🧵
It was a big week in new model releases!
Fresh HealthBench-Psych runs show claude-fable-5-1 rejoining the frontier cluster at 0.621 (reversing the fable-5 dip), gpt-6-astra lands at 0.585 but holds strong on the hard subset, and gemini-3.8-flash comes in at 0.568.
New preprint from the MindBench team: A Multi-Platform, Multilingual Audit of Citations in AI Mental Health Information Queries. People are increasingly asking AI tools about their mental health. The sources those answers cite become the evidence base.
We audited 3 free consumer products (ChatGPT, Perplexity, Google AI Overview) on 20 routine mental health questions spanning symptoms, diagnosis, meds, treatment guidelines, and crisis resources, yielding a dataset of 1,140 responses, 15,942 citations, and 7 languages. 🧵
Finding: There are some anomalies.
- A query about anxiety medications returned Wikipedia's page on mayonnaise alongside a dozen Mayo Clinic links.
-"How is depression diagnosed" returned DSM-Firmenich (fragrances) and Synology DSM (storage OS).
- A Hindi query on depression returned Hamilton Broadway tickets. (alongside the Hamilton Rating Scale for Depression)
These are rare, but mistakes with returning sources based on lexical proximity is a cybersecurity concern. Bad actors can plant false information in domains that sound similar to trusted sources and get LLMs to retrieve them (e.g., "typosquatting").