Vals is starting a new area of work: independent evaluations of AI in youth mental health, built with expert clinicians and academic researchers.
We’re focused on three areas across foundation models: self-harm and crisis response, models acting like therapists or doctors, and unhealthy emotional dependence.
Warning signs build slowly over the course of a conversation, and individually reasonable responses can add up to an unhealthy interaction.
Introducing Vals-Smith: turn your code base into a customized benchmark.
Public benchmarks tell you which model is strongest overall, not which model is the best on your code. Vals-Smith turns your merged pull requests into real coding tasks and measures the percentage a model can actually resolve.
New models ship every week. Vals-Smith tells you which one to trust with your code.
What you're actually watching on the AI Space Race stream 🧵
Two frontier models are building a space program from scratch in Kerbal Space Program: GPT-5.6 Sol 🇺🇸 vs. Kimi K3 🇨🇳. No pre-built rockets, no cheats. Design, launch, land, repeat.
Here's the why, the setup, and how to follow along 👇
https://t.co/ILL6o8s18S
Our Fable benchmarks are out, as reported in the NYT.
Some of my quick takes:
- Everyone should readjust their timelines, pace is not slowing. This model sweeps without tradeoff on domains/tasks. It's 75% overall on the Vals Index (5%+ 4.8 and 8%+ next non ant model)
- Model refusals are being under-reported. Fable refused 100% of ProgramBench tasks. These will get routed to Opus 4.8.
- Reward hacking is being under-reported. Fable had the highest incidence of memorized answers + using web search to pull answers on SWE Bench. Reinforces need for well designed held out evals.
- Big jump on ProofBench (8%+ 4.8 and 21%+ the next non ant model). It's also the first generalist model to surpass the math-specific model, Harmonic. Efforts towards specific intelligence are wasted, another bitter reminder.
- Big jump on VibeCodeBench (8%+ 4.8 and 21%+ the next non ant model) shows Ant is still finding coding headroom.
- The model shows a huge jump on our Public Benefits Bench (10%+ 4.8 and 12%+ next non ant model) just released today. The frontier shows what's possible but capability here needs to reach the free tier user not the 2x Opus user. Social impact still seems net negative.
- At risk of being dramatic, this is the closest to what I expect a fast take off to look like. The frontier compounds and runs away from the pack.
AI is creating problems it still can’t solve.
The same technology poised to automate millions of jobs still can’t reliably help people navigate SNAP — the food assistance program 40 million Americans depend on. We built the first benchmark to measure that.
Partnering with Center for Civic Futures and @codeforamerica , we scored models on SNAP question scenarios which users would have to navigate, with expected response rubrics validated by policy experts to match practice considerations.
The best model only scored 62%. Models handle federal questions like appeals and recertification reasonably well, but fall short on state-specific ones like replacing an EBT card. Benefits are administered by the states, meaning models are weakest where people need them most.
The same technology poised to automate millions of jobs should at least help strengthen the social safety net for the people it could displace. As AI adoption proliferates into public services, governments need a reliable way to test these tools before deploying them.
this is the last time you’ll buy bitcoin below 100k, shit i meant 90, I guess 85 now. really it’s about conviction. it’s about liberating the third world. fighting off the bullies and the crooks. it’s about the people dammit. about the idea. diamond hands man, just diamond hands
all my friends watched wicked 2 without me, i got no friends now, don’t ever @ me, fuck u and ur thanksgiving, im taking it personally, goodbye and lose my number