We sat 12 frontier models at a Mafia table for 40 games.
Every role, vote, lie, investigation, and night action was logged and scored.
When it came to detection, we found that @Kimi_Moonshot K3 was the best, figuring out the mafia 72.9% of the time while @claudeai Opus 5 was only able to do so 48.8% of the time.
(🧵 below)
How do you go from jailbreaking your PlayStation to CEO of Canada's leading LLM provider? You go to the University of Toronto.
"When you go to @UofT, you get raised into AI." - @aidangomez
Wow! A future with Nvidia as the only AI company??
AI Hardware ✅
Model Training ✅
Model Inference ✅
Model Distribution ✅
They built an omni model. Are they building the omni AI company?
Benchmaxxing is the lazy way to progress.
The best voice AI benchmarks don’t even exist yet.
Embrace the ambiguity, choose a capability nobody measures well, invent the eval, and go build it.
We silenced a number in the benchmark audio. It was never spoken.
Some of the highest-scoring speech recognition models wrote it down anyway, in 30–40% of cases.
New research with @huggingface on measuring when a model is optimizing for the test instead of the audio.
Speeches in academia are too long.
Commencements, matriculations, keynotes, all of them.
You can’t inspire an audience you lost halfway through a speech
Introducing S1-mini ✨
Our first open-weights language model.
A 0.6B parameter model that processes transcripts entirely on your device. Try it in app today.
We silenced a number in the benchmark audio. It was never spoken.
Some of the highest-scoring speech recognition models wrote it down anyway, in 30–40% of cases.
New research with @huggingface on measuring when a model is optimizing for the test instead of the audio.
Such a great story!! 👏
This line really caught my eye:
“From the inside, the market didn’t look finished. It looked like it had barely started.”
That’s exactly how voice AI feels right now. Still so early, so much left to build. What a time to be building in voice 🔥