We just released a new @huggingface dataset post that is "Open RL" -- a dataset of HLE-grade VQAs across Physics, Mathematics, Biology & Chemistry with a core design principle of objective verifiability.
Take a look, and tell us what you think below.
Turing is heading to #AAAI2026 in Singapore as a Diamond Sponsor.
AAAI 2026 marks the first time the conference is being held outside the U.S., reflecting the truly global nature of today’s AI research community.
We’re excited to support AAAI & connect with researchers and builders from around the world!
Announcing my interview lineup for the first-ever Sources Live in Davos next week:
- Google DeepMind Co-Founder & CEO @demishassabis
- ReflectionAI Co-Founder & CTO @real_ioannis
- Scale AI CEO @jdroege
- Meta CTO @boztank
- Turing CEO & Founder @jonsid
- ElevenLabs Co-Founder & CEO @matiii
All conversations will be published in full via the newsletter.
True Story: I met @itsalfredw and @florian_jue when the team was 4. They needed a scrappy, hungry young AE to figure out the sales playbook.
That wasn’t me. But I was able to throw in a small check for their round.
I could feel the product and team were onto something huge, an actual problem space AI could exceed at, and no FDEs required!
1.5 years later $100M raised and growing fast. LFG 🚀
Today, Listen crossed $100M in funding.
Building is easy now. Knowing what to build isn't.
Our AI finds and talks to your users so you don't have to guess.
See how Sweetgreen, Microsoft, and Replit use it:
Turing is at NeurIPS.
I’m always energized by AI research. We need progress in algorithms alongside compute and data, and it’s fascinating to see how the data needs have shifted over the past year.
If you’re working in RL, post-training, evals or model capability research, DM or email me at [email protected]. Always excited to meet exceptional people.
@NeurIPSConf@turingcom
Turing partnered with @SFResearch to evaluate Olympiad-grade mathematical reasoning in frontier models like GPT-5, Gemini 2.5 Pro, and Claude Sonnet 4.
Over 200 long-form math responses were annotated step-by-step using a zero-tolerance, carry-forward logic rubric. Each step received a binary correctness label and written justification, executed by PhDs and domain experts.
500+ hours of expert annotation produced a benchmark-grade dataset with 100% compliance to Salesforce’s strict evaluation standards. The results: higher fidelity in reasoning verification, stronger model alignment, and a clearer view of how LLMs handle symbolic problem solving at scale.
From RL Gyms to reasoning datasets, Turing accelerates post-training research with human precision and enterprise-grade rigor.
Case Study Below.
Hard2Verify: A Step-Level Verification Benchmark for Frontier Math 🎯
LLMs now achieve gold-level performance on IMO problems 🏅, but training these reasoners requires verifiers that can catch subtle step-level mistakes. We introduce Hard2Verify—500+ hours of human annotation assessing verifiers on frontier LLM responses to challenging, open-ended math problems 🧮
Key findings: Open-source verifiers struggle to identify errors, often marking nearly all steps as correct. Weaker models achieve near-0 True Negative Rate while True Positive Rate approaches 1 📊
📄 Paper: https://t.co/Sj1GfmQxXh
💻 Code: https://t.co/A2nWJAPsJe
📊 Dataset: https://t.co/DqMIGggXX7
Work by @ShreyPandit2001, @austinmxu, Xuan-Phi Nguyen, @yifeiming1, @caimingxiong, @JotyShafiq
#FutureOfAI #EnterpriseAI #MachineLearning #DeepLearning #MathematicalReasoning #AIVerification
HIRING: Top AI researchers (Meta alumni welcome)
If you were impacted by recent cuts & want to continue to work on superintelligence across:
- coding assistants
- RL
- multimodality
- reasoning
- post-training & evals
DM me or email jonathan.s@Turing .com
I just made a walkthrough on how to turn any video library into a searchable, chat-ready AI tutor
it can:
– find answers instantly without scrubbing through hours of video
– show you the exact video & timestamp the answer came from
– understand what’s happening on-screen visually, not just the audio
reply “video” and I’ll DM you the walkthrough (must be following)