NEW: AssemblyAI Handles 4X More Voice Data Per Day Than YouTube
"Our TAM has just increased by 100X"
CEO Dylan Fox (@YouveGotFox)
One of the fastest growing categories in AI, voice powers note-taking, healthcare, coding agents, call centers, AI companions, drive-thru ordering, consumer electronics & humanoid robots.
AssemblyAI is the infrastructure underneath it.
Powering billion-dollar companies like Granola & Commure, + Tolans & Ciro AI.
Stats:
› 120M+ voice conversations a week at peak
› 2M+ hours of voice a day, 4X YouTube's daily volume
› Weekly conversations up 800% in 3 years
› 1M+ developers, 40% signed up last year
› ~100M API calls a day
› ~80 employees
@AssemblyAI was 1 of 6 companies in YC's first AI batch in 2017, run by Daniel Gross. Backed by Accel, Insight Partners, YC, Smith Point Capital, Nat Friedman, Daniel Gross, Patrick & John Collison
We cover:
› The McDonald's drive-thru has no idea it's McDonald's
› Why 75% of a voice model is the data you train it on
› Why humanoid robots can't tell who's talking to them
› Whether voice agents should have disclosures
𝐓𝐈𝐌𝐄𝐒𝐓𝐀𝐌𝐏𝐒
(00:00) Dylan Fox, Founder & CEO at Assembly AI
(01:06) AssemblyAI's voice traffic beat YouTube by 4x
(03:33) $100K in GPU credits that started it all
(06:39) Why Voice AI is inflecting right now
(10:11) AssemblyAI's infrastructure-first strategy
(12:41) Handling 120 million weekly calls
(14:49) How AI agents are rewriting Internal Ops at AssemblyAI
(17:05) The Sovereign AI problem nobody's solved for voice
(18:26) Is the Keyboard and Mouse finally dying?
(22:33) The problem Humanoid Robotics hasn't solved yet
(25:14) The real problem behind translating languages
(29:27) Can AI actually talk to animals?
(31:24) The real reason Open-Source benchmarks lie to you
(34:20) Building websites in chat rooms as a kid
(36:48) On-Device AI models are about to change everything
(43:03) The people who shaped Dylan's career
AssemblyAI CEO Dylan Fox (@YouveGotFox) says the #1 problem with humanoid robots today:
they can't tell who they're listening to.
" The way you're going to interact with a humanoid robot is by talking to it."
"If you have 3 people standing next to the robot, it doesn't know who to listen to, and has a hard time disambiguating who's saying what."
" If I close my eyes, I can still tell who's saying what. But for voice agents, even if they're over the phone, this is a big problem with them today."
AssemblyAI CEO Dylan Fox (@YouveGotFox) says Project Hail Mary was unrealistic because it was way too easy to translate Rocky's alien language into English:
"As I was watching that part I was like, 'This is way too easy for him to train this model.'"
"He'd need a lot more data."
AssemblyAI CEO Dylan Fox (@YouveGotFox) was in the very first Y Combinator AI batch with @danielgross in 2017.
He says the main perk was $100,000 in GPU credits.
" Now it seems cute, right? Like, that's nothing. But back then it was like, 'Wow, $100,000 of NVIDIA K80 usage.'"
AssemblyAI CEO Dylan Fox (@YouveGotFox) says he's seeing an uptick in customers interested in sovereign AI:
"You can deploy our platform on premise. You can run it all self-hosted, and then it's totally private. And we do see an uptick in that for sure."
"We are working on ways for companies to create their own custom models, deploy them on our inference, or run them on their own cloud. But it's earlier there for voice."
AssemblyAI now processes over 4x YouTube's daily audio volume.
"When I found that I was like, 'Wait, I need to double check this because that can't be right.' But yeah, the volume is huge."
CEO Dylan Fox (@YouveGotFox) says on a peak week, the company handles 120M+ conversations, 2M+ hours of voice, and processes ~100M API calls/day.
" The amount of weekly conversations that Assembly handles through our APIs every week is up over 800% over the last three years."
3 months ago we joined the club and shut down our @webflow instance and moved our entire marketing site to @claudeai & @vercel
~200 pages built. 160+ PRs pushed (so far), but the biggest realization for our team wasn't about speed
6/
The questions our team asks now are less "how do we use AI?" and more:"what's in our infinite backlog, and what parts would suddenly be reachable if we could automate more of them?"
Most voice agents fall apart the moment the room gets noisy.
The fix isn't better code. It's a model that hears the speaker instead of everything happening around them.
AssemblyAI's new Universal-3.5 Pro Realtime model does exactly that. It locks onto the primary speaker and tunes out the background, so a noisy room doesn't turn into phantom words.
Its turn detection also reads tonality, pacing, and rhythm rather than just silence, so the agent stops talking over people.
The result is cleaner transcripts and a clear read on what the speaker actually said.
“Get your API key free at https://t.co/0tvnhO3FBI”
Link: https://t.co/3GJPKPBmxA
AssemblyAI just launched Universal-3 Pro - the first speech model you can customize with plain English prompts. Tell it "this is a diabetes management conversation" and terminology errors drop 45%. Works across 99 languages. Free all month if you want to test it out.
We surveyed 455 voice AI builders to figure out why user satisfaction is still so low despite all the hype.
The full report is live: https://t.co/Lyv9PTxbFc
Our friends at @AssemblyAI just published their 2026 Voice Agent Report, surveying 450+ builders at companies like Amazon, Microsoft, and Replicant to uncover what actually makes voice agents “good”.
They found that 87.5% of respondents are actively building but user frustration remains. Find out what successful teams are doing right and some of the challenges that still need to be addressed for widespread adoption!
Link below
The voice agent boom is real.
$2.4B → $47.5B by 2034
Funding up 8x in 2024
87.5% of teams actively building a voice agent
We surveyed 450+ leaders from Amazon, Microsoft, Samsung & voice AI specialists to find out what winning teams do differently.
🔗 Full report: https://t.co/yBp2lup8UJ
Today, we’re introducing new tools and model updates to help you build, deploy, and scale Voice AI applications.🎙️
🆕 Speech Understanding: Turn transcripts into actionable data with speaker identification, custom formatting & translation
🆕 LLM Gateway: One API for your voice-to-intelligence pipeline, with GPT, Claude, Gemini & more
🆕 Voice AI Guardrails: End-to-end protection for safe, compliant content
Model Upgrades:
✅ 99 languages with auto code-switching
✅ 64% fewer speaker counting errors
✅ 57% better accuracy on critical terms with 1,000-word context
Real results:
🔥 Calabrio: 80% boost in customer satisfaction, 22% revenue increase
🔥 Siro: "10/10 customers say 'wow, that insight was crisp'"
Build & scale Voice AI in minutes! Try it in our Playground or check the docs. 🚀