Seeing millions of learners around the world use Speak to build fluency and confidence has been the most rewarding part of the journey.
Crossing $100M ARR is an exciting milestone and one step toward our goal of building the world’s best superhuman AI language tutor.
What's next? We’re deepening our push in the US, expanding into new languages and growing our enterprise arm. Join us, we're hiring!
https://t.co/4tLshWo1Kc
Amazing technical writeup worth reading in full - kudos to the gpt-live team on this work! It's clear that full duplex/bidi is the future of voice and it's really cool to peek under the hood at the complex systems engineering that makes it all "just work". Some observations/notes for voice agent builders:
- The whole post is a huge flex, but especially turning the “improve startup latency” ticket into a new webrtc handshake protocol. Generational scope creep. @juberti what was the original estimate on that issue? :)
- The voice must flow. Nothing matters but continuous realtime audio flow and smoothness. Be extremely clear in your voice agent’s system design about division of responsibilities between the realtime band and async channels/sidecars, and treat the async RPC boundary as a first-class conversational & UX design principle.
- Many other details are direct consequences of optimizing for voice fluidity: prompt caching and avoiding audio-blocking prefills, live context compaction with blue-green-style cutovers, smoothly handling long-running voice sessions (that are much longer than gpt-live’s context window)
- Future models and versions of gpt-live will continue to become more intelligent and handle more tasks natively in the realtime band. As inference speed improves, the model will also be able to spend more tokens thinking and delegating within the realtime response latency budget. This means both the actual delegation and “perceived delegation” boundary will push out.
- Interesting details on the turn-transcript heuristics, e.g. “we prioritize coherence in the displayed assistant responses even when the user speaks in the middle.” The full bidirectional/duplex model arch breaks free of modeling the convo as a traditional ordered list of chat messages, but heuristic/projected semantic turn messages are still critical for most voice agent harness systems you’d want to build around the core live model.
- Realtime full-duplex sessions load your voice agent application server’s CPU (stream handlers, queues, network paths) more heavily than almost any other ordinary application type. Think about the fact that each full-duplex session processes ~100 audio frames per second across both directions. This is a very nontrivial driver of higher infra/capacity costs when scaling realtime voice experiences.
- The bar has been set extremely high for long-running voice agent infra providers, with a lot of additional complexity to support very smooth realtime voice model inference with an explicit temporal dimension. It’ll be interesting to see the industry patterns that emerge here.
"The success of my [agentic influencer management dashboard] changed my career — I went from a marketer who manages a process to one who builds the systems that run it."
Naoki, Speak's Influencer Marketing Manager, shares his learnings on how non-eng roles can leverage agentic tools to stand out in their roles!
https://t.co/Vl2tBR6Sxf
Thank you to everyone who stopped by our booth last week at #ACL2026! We loved meeting with thought leaders at the intersection of ASR, LLMs, and NLPs.
The conference was full of thoughtful conversation, boundless coffee, and invigorating sessions on some of the latest challenges we're tackling in modern ASR for pronunciation training.
If solving problems at the intersection of Voice/ASR, LLMs and EdTech is interesting, come join us!
Happy Independence Day from Speak! If you're at #ACL2026, swing by our booth (Harbor Foyer Booth #4) to see what we're building, grab some swag and meet our eng/ML team.
We're hiring across our EPD team including: Voice ML engineers, Assessment ML engineers, AI Product engineers and Engineering Managers!
Tool of The Day - Day 19
name - Speak
→ what it does
an AI language learning app built around helping people actually speak.
instead of focusing mostly on vocabulary, streaks, or passive lessons, @speak pushes users into real conversations from day one.
it combines ┐
- structured lessons
- AI tutoring
- speech recognition
- pronunciation feedback
- roleplay scenarios
- personalized practice
into one speaking-first system.
the goal of that system is to learn a language by using it.
→ why people care
a lot of people spend months or years learning languages and still freeze when it's time to talk.
they understand it in videos.
they know the grammar.
they recognize the words.
but can't make a conversation with it.
Speak removes that gap by creating low-pressure practice where people can talk repeatedly, make mistakes privately, and get immediate feedback.
so instead of waiting for a tutor or language partner, users can practice whenever they want.
→ how people actually use it
people usually open Speak for short daily sessions.
common usage ┐
- practicing English before interviews
- preparing for travel conversations
- improving pronunciation
- simulating real-world situations
- replacing expensive tutors
- building confidence before speaking with native speakers
- maintaining language habits
many users treat it like a daily gym session for speaking.
→ extra proof
- founded by Connor Zwick (@connorzwick) and @adhsu
- backed by OpenAI Startup Fund, Accel, Khosla Ventures, and other major investors
- crossed $100M+ ARR publicly in 2025
- reached 10M+ learners across 40+ countries
- users reportedly spoke over 1B sentences in 2024
- scaled from YC W17 into a unicorn-valued company
Such a pleasure hosting @youtubejocoding at our SF HQ for a behind-the-scenes office tour! If you see yourself working at an AI-native startup with some of the sharpest minds in ASR, come check us out 🤓
https://t.co/tY7c0nl92E
"If I wasn't sure AI was going to change how everyone works before I joined, I am now."
Our newest engineer's 90-day reflection on building @ Speak and what it's actually like when the whole company is AI-pilled!
If you're an AI-pilled operator/eng/GTM, check us out 🤓
https://t.co/UveHieWCaz
Great first-hand account on what it's like to build @ Speak as an AI-native PM!
If you're excited to do your life's work with a diverse, genuine, and fun team... slide into our Ashby 😉
https://t.co/3CKfCPYqpN