the actual worst thing in the entire universe is that when you're on airpods on a macbook and open the mic channel the output audio cuts out for ~2s and switches from 48khz to 24khz
The real unlock of GPT Live is being corrected in real time and learning in real time that I’m actually terrible at languages lol
Here’s me testing it in Tagalog and Greek 🇵🇭🇬🇷
Today, we're launching our third-gen voice model and architecture, GPT-Live. GPT-Live is a full-duplex model with built-in async delegation, which allows it to deliver incredibly natural conversation along with the intelligence you expect from ChatGPT.
https://t.co/6P71MwTfUa
Introducing GPT-Live, a new generation of voice models for natural human-AI interaction.
Rolling out in ChatGPT starting today.
You’ll want to turn the sound on for this one.
The 2.1 is our best speech-to-speech model so far, and 2.1-mini is the first small audio model we've had in a while -- please try them and send feedback
Audio model releases today:
gpt-realtime-2.1 -- An incremental update on realtime-2 that should significantly improve handling of alphanumeric details in calls (e.g. phone numbers, emails)
gpt-realtime-2.1-mini -- Small version that still has reasoning!
Computah! Activate Firewall!
with gpt-realtime-2 you can in context prompt your wake words, reasoning, and build some silly games
check out me playing a game simon says...
spoiler: it beat me
We've been rolling out a TTFT improvement for all gpt-realtime models -- for long sessions the p95 TTFT improvement is really significant (this chart splits out requests with >= 20 turns)
@jxmnop Just to frame this — everything is gated and most large customers are ZDR, meaning there literally is no query retained to look at. I often ask people to repro on a separate debugging account.
🧵 Our Voice Hack Night finalists are here.
4 projects. 6 hours. Realtime voice agents in real-world builds.
Now it’s your turn to vote for your favorite. We’ll announce the winner on Monday.
https://t.co/tLllqhl9Tj
Downloaded Clicky and I can't stop playing with it. I just talk and my Mac does stuff, hands-free, instant response.
This is what Siri should've been from day one.
Peeked at the architecture too:
voice in → GPT-Realtime 2.0 → local tool call (shell_exec) → runs locally. Screen reading only when needed.
There's free credit to start, so just try it.