This is a big deal.
Most of the voice AI community says that speech-to-speech models aren’t ready for production use cases.
Two reasons:
1. Reliability (more hallucinations)
2. Cost (s2s models are expensive)
On (1), Grok’s Voice Agent API is already running at large scale across Grok’s apps, in Tesla vehicles, and in call centers. There’s more work to do, of course, but progress is being made quickly.
On (2), you get SOTA performance for $0.05/minute, which meets or beats aggregate cascade model (STT+LLM+TTS) pricing.
Excited to partner with @xai on this launch — you can build a custom Grok Voice Agent with workflows, tool calling, the whole shebang, in a few lines of @livekit code.
The future of speech-to-speech is bright!
Devs building in voice and video AI, we’ve got something special cooked up for next week — on 4/30, we’re doing a fireside chat with @juberti in SF.
Justin is a legend. He created the WebRTC protocol, led dev for Google Meet and Stadia, started @FixieAI, and is now the Head of Realtime AI at @OpenAI.
Come hang out with us @fdotinc, there will be lots of food and good vibes!
Announcing LiveKit Agents 1.0 and a $45M Series B
Back when we launched ChatGPT Voice Mode with OpenAI, voice AI was not a thing. Now it's a whole ecosystem of companies, products, and tools.
LiveKit’s infra for building and running voice AI agents is also at scale: over 100K developers use LiveKit Cloud and collectively doing over 3 billion calls a year in production.
Today we're announcing Agents 1.0:
🧠 Multi-agent orchestration engine
🌍 New turn detection model (13 langs, <25ms CPU inference)
📞 Robust telephony stack (used by 25% of 911 dispatch and more)
🧱 Cloud Agents deployment platform for running agents at the edge
We’ve also raised a $45M Series B led by Altimeter to continue building towards an all-in-one platform for AI agents that can see, hear, and speak.
⌨️🖱️ 👉 📹🎤
Sound on for this one.
We wired up @OpenAI's new STT and TTS models into a single voice agent. The results are super fun! Link to a live demo in the next tweet.
Today we’re launching our first homegrown AI model: an open source turn detection model for building voice agents.
Instead of relying solely on voice activity detection (VAD), which only considers when a user is speaking, our model also considers what has and is being said in the context of a conversation and predicts when a user is finished expressing their thoughts before the agent responds.
Conversations with AI voice agents using this new model flow much more naturally without constant interruptions from the AI— check it out (more videos, details, and code in the thread):
Put your speed drawing skills to the test with LivePaint - A fun game that pits you against your friends to draw for a realtime AI judge and looks like Windows 98.
🖌️ ➡️ https://t.co/OZ8enSzOia
🧵 for more!
We had a blast at @OpenAI's DevDay! Here are our favorite features that were just announced, plus fresh code samples from @shayneparlo to help you start building with them today:
I think many companies would benefit from doing annual product launches. Beyond the marketing benefit, the forcing function is often a game-changer for an organization
For the past few months the @livekit team has been working with OpenAI to give developers access to the same technology that powers Advanced Voice in ChatGPT.
With the new Realtime API, you can build voice AI that understands the nuances of human speech and responds in 300ms with humanlike expression.
Under the hood, the Realtime API is built on websocket, which works well for server-to-server communication. But for delivering voice data to client devices, websocket isn’t the best choice. It’s not resilient to network congestion and wasn’t designed for transporting media in real time.
This week, we released the Multimodal Agent API in our Agents framework. It completely wraps OpenAI’s Realtime API, abstracts away the raw wire protocol, and provides an ultra-low latency WebRTC transport between GPT-4o and your users’ devices
We also released:
Hooks, components, and visualizers for building voice AI frontends on web and mobile.
“AI needs UI” — @AndrewYNg
AgentsJS — build voice AI applications in Node environments:
https://t.co/IYj1yr6ckE
A Realtime API playground:
https://t.co/c6hA9XDQpk
We wrote some words too:
Build your first app with the Realtime API: https://t.co/hMRrdsRmos
Blog post with more details: https://t.co/t1jdWkaRPr
We’re going to see a totally new class of voice AI applications get built with this API. I’d love to hear about what y’all use this API for and help in any way. DMs open!
Hybrid Custody on @flow_blockchain just clicked for me, and it's a game changer...
It could have a phenomenal impact in bringing in the mainstream to web3
Here's a simple explanation for anyone curious 🧵
We've announced our first Flow Hackathon! Come join us for a global competition with $500k in prizes across a range of tracks, including:
- Best mobile experience
- Best use of walletless onboarding
- Extending the ecosystem
Learn more at https://t.co/AO11VOTVF1
We’re excited about the experiences that hybrid custody will unlock for users on @flow_blockchain and if you’re interested in building them with us, please reach out! Read on for more:
https://t.co/wn9aaPlv9C
It’s no secret that moving from Web2 to Web3 can be overwhelming - if we’re to help mainstream users embrace blockchain technology, we need to illuminate a path that’s accessible and enables real ownership over their digital worlds. 🧵
In this hybrid custody state, the user can simultaneously use their deck of cards seamlessly within the game while being free to take the deck of cards and list it for sale on a marketplace. All without needing to transfer the deck between accounts!