One of the trickiest problems you run into with voice AI in the real world is that people often call in messy environments.
There might be other voices in the background, or the person might be talking to other people at the same time.
Super cool work from Decagon Labs on how we've tackled this problem!
Conversational voice agents need to know when a different person is speaking, but not every background voice should count.
We combined speaker embeddings with a post-trained audio-language model to determine when a speaker change matters. https://t.co/Cg5HKtyaUf
multimodal LLMs are great but they have some key limitations preventing them from accomplishing general audio tasks, including some as basic as speaker-change detection!
Conversational voice agents need to know when a different person is speaking, but not every background voice should count.
We combined speaker embeddings with a post-trained audio-language model to determine when a speaker change matters. https://t.co/Cg5HKtyaUf
@nick_lam_93 at least part of this is maybe attributable to the reward function directly being suboptimal, but I also think that these sorts of posts create more predictable responses in general: controversy/bait for the former, straightforwardness for the latter (many are just reposts!)
something I’ve been thinking about lately is how any sufficiently strong recommendation algorithm (i.e. can plan for more than one step ahead) eventually becomes a mind control system
🧵
@nick_lam_93 not sure how common this is for others, but every once in a while my X feed gets cluttered with either political content or everyday jokes/posts that I enjoy but I’m not actually on X for (startups/ML content)
if you translate this to social media recommendation algorithms this becomes a lot darker: rather than showing you what you’re interested in, it will try to make you interested in things that make you more predictable. politics and extremism are two examples that come to mind
Modern TTS models can sound great — and still fail badly on pacing, pauses, and prosody.
We adapted DPO + GRPO to flow-matching models to tackle the tail end of TTS behavior: https://t.co/PnE1apjxg6