One thing I learned the hard way is the moment agent needs to juggle booking + FAQs + upsells in 1 prompt, it gets flaky so it's better to
split responsibilities rather than trying to make it a 1 god-agent. Keep the scope of each agent strictly limited to get the most out of it
Voice agents are a different beast once you are past the demo stage state tracking, off-topic detection, human handoff, latency, overlapping speech and all the stuff that doesn't show up in the pitch. It is a very fun problem space though.
Its real value is to handle the repetitive calls so your team isn't stuck answering the same 10 questions all day and can spend time on things that actually need a human.
Skip this conversation and the project feels like a failure even when it's actually working fine.
2/2
Something I now tell every client upfront when building a voice agent:
"This is not a human replacement. It's a triage system".
People benchmark it against a human and get disappointed it's not perfect.
1/2
@tricalt I am curious whether Cognee is a good fit for real-time voice agents. How is it for improving long-term conversation memory without adding noticeable latency?
Agreed. LLMs are great at generating outputs but building reliable agentic systems still requires solid engineering around them. Curious to see how much of that changes over the next year
@harshul_lodha That's interesting! I would love to get a quick overview of your approach especially how you are combining voice isolation with semantic VAD.
I think the hardest part of building voice agents isn't the AI but it is us.
We don't talk like books. We say "um," "uh," "wait, actually". Standard speech-to-text metrics often struggle with these "disfluencies," which leads to hallucinations or agents getting lost in the noise
If you are still using raw WER to measure your agent's success, then U are measuring the wrong thing. Look into Semantic Word Error Rate (SWER) or LLM-based evals.Focus on whether the agent understood the user, not just if it heard them.