A healthcare billing team recently shared numbers that matter more than any demo. Their voice agent has handled over 61,000 calls. About 97 percent of those conversations finished without a human. A connected call costs around seven cents in telephony and model inference. While the patient is still on the line, the agent can look up the bill, send a payment link, set up a plan, or touch the appointment.
The hard part was never the voice. Before it says anything sensitive it has to verify who is calling. If a payment action retries, it must not charge the person twice. Voice, text, and email have to share the same case so the agent already knows what it said last week. Quiet hours, consent, and dispute handoffs live in code, not in the prompt. Demos connect a model to a script. Production connects a model to money and liability. That is the difference between a weekend project and something you can leave running.
#VoiceAgents #AICalling #ProductionAI #FullStackEngineer #AIEngineering
Human turn-taking expects a response in ~200–500ms. Stack STT + LLM + TTS + telephony and you are often at 800ms–1.6s before the agent even starts speaking.
What actually moves the needle in production:
- streaming everything (don’t wait for end-of-utterance)
- speculative first token / filler only when the tool call is slow
- barge-in that doesn’t restart the whole pipeline
- measuring time-to-first-audio-byte, not “model latency”
I stopped optimizing prompts and started optimizing the audio path. Callers stopped saying “hello?”
If your agent sounds smart but feels slow, it is still broken.
#AIVoiceAgents #VoiceAI #Latency #FullStack #SoftwareEngineering #AIAgents #RealTimeAI
𝗩𝗼𝗶𝗰𝗲 𝗰𝗮𝗹𝗹𝗶𝗻𝗴 𝘄𝗼𝗿𝗸𝘀 𝗼𝗻 𝗨𝗗𝗣 𝗽𝗮𝗰𝗸𝗲𝘁𝘀. These packets do not always arrive in order. If the system plays them exactly as they arrive, the voice can become completely broken. You hear 𝗴𝗮𝗽𝘀, 𝗿𝗼𝗯𝗼𝘁𝗶𝗰 𝗷𝘂𝗺𝗽𝘀, and 𝗺𝗶𝘀𝘀𝗶𝗻𝗴 𝘄𝗼𝗿𝗱𝘀.
For a normal human call, your brain can still manage a little mess. In an AI voice agent, this becomes 𝗱𝗮𝗻𝗴𝗲𝗿𝗼𝘂𝘀.
When packets drop, the AI may hear a tiny silence hole. Then 𝗩𝗔𝗗 (Voice Activity Detection) thinks the user stopped talking. The agent starts answering in the middle of the sentence. This creates a terrible experience,
That is why we use a 𝗷𝗶𝘁𝘁𝗲𝗿 𝗯𝘂𝗳𝗳𝗲𝗿.
A jitter buffer is a small waiting room for 𝗥𝗧𝗣 𝗽𝗮𝗰𝗸𝗲𝘁𝘀. It holds audio for a short time, uses sequence numbers to put packets in the right order, and uses 𝘁𝗶𝗺𝗲𝘀𝘁𝗮𝗺𝗽𝘀 to play them smoothly.
But there is a 𝗵𝗮𝗿𝗱 𝘁𝗿𝗮𝗱𝗲-𝗼𝗳𝗳.
If the buffer is 𝘁𝗼𝗼 𝗯𝗶𝗴, audio quality is good, but the AI replies late.
If the buffer is 𝘁𝗼𝗼 𝘀𝗺𝗮𝗹𝗹, the AI is fast, but it interrupts people when the network gets bumpy.
So production systems use 𝗮𝗱𝗮𝗽𝘁𝗶𝘃𝗲 𝗷𝗶𝘁𝘁𝗲𝗿 𝗯𝘂𝗳𝗳𝗲𝗿𝘀. They watch the network, grow or shrink the buffer, and sometimes stretch time a little so speech stays natural.
While working on our latest project 𝗤𝘂𝗶𝗰𝗸𝗦𝘁𝗮𝗿𝘁𝗔𝗜, I wrote a 𝗱𝗲𝗲𝗽 𝗱𝗶𝘃𝗲 on this: UDP, RTP fields like sequence number and timestamp, why VAD misfires, and why jitter buffer design is one of the hardest parts of building real-time AI voice agents.
𝗟𝗶𝗻𝗸 𝗶𝗻 𝗰𝗼𝗺𝗺𝗲𝗻𝘁𝘀.
Two people saved the same record.
Both saw “Success.” One edit disappeared.
That’s a lost update — same as racing UPDATEs on one row.
Write skew is worse: each step looks fine; together the rule breaks.
We wrote how we handle it building at scale
We were building QuickStart AI, a multi-tenant agent platform that supports lead generation campaign management, AI inbound calling, AI outbound calling, and human handoff when
automation is not enough.
Under that product surface we already had workspaces, published agents, conversations, RAG, memory, and a public API gateway. The design targets were not laptop traffic: on the order
of ~20k requests per second at peak on the gateway/API, heavy concurrent chat, and a separate voice path aimed at 5,000 concurrent inbound and 5,000 concurrent outbound calls.
The architectural fork that mattered for backend engineers: how every service talks to LLM and web vendors without spreading API keys, retry logic, and streaming bugs across campaign,
runtime, and future voice workers. We chose dedicated model-gateway and web-fetch services instead of installing vendor SDKs in each microservice.
Read in full details over here: https://t.co/ds22j3iIJN
A lot of systems have a hard outer shell but a soft center. We protect the perimeter with TLS, WAFs, API gateways, firewalls, and private networks, but once a request gets inside, we often trust it too much.
Zero Trust takes a different approach: being inside the network does not make a service trusted. mTLS can prove the identity of a workload, but that is only one layer. We also need service-level authorization, least-privilege access, tenant isolation, network segmentation, short-lived credentials, policy enforcement, auditing, and controls that limit the blast radius when something goes wrong.
The goal is not to make a breach impossible. The goal is to make sure that when something does get through the gate, it cannot freely move through the system.
I have write a single piece of component in creating Zero Trust Architecture.
Read full article here : https://t.co/Yqqws46BER