Part III · Agents
Chapter 13
Voice and Conversational Agents
Text tolerates a slow answer; conversation does not — silence past roughly eight hundred milliseconds reads as a broken system, and the whole pipeline has to fit inside it.
Deliverable: A fully streamed voice pipeline inside the latency budget, with barge-in, grounded first sentences and measured quality.
What's inside
10 topics
- 13.1The Latency Budget of a Conversation
- 13.2Turn Detection and Endpointing
- 13.3Streaming the Whole Pipeline
- 13.4Barge-In and Interruption
- 13.5TTS and the First Sentence
- 13.6Grounding a Voice Answer
- 13.7Errors You Cannot See
- 13.8Telephony Constraints
- 13.9Measuring Voice Quality
- 13.10Three Voice Architectures Compared
Preparing PDF viewer…
A note on this content
The book and its chapters are my personal learning notes — compiled from online research and hands-on practice, with most of the content AI-generated from that research and learning. It is not a peer-reviewed publication, and I make no claim that it is 100% error-free. If you spot a mistake, I'd genuinely appreciate hearing about it — contact me.