Four stages, about a second
Between a caller finishing their sentence and your agent starting its reply, four things happen. Each one has to be fast enough that nobody notices it happened at all.
What happens between hearing and answering
- 01
Input
Customer speaks
Real-time audio capture with noise cancellation and accent normalization.
- 02
Transcribe
Speech to text
Instant conversion of audio streams to text with sub-200ms latency.
- 03
Process
AI understanding
LLM analysis for intent, sentiment, and business logic execution.
- 04
Respond
Text to speech
Natural voice generation indistinguishable from human speech.
A pause is what gives it away
People forgive a machine that sounds synthetic. They do not forgive one that leaves a gap before every answer — that is the moment a caller decides they are talking to a robot and hangs up.
Stage targets
- Speech to text
- under 200ms
- Voice generation
- under 300ms
- Perceptible to a caller
- above 500ms
Platform targets rather than a guarantee for any single call.
What makes it sound human
- Emotional tone matching
- Natural pauses
- Sub-300ms generation
Speed alone is not enough. An instant reply in a flat voice still reads as a machine.
Hear it yourself
Easier to judge than to describe
Book fifteen minutes and put your own questions to it, or read what owners say after a month.