captured from the deployed service · gpt-realtime-mini · 2026-08-07
This is one genuine call against the live backend, with the audio split apart by the server's own clip announcements. Two different voices play down the same socket, and on this call the pre-rendered clips outlasted the model roughly four to one.
Proportion of audible speech on a single call.
Everything the model itself said, concatenated. The voice you selected.
The recording-consent notice (17.4s) plus the 911 clip (17.6s). Pre-rendered, played from disk, never generated live.
The take you picked, for comparison against the deployed agent above.
The clips are deliberately a different voice: when the system interrupts to say something the practice is accountable for, a caller should be able to hear that something changed. They are also rendered with an explicit “recorded announcement on a telephone line, do not sound cheerful” delivery instruction — which is why they sound flatter and more synthetic than the agent. That was a deliberate choice, and it is a reversible one.