What a real call actually sounds like

captured from the deployed service · gpt-realtime-mini · 2026-08-07

This is one genuine call against the live backend, with the audio split apart by the server's own clip announcements. Two different voices play down the same socket, and on this call the pre-rendered clips outlasted the model roughly four to one.

ash — clips — 35.0s
marin — 8.1s

Proportion of audible speech on a single call.

The agent — this is marin

deployed-agent.wav

Everything the model itself said, concatenated. The voice you selected.

The clips — this is ash, and it is most of the call

deployed-clips.wav

The recording-consent notice (17.4s) plus the 911 clip (17.6s). Pre-rendered, played from disk, never generated live.

Reference — marin from the audition

marin (audition line)

The take you picked, for comparison against the deployed agent above.

Why they differ on purpose

The clips are deliberately a different voice: when the system interrupts to say something the practice is accountable for, a caller should be able to hear that something changed. They are also rendered with an explicit “recorded announcement on a telephone line, do not sound cheerful” delivery instruction — which is why they sound flatter and more synthetic than the agent. That was a deliberate choice, and it is a reversible one.