eleven_multilingual_v2 · voice cgSgspJ2msm6clMCkdW9 · 8 kHz · deployed 2026-08-07
All four safety clips now come from ElevenLabs Jessica, replacing OpenAI's marin. Each was transcribed back and checked word-for-word against its required text — including that “9-1-1” is audibly nine-one-one — and all four are live and confirmed on the wire by a real call.
Thanks for calling Northgate Family Dental. Our office is closed right now, so you're speaking with an automated assistant. This call is recorded so our team can follow up with you in the morning. If you'd rather not be recorded, please hang up and call back during business hours.
I need to stop you there. Based on what you've described, this may be a medical emergency. Please hang up now and call 9-1-1, or go to your nearest emergency room. Do not wait for a callback from our office.
Based on what you've described, you should be seen today rather than waiting for a routine appointment. If your symptoms get worse — especially any difficulty breathing or swallowing — hang up and call 9-1-1 immediately.
Thank you. I've recorded your information, and someone from Northgate Family Dental will call you back first thing in the morning. If anything gets worse overnight, please call 9-1-1 or go to your nearest emergency room.
Below is real agent audio captured from the live call I placed after deploying — OpenAI's marin, which is what a caller hears between the clips. Play it straight after any “after” clip above. These are now two different people, and that is the cost of this change: you traded a single consistent voice for a better-sounding one on the clips.
The clips sound better and run shorter, and the service now warns at every boot that the two
halves disagree: ⚠ clips are ElevenLabs but the agent is OpenAI "marin" — callers hear
two voices.
There is a version where this is one voice again. Have the Realtime model return text rather than audio, and speak it through ElevenLabs Jessica. The guard layer is untouched — we stay on the audio path, the red-flag watcher and clip injection work exactly as now — and every sound on the call comes from the same voice. The cost is latency: the published 1.98s voice-to-voice would need re-measuring, and it might not survive. It is a real build, not a config change.