Voice cloning is five to eight seconds too slow for a live client session
A health practitioner put her cloned voice on the shelf over a five-to-eight-second delay. What to test before you buy a voice tool.
Cloned voice is not yet fast enough for a live client conversation. Dr. Nikki Siso, a certified holistic health practitioner and founder of Conscious Kitchen, puts the delay at "probably a good five to eight seconds" between a client speaking and the system answering. That is enough for her to describe the exchange as jerky, and it is why her own cloned voice is not in the fifteen-skill AI assessment she built in Claude.
What would she use a cloned voice for?
A spoken assessment a client walks through their own house to complete. Siso told Artificial Reality, the AI buyer-evidence podcast hosted by Galina Fendikevich, what she wants:
"I would love if AI could actually use my voice in real time to have a dialogue with someone."
Fendikevich sketched the use case in the interview: a client moving through their home during the assessment, asked out loud to open a cabinet or a makeup drawer and describe what they see, instead of typing answers into a form. Siso's own description of the payoff avoids overclaiming: "That would make the assessment feel more human is a funny thing to say — but more connected might be a better word."
What is actually blocking it?
Latency, not voice quality. Siso is clear that the cloning itself is available and names the category:
"Yes, there's ElevenLabs, there are ways to do voice cloning, but the speed of the conversation, there's a very long delay ... Right now it's probably a good five to eight seconds. It's a long enough delay to where it's jerky."
This is a specific, measurable blocker rather than a vague reservation, and it is the reason a paying operator with a built product has left a feature on the shelf. Artificial Reality notes that her five-to-eight-second figure is her own estimate from working with the tools, not a benchmark.
Why does it have to be her voice and not a good voice?
Because in coaching, delivery is part of the intervention. Siso's argument is that the same information lands differently depending on who says it and how:
"The way I say things or the way you say things might land for someone differently. You could have heard the exact same thing from someone else and it didn't register. And then when I say it the way I say it, somehow you're like: oh my God."
"So capturing that tone is very important per coach. Really important to get that right. And there's a lot of room there — to really embody, which is so humanistic, but: how do we get these AI tools to embody the human experience?"
She is also explicit that the requirement is not specific to health. Any coach downloading their knowledge into a system, she says, needs it to come back out in their own voice, "if you're a coach of peak performance, or any kind of coach, it doesn't matter."
What should you check before buying a voice tool for client work?
These checks are Artificial Reality's analysis, built from the blockers Siso describes rather than from vendor claims:
- Time the full turn, not the synthesis. The number that matters is from the client finishing a sentence to hearing a reply, including transcription and model response, not how fast the audio renders.
- Test it with interruptions. A real client talks over the question. Ask whether the system handles being cut off, because a scripted demo never is.
- Test it on a phone, on the move. Siso's use case is someone walking around a house, which is a different network condition from a desk.
- Check that tone survives length. Her stated requirement is a voice that "sticks and stays" and does not fade over a long session. Run a long one.
- Ask where the voice model and the recordings live. A cloned voice used in client sessions is both biometric data and a brand asset.
- Have a written answer for what happens when the client asks whether it is you. Decide the disclosure before you deploy, not after.
Is there a version worth running today?
The asynchronous half. Nothing about the latency problem stops a cloned voice from carrying recorded guidance, summaries or check-in messages, where a five-to-eight-second gap is invisible because nobody is waiting in a conversation. Live dialogue is the use case that breaks; one-way delivery is not.
Siso pairs that with a boundary worth copying. She automates the noticing, so that a client reporting a bad day in a program triggers a prompt to her, and she keeps a person at the end of it:
"There has to be a human at some point in the chain. So even if AI did the automatic prompting of, hey, what'd you eat today — your reply, if you knew, was gonna get seen by a human, then it challenges your integrity to not reply. That's what's missing."
So the decision this week is narrow. Buy voice cloning for what you would otherwise record yourself, time the full turn before you buy it for anything live, and keep your name on the part a client is answering to.
Where this comes from
S1E7: How a non-tech founder built a 15-agent AI business on Claude — the full interview with Dr. Nikki Siso. Listen or watch: YouTube, Spotify or Apple Podcasts.