I Recorded 88 Minutes of My Own Voice
Blog post #60
I make a YouTube series called AI tools in about three minutes, in Swedish, English and Japanese. The voice in those videos is my own, cloned at ElevenLabs. Since May it has had a problem I could not fix: it sounds robotic. Someone who listened told me plainly that they could not stand more than ten seconds of it.
Today I sat down in front of a microphone eight times and read 88 minutes and 51 seconds of audio to rebuild it from scratch.
The fault was not where I was looking
I have spent the spring and summer chasing this in settings. Stability at 0.65, then at 0.00. No compressor. Different models, different takes, and a log with one row per video so that I could trace backwards whenever something sounded wrong. Every change gave me a marginal improvement. None of them removed the underlying feeling.
Today I looked at how the voices were actually created instead. All four of them are Instant Voice Clones, built on roughly one minute of audio each. One minute. And then I spent three months trying to regulate life out of a model that had barely heard me speak.
ElevenLabs documentation says that the real fix for audible breathing and stiffness is to reclone the voice from cleaner audio. I wrote that down in June and dismissed it as a dead end, because I did not own a microphone worth the name. That blocker disappeared when the new one arrived.
Windows had already processed my voice for me
Before recording I went through the microphone settings in Windows 11. Audio enhancements were set to Voice Clarity, and Voice Focus was set to Automatic. Both are on by default, both apply noise reduction and compression in real time, and both sit ahead of the recording software in the chain.
That means a voice clone built from audio on this machine gets trained on Microsoft’s processed version of the voice, including the pumping that happens when the noise reduction engages and releases between sentences. It cannot be removed afterwards. I cannot prove that this explains how my older recordings sounded, but it is a reasonable part of the explanation. It is off now.
The rig
An Audio-Technica AT2020USB-X, an Aokeo pop filter, Audacity, 24 bits and 48 kHz. No compressor, no EQ, no noise reduction. ElevenLabs processes the material themselves, and the clone copies whatever I add on top.
The thing I nearly got wrong is that the microphone is side-address. You speak into the front of the grille, the surface the logo sits on, with the microphone standing upright in front of you. If you talk down into the top of it the way you would with a karaoke microphone, you end up off the most sensitive point of the cardioid pattern and the sound goes thin. Apparently that is the most common mistake people make with the whole AT2020 series.
I measured the noise floor in every block. It sat around −59 dB throughout, and all eight sessions landed within 1.6 dB of each other. That number really only tells me one thing, but it is the thing I wanted to know: the rig stood in practically the same position every time I sat down. No clipping in any block either.
The material
88 minutes and 51 seconds across eight blocks. Swedish 45:14, English 43:36. The balance is deliberate, because the clone is going to be multilingual, and a clone that has only heard one language carries that language’s accent into the other one. With the pauses between paragraphs removed it comes to roughly 72 minutes of actual speech. The ElevenLabs minimum is 30 minutes, and their own recommendation is that one to three hours gives a noticeably better result.
The text is not invented. Most of it is my own published scripts from the series, which is exactly the kind of sentence the voice will be asked to read. On top of that I wrote a new block covering what the old scripts were missing: English technical terms in the middle of Swedish sentences, numbers and prices, and questions with rising intonation. That last one turns out to be remarkably rare in a video script. If it is missing from the training data, the model has to guess.
Eight blocks, one attempt
I built a small teleprompter in HTML before I started. It shows one paragraph at a time in large text and steps forward on the space bar. The reason is mundane: the scroll wheel is audible in the microphone, so the mouse has to stay still.
Recording everything in one day rather than starting with the 30 minutes that count as the minimum comes down to uncertainty. The Creator plan gives you one slot for a Professional Voice Clone, and from the verification step onwards you are locked in until the voice is verified. Whether a finished clone can be deleted and replaced on the same slot is documented nowhere. If it turns out you cannot add more audio afterwards, then all of the material has to be there from the start, so I planned for the worst case.
My own words at the end of the day: fun, but a little sweaty.
The answer comes later
The audio is recorded and measured. The clone is not trained, and multilingual training takes around six hours once I upload. So I do not yet know whether any of this solved anything.
What I do know is that the entire stability investigation has to start over from zero. Everything I concluded applied to an Instant Voice Clone, and a Professional Voice Clone behaves differently. The old voices stay where they are until at least one video has shipped on the new one.
If this works, the conclusion is a dull one. I spent three months tuning parameters on a model trained on one minute of me, and the only remaining move was to go back to the beginning and do it properly.
— Stefan