How Much Audio Do You Need to Clone Your Voice?
With ElevenLabs, an instant voice clone needs 1–2 minutes of clean audio, and a professional clone needs 30 minutes to 3 hours. Quality matters more than length: one speaker, no background noise and a consistent delivery beat a long, messy recording. For instant cloning, going past 2–3 minutes gives little improvement and can make the clone less stable.
The rest of this guide covers why more audio doesn't always help, what makes a clone sound off, which type you need and how to record a clean sample at home. The figures come from ElevenLabs' documentation, last checked in September 2026.
1. The short answer
ElevenLabs offers two kinds of voice clone, and they ask for very different amounts of audio.
| Type | Audio needed | Time to ready | Plan needed | Best for |
|---|---|---|---|---|
| Instant | 1–2 minutes of good audio (recommended) | No model training. It makes an educated guess from existing training data. | Starter and up | Most voices |
| Professional | 30–180 minutes of good audio | Fine-tuning usually takes 3–6 hours, longer when the queue is busy. | Creator plan and above | Your own voice, on a dedicated trained model |
Free and Starter plans have no professional clone slots, so a professional clone means the Creator plan or above.
2. Why more audio doesn't always help (instant cloning)
It's natural to assume a longer recording makes a better clone. For instant cloning, that isn't how it works.
Instant cloning doesn't train a custom model. It makes an educated guess from existing training data, which works well for most voices but can struggle with very unique voices or uncommon accents. In practice, that looks like this:
- 1–2 minutes is the recommendation. That's the amount of good audio ElevenLabs suggests for an instant clone.
- Past 2–3 minutes, gains are small. More audio gives little improvement and can make the clone less stable.
- Results vary by voice. Some users get excellent results from 30 seconds, and some get worse results from 10 minutes.
- Only total runtime counts. The number of clips doesn't matter.
So put your effort into making one short sample clean, not into making it long.
3. Why your clone sounds robotic, and how to fix it
When a clone sounds off, the sample is the first place to look. The AI copies everything in it, including ums, breaths and pauses, so whatever is in the recording carries into the voice.
- More than one speaker. Use a recording with one speaker only.
- Background noise or room echo. Re-record somewhere with neither. A closet or a blanket fort works as a DIY booth.
- Uneven volume or tone. Keep volume and tone consistent from start to finish.
- Long silences. Keep them out of the sample.
- Mixed delivery. Don't blend animated and calm reading, or switch accents, in one sample. Pick one delivery and stay in it.
- Ums, breaths and pauses you don't want. They get copied too. Record a cleaner take if you don't want them in your clone's speech.
- Too much audio for an instant clone. Past 2–3 minutes you get little improvement and possibly a less stable clone. Try your best 1–2 minutes instead.
- A very unique voice or an uncommon accent. Instant cloning makes an educated guess and can struggle here. A professional clone trains a dedicated model, so it may be worth a look (see the next section).
4. Instant or Professional: which do you need?
Choose instant if
- You have 1–2 minutes of clean audio and don't want to record much more.
- You'd rather skip a training step. Instant cloning doesn't train a custom model.
- Your voice and accent aren't unusual. Instant cloning works well for most voices.
Choose professional if
- You want a dedicated model trained on your voice, not an educated guess from existing training data.
- You can record 30–180 minutes (30 minutes to 3 hours) of consistent, clean audio.
- You're on the Creator plan or above. Free and Starter plans have no professional clone slots.
- The voice is your own. Professional clones can only be made of your own voice, even with another person's consent, and verification is required.
- You can wait. Fine-tuning usually takes 3–6 hours, and longer when the queue is busy.
- Your voice is very unique or your accent is uncommon, the cases where instant cloning can struggle.
If you're unsure, start with instant. It needs the least audio and has no fine-tuning wait. If the result isn't close enough and you meet the plan, audio and verification requirements above, a professional clone is the next step.
Ready to clone your voice?
Record a short, clean sample and hear how an instant clone sounds.
Try voice cloning with ElevenLabs →5. How to record a clean sample at home
Step 1 — Find a space with no noise or echo
You need no background noise and no room echo. A closet works as a DIY booth, and so does a blanket fort.
Step 2 — Set up your microphone
Place the mic about 20 cm (two fists) from your mouth, use a pop filter, and speak slightly at an angle to the mic.
Step 3 — Record one consistent take
Audacity is a free option for recording. Record one speaker only, keep your volume and tone steady, and avoid long silences. Choose one delivery, all animated or all calm, and keep the same accent throughout. Remember the AI copies everything, including ums, breaths and pauses.
Aim for 1–2 minutes for an instant clone, or 30–180 minutes for a professional clone.
Step 4 — Export as MP3
MP3 at 192 kbps or higher is recommended. WAV usually doesn't improve quality, so you don't need it.
Step 5 — Listen back and keep your originals
Play the file through once before you upload it. Listen for background noise, echo, changes in volume and long silences, and re-record if you hear any. Then keep your original samples: clones stay in your ElevenLabs account and can't be exported as files.
6. FAQ
How much audio do I need for an instant voice clone?
ElevenLabs recommends 1–2 minutes of good audio for an instant clone. More than 2–3 minutes gives little improvement and can make the clone less stable. Some people get excellent results from 30 seconds, and some get worse results from 10 minutes.
How much audio do I need for a professional voice clone?
A professional clone needs 30–180 minutes of good audio (30 minutes to 3 hours). It trains a dedicated model, and fine-tuning usually takes 3–6 hours, or longer when the queue is busy.
What plan do I need for a professional voice clone?
Professional cloning is available on the Creator plan and above. The Free and Starter plans have no professional clone slots.
Does the number of audio clips matter?
No. For instant cloning, only the total runtime of your audio matters, not how many clips it is split into.
Can I change my clone's accent or tone after creating it?
No. A clone's accent and tone can't be changed after you create it. Change the samples instead.
Can I make a professional clone of someone else's voice?
No. Professional clones can only be made of your own voice, even with another person's consent, and verification is required.
Can I export my cloned voice as a file?
No. Clones stay in your ElevenLabs account and can't be exported as files, so keep your original samples.
Should I record in WAV or MP3?
MP3 at 192 kbps or higher is recommended. WAV usually doesn't improve quality.
Record your sample, then clone your voice
For an instant clone, 1–2 minutes of good audio is the recommended amount.
Clone your voice with ElevenLabs →Or read the full ElevenLabs review · how it compares to alternatives