What makes a good sample
Five to fifteen seconds of one person speaking normally, with no music, no second voice and as little room echo as possible. A phone recording in a quiet room beats a studio recording with a backing track.
Longer is not better. Thirty seconds of speech with a cough and a door closing in it produces a worse result than eight clean ones, because everything in the sample is treated as part of the voice.
Use it on a voice you have the right to use
Your own, an actor who agreed, or a synthetic one. Cloning somebody else’s voice to make them appear to say something is impersonation, and in a growing number of places it is also a crime.
This is a tool for narrating your own work, keeping one voice across a series of videos, and reading text aloud. It is not the other thing, and nothing about it being easy makes the other thing acceptable.
Your voice sample, and how long it stays
Both the sample you upload and the audio it produces go through a server — this is one of the AI tools here that works that way. The sample goes straight from your browser to storage on an address signed for five minutes.
Everything is deleted after about three hours: the sample, the result, all of it. No voice is stored, none is reused for anything, and there is no profile for one to be attached to.
Frequently asked questions
What if I do not upload a sample?
You get a clear default voice reading your text — ordinary text to speech, with nothing to prepare.
Which languages can it speak?
It follows the language of the text you type, and the accent tends to follow the sample. A short English sample reading Chinese will sound like someone reading a second language, which is sometimes what you want and usually not.
How long can the text be?
A few paragraphs at a time. For something longer, do it in sections and join the audio afterwards — the audio joiner on this site does that in your browser, with nothing uploaded.
What audio files can I upload?
MP3, WAV, M4A, AAC, OGG, FLAC and WebM, up to fifty megabytes — a much larger allowance than the image tools get, because a recording of somebody talking is legitimately tens of megabytes and a photograph that size is a mistake. A voice memo straight off a phone is exactly the right thing.
Something wrong with this tool, or an idea for it? Tell us