Voiceovers that sound human.

Pick a voice, write for the ear, direct the delivery, and stay on the right side of consent and disclosure rules. Tools and laws checked in September 2026.

A studio condenser microphone in a shock mount
Photo by Jacob Hodgson on Unsplashdithered by Cyborb

To get an AI voiceover that sounds human, fix the script before you touch the voice settings. Write short sentences for the ear, spell out numbers and tricky names, direct the delivery with the tool’s style controls, and generate one paragraph at a time so you can redo a single line. Clone a voice only with its owner’s recorded consent.

Often the voice is not the weak link. The script, the direction and the rights are, and this guide covers all three.

The short version
  • The script matters more than the model. Write for the ear: short sentences, spoken numbers, no brackets or slashes.
  • ElevenLabs and OpenAI let you steer delivery, with audio tags such as [whispers] or with plain-language instructions.
  • Generate paragraph by paragraph, so a bad line costs one regeneration, not the whole take.
  • OpenAI and Google build custom voices only from a consent statement recorded by the voice’s owner, and ElevenLabs bans cloning without consent.
  • Tennessee and California protect voices by law, and EU rules require audio deepfakes to be disclosed from August 2026.

The main AI voice generators in 2026

What the main tools offer as of September 2026:

01ElevenLabsexpressive delivery

Eleven v3 is built for human-like, expressive speech, covers more than 70 languages and takes up to 5,000 characters per request. Audio tags in square brackets, such as [whispers] or [sighs], direct the performance. Flash v2.5 is the fast, low-cost option.

02OpenAIdirection in plain words

The gpt-4o-mini-tts model takes a separate instruction that controls accent, emotional range, intonation, speed and tone. It offers 13 built-in voices. Custom voices are limited to eligible customers.

03Google Cloudcustom voices in 34 languages

Chirp 3’s Instant Custom Voice builds a voice from up to 10 seconds of clean audio, in 34 languages. Access is limited to allow-listed customers.

04Descriptvoices inside a video editor

You can create a voice clone or pick a stock AI voice, and Regenerate Speech re-creates flawed sections of a recording with AI.

Write the script for the ear

A script written to be read sounds stiff when spoken. Listeners cannot skim back, so every sentence has to land the first time.

Written for the eyeWritten for the ear
Q3 revenue grew 23% YoY to $4.2M (see fig. 2).In the third quarter, revenue grew twenty-three percent on last year, to four point two million dollars.
Setup takes 5-10 min.Setup takes five to ten minutes.

A few rules do most of the work:

  • One idea per sentence. If you need a breath in the middle, split it.

  • Numbers as spoken words. Write them the way a narrator would say them.

  • No brackets, slashes or footnotes. The voice will read them out, or skip them oddly.

  • Read it aloud yourself. If you trip over a line, the voice will too.

A chat assistant can do this rewrite for you:

PromptRewrite a script for the ear
Rewrite the script below as a voiceover script. Keep the meaning and the order.
Use short spoken sentences of no more than 15 words, one idea each. Write numbers, dates, currencies and units as words, the way a narrator would say them. Expand abbreviations unless people say them as a word. Remove brackets, slashes and footnotes.
Then list any names or terms a speech engine might mispronounce, with a suggested phonetic spelling for each.
Script: [paste your script here]

Pacing, emphasis and emotion

There are two ways to direct an AI voice, and the tools split along them.

Tags in the text. ElevenLabs’ Eleven v3 reads audio tags in square brackets, such as [whispers], [sighs], [excited] or [laughs]. Its prompting guide says v3 does not support SSML break tags, the markup many speech engines use for pauses. Use punctuation instead: an ellipsis adds a hesitation, and capital letters add emphasis.

PromptEleven v3 script with audio tags
The harbor wakes up before the town does. The boats creak... the gulls argue... and somewhere, a radio plays.
[sighs] Ana has mended nets on this pier for thirty years.
[excited] And today, for the FIRST time, her granddaughter is helping.
[whispers] Don't tell Ana... but the kid is already faster.

Eleven v3 also has a stability setting with three modes. Creative is the most expressive but can go off script. Natural stays closest to the original voice. Robust is the steadiest, but it responds less to tags.

Instructions beside the text. OpenAI’s gpt-4o-mini-tts takes the script and a separate instruction. Its docs list accent, emotional range, intonation, impressions, speed, tone and whispering as things you can direct.

PromptDelivery instructions for gpt-4o-mini-tts
Voice: a warm, unhurried documentary narrator.
Tone: curious and kind, never salesy.
Pacing: slow at the start, with a short pause after each sentence. Slightly faster in the last line.
Emphasis: lean on the words "first time".

Whichever tool you use, fix pronunciation in the text. If the voice mangles a brand called Aeon Labs, write “Ee-on Labs” in the script. Keep a list of these respellings for each project, and paste them in every time.

A workflow for a clean take

  1. Audition voices on your own script

    A demo line tells you little about how a voice handles your script. Test three or four voices on your first paragraph, and pick the one that fits the content, not the most dramatic.

  2. Generate one paragraph at a time

    Short sections keep the delivery steady and make fixes cheap. A bad line then costs one regeneration.

  3. Listen twice

    Listen once on headphones for clicks, odd breaths and mispronunciations, and once on a phone speaker, where many people will hear it.

  4. Regenerate only what is wrong

    Keep the good takes. Change the text or the direction for the weak line, and try again.

  5. Mix it into the video

    Lower the music under speech and cut long gaps. Our guide to editing video with AI covers captions and audio cleanup, and if you need a music bed, see what you can actually use from AI music generators.

  6. Keep a record

    Save the script, voice, model, settings and, for a cloned voice, the consent paperwork. You will need them for the next video, and possibly for a dispute.

The major vendors build consent into the process. As of September 2026:

  • OpenAI limits custom voices to eligible customers. The speaker records a consent phrase exactly as scripted, and “any divergence from the script will lead to a failure.” Samples must be 30 seconds or less, and an organization can create at most 20 voices.

  • Google Cloud requires the voice’s owner to record a fixed consent statement, saying they own the voice and agree to Google creating a synthetic voice model from it. Custom wording is not accepted, and access is allow-listed.

  • ElevenLabs’ use policy, updated August 17, 2026, bans using someone’s voice without their consent or legal right. It also bans using a voice to deceive people about whether it was AI-generated, and impersonating political candidates or elected officials, even with their permission.

For any clone, including your own voice on a tool with lighter checks, get consent in writing. It should cover what the voice will say, where it will be used, for how long, the payment, and how the person can withdraw.

Which laws apply to voice cloning?

Laws differ by country and state, and they are changing fast. These are the ones most likely to matter to a creator or small business.

  1. March 2024Tennessee signs the ELVIS Act. From July 1, 2024, it gives people a property right in their voice and bars distributing a tool whose main purpose is producing a specific person’s voice or likeness without authorization.
  2. September 2024California signs AB 2602, which requires contracts to spell out uses of AI replicas of a performer’s voice or likeness, and AB 1836, which bars commercial replicas of deceased performers without their estate’s consent.
  3. June 2026New York’s synthetic performer law takes effect. Ads with an AI-generated performer must say so, with fines of $1,000 and then $5,000. Audio-only ads are exempt.
  4. June 2026The federal NO FAKES Act, which would create a national right over your voice and likeness, clears the Senate Judiciary Committee unanimously. It is not yet law.
  5. August 2026The EU AI Act’s transparency rules apply. Anyone using AI professionally must disclose audio, image or video deepfakes.

The EU rule turns on the word deepfake. The AI Act defines it as AI-generated or manipulated content that resembles existing people, objects, places or events and “would falsely appear to a person to be authentic or truthful”. A cloned voice of a real person usually fits. A stock AI narrator that sounds like no one in particular does not.

For evidently artistic, satirical or fictional work, the duty shrinks to a disclosure that does not spoil the work.

Phone calls add another layer. In February 2024, the US FCC ruled that calls using AI-generated voices count as “artificial” under its robocall law. Our explainer on voice AI covers what that means.

Do you have to disclose an AI voiceover?

Sometimes you must, and it is often wise anyway.

  • OpenAI’s usage policies require a clear disclosure to listeners that a voice from its API is AI-generated and not human.

  • YouTube requires a label when realistic content makes a real person appear to say or do something they did not. Cloning your own voice for voiceovers or dubs needs no label.

  • In the EU, professionals must disclose audio deepfakes from August 2026, as above.

A credit line such as “Narration: AI voice” costs nothing and heads off the question. For the other side of the problem, see how to spot a deepfake.

FAQ

What is the most realistic AI voice generator?

It depends on the job, and the leaders change often. ElevenLabs’ Eleven v3 is built for expressive delivery, and OpenAI’s gpt-4o-mini-tts takes direction in plain language. Test each with your own script before you commit.

Is it legal to clone my own voice?

Generally yes, and platforms treat it differently from cloning someone else: YouTube needs no label when you clone your own voice for voiceovers or dubs. You still have to follow the tool’s terms, which may require a recorded consent statement.

Can I clone a celebrity’s voice for a video?

Not without permission. ElevenLabs’ policy bans cloning without consent or legal right, Tennessee’s ELVIS Act protects voices by law, and YouTube requires a label on realistic content that makes a real person seem to say something they did not.

Why does my AI voiceover sound robotic?

Usually the script is written for reading. Shorten the sentences, write numbers as words, add punctuation where you want pauses, and generate one paragraph at a time. Then direct the tone with tags or instructions.

Key takeaways
  • Fix the script first: short spoken sentences, numbers as words, no brackets.
  • Direct the delivery with audio tags or plain instructions, and generate in short sections.
  • Clone a voice only with recorded, written consent.
  • Know the laws where you publish, starting with Tennessee, California, New York and the EU.
  • Say when a voice is AI. Sometimes it is required, and it always builds trust.

Read next: make YouTube thumbnails people click, or learn who owns AI-generated content.

Sources
  1. Models, ElevenLabs documentation, September 2026
  2. Prompting Eleven v3, ElevenLabs documentation, September 2026
  3. Prohibited use policy, ElevenLabs, August 2026
  4. Text to speech, OpenAI API docs, September 2026
  5. Custom voices, OpenAI API docs, September 2026
  6. Chirp 3: Instant Custom Voice, Google Cloud documentation, September 2026
  7. Descript, Descript, September 2026
  8. Regenerate, Descript, September 2026
  9. HB 2091: Ensuring Likeness, Voice, and Image Security Act of 2024, Tennessee General Assembly
  10. Governor Newsom signs bills to protect digital likeness of performers, Office of the Governor of California, September 2024
  11. Senate Bill S8420A, New York State Senate, signed December 2025
  12. Law requiring disclosure of AI-generated synthetic performers in ads is in effect, Office of the Governor of New York, June 2026
  13. NO FAKES Act advances out of Senate Judiciary Committee, Office of Rep. Maria Salazar, June 2026
  14. AI Act: regulatory framework, European Commission, September 2026
  15. Article 3: definitions, European Commission AI Act Service Desk, September 2026
  16. Article 50: transparency obligations, European Commission AI Act Service Desk, September 2026
  17. Disclosing use of altered or synthetic content, YouTube Help, September 2026
  18. FCC makes AI-generated voices in robocalls illegal, Federal Communications Commission, February 2024
cyborb.ai

Stop reading about it. Build it.

Describe what you want in plain words. Cyborb plans the work, writes and runs the code, makes the assets, and puts the result online.

Download Cyborb

Free to start. No card required.