Natural AI Voiceover & TTS Script

Polish written drafts into clear, ear-friendly scripts for AI voiceover.

Prompt · 2 variables

I want to convert the draft manuscript below into an AI voiceover script for a {{video length}} video. The audience will only follow the content by listening while watching the screen, without rewinding.

Please rewrite this draft into an ear-friendly narration script. Keep the original meaning intact while restructuring the sentences. Keep each sentence around 10 to 15 words, expressing only one idea per sentence. Unpack any text enclosed in parentheses or special symbols into full spoken sentences, and replace complex jargon or abbreviations with plain language. Write out all numbers, dates, units, and acronyms phonetically as they should be pronounced (e.g., "40%" → "forty percent", "3/1~3/14" → "March first through March fourteenth").

Provide the output in three sections: ① Polished Script — separate paragraphs with a blank line and prepend an estimated read time to each paragraph in (00:12) format. ② Key Changes — list up to 5 major phrases changed from the original. ③ Total Estimated Read Time. (Calculate based on an average speaking pace of 2.5 to 3 words per second).

Do not add new information, unsubstantiated numbers, interjections, or promotional hype not present in the draft. Avoid ending three consecutive sentences with the identical grammatical ending. Use commas where natural breathing pauses occur, but limit them to at most two per sentence.

After rewriting, review the script by simulating a read-aloud test. Identify and fix any tongue-twisters, repetitive sounds, or awkward phrasing, and output only the finalized script. If the estimated total runtime exceeds {{video length}}, trim less critical sentences to fit the time limit.

Draft Manuscript: """ {{draft manuscript}} """

Copy, then paste here · ChatGPT and Claude open with the prompt filled in Open in ChatGPT ↗Open in Claude ↗Open in Gemini ↗ Edit in builder Download classroom card

Why it is written this way

Context
I want to convert the draft manuscript below into an AI voiceover script for a {{video length}} video. The audience will only follow the content by listening while watching the screen, without rewinding.
Task
Please rewrite this draft into an ear-friendly narration script. Keep the original meaning intact while restructuring the sentences. Keep each sentence around 10 to 15 words, expressing only one idea per sentence. Unpack any text enclosed in parentheses or special symbols into full spoken sentences, and replace complex jargon or abbreviations with plain language. Write out all numbers, dates, units, and acronyms phonetically as they should be pronounced (e.g., "40%" → "forty percent", "3/1~3/14" → "March first through March fourteenth").
Format
Provide the output in three sections: ① Polished Script — separate paragraphs with a blank line and prepend an estimated read time to each paragraph in (00:12) format. ② Key Changes — list up to 5 major phrases changed from the original. ③ Total Estimated Read Time. (Calculate based on an average speaking pace of 2.5 to 3 words per second).
Constraints
Do not add new information, unsubstantiated numbers, interjections, or promotional hype not present in the draft. Avoid ending three consecutive sentences with the identical grammatical ending. Use commas where natural breathing pauses occur, but limit them to at most two per sentence.
Self-check
After rewriting, review the script by simulating a read-aloud test. Identify and fix any tongue-twisters, repetitive sounds, or awkward phrasing, and output only the finalized script. If the estimated total runtime exceeds {{video length}}, trim less critical sentences to fit the time limit.
Input
Draft Manuscript: """ {{draft manuscript}} """

When creating an AI narration script, the most common pitfall is pasting written text directly into a text-to-speech (TTS) engine. Written text reads differently than spoken word. Long sentences with parentheses cause machines to rush through without breathing, and notations like "3/1~3/14" often get mispronounced.

The rules defined in the Task are all ear-first: short sentences, one thought per line, unpacked parentheses, and fully spelled-out numbers. Phonetically transcribing numbers and symbols eliminates pronunciation guesswork. Rewriting the text upfront is far more reliable than tweaking voice settings later.

Requesting estimated read times per paragraph in the Format ensures alignment with the target video length. Finding out a script runs 10 seconds over after rendering requires tedious video editing. Including the list of key changes lets you verify at a glance that no factual meaning shifted.

The Constraints prevent the AI from fabricating promotional fluff or ungrounded statistics. Finally, the Review step prompts the AI to simulate an audible read-through, catching tongue-twisters and awkward phrasing that look fine on paper but sound stiff when read aloud.

Unfamiliar terms? See Aha AI: output-format, self-consistency

Compared with a bad example

Common bad example

Please polish this draft to sound natural for a TTS voiceover. (Draft text here)

Simply asking to "sound natural" usually results in minor sentence trimming. Numbers, dates, parentheses, and abbreviations remain untouched, leading to robotic mispronunciations. Without duration constraints, you only discover the script is too long after recording, forcing video timeline edits.

Variations

Two-Person Dialogue Script

Two-Person Dialogue Script

I want to create a dialogue voiceover with two alternating AI voices for a {{video length}} video. Please convert the draft manuscript below into a two-person script featuring a Host and an Expert. Start a new line with the speaker's name whenever the speaker changes. Ensure neither speaker speaks for more than three consecutive sentences, and include an estimated read time for each section.

Draft Manuscript: """ {{draft manuscript}} """

Monologues can make detailed explanations feel tedious. A conversational dialogue format keeps listeners engaged and makes complex information easier to digest.

Trimming to Strict Video Length

Trimming to Strict Video Length

Please trim the draft manuscript below to fit precisely within a {{video length}} runtime. Calculate the target word count based on an average speaking pace of 130–150 words per minute, state this target word count first, and then rewrite the script to fit. Rather than densely packing sentences, cut non-essential points entirely, and list the omitted points at the end.

Draft Manuscript: """ {{draft manuscript}} """

Over-condensed sentences are harder to follow by ear. Removing low-priority sentences entirely allows the remaining script to breathe and sound natural.

Model notes

Different TTS platforms handle commas and periods with varying pause lengths, and speaking pace differs across AI voices. The average speaking speed is an estimate, so generate the first paragraph in your TTS tool to benchmark the exact pace. While some tools support SSML tags for pitch and speed, formatting syntax varies widely—adjusting pacing directly on the platform UI is usually safer than embedding tags in the script.

Related prompts

Last updated 2026-09-02 · Found a mistake? Let us know