Why it is written this way
The most common mistake when making AI short-form videos is simply asking: "Make a 30-second video." Current generative AI models are optimized to create clips that last only a few seconds. Asking for a full-length video at once leads to disjointed scenes or a video where only the first few seconds look coherent. Breaking the video down into achievable units solves half the battle.
The Task section separates narrative planning from prompt writing. By defining what each shot shows first, the storytelling remains clear and structured before visual descriptors are added. Trying to generate prompts directly without outlining the narrative leads to flowery English prompts with a chaotic story sequence.
The list of consistent elements in the Format section is what makes the stitched clips feel like a single cohesive video. Because video generators generate each clip from scratch, character clothing, hair, and lighting will shift unpredictably between cuts. Repeating identical descriptive anchors across every prompt remains the most reliable way to maintain visual continuity.
Finally, the Review step prompts the AI to audit its own output before finalizing. It catches continuity errors and verifies that the total runtime matches the target length, saving you from having to restart the generation process later.
Unfamiliar terms? See Aha AI: output-format, hallucination
Compared with a bad example
Write a prompt for a 30-second Shorts video about making breakfast in a studio apartment.
A prompt like this yields a single, long descriptive paragraph without shot boundaries. Feeding that into a video tool results in a 3-to-5-second clip that captures only the beginning. Re-generating creates a completely different person and kitchen each time, making it impossible to edit them into a seamless video.
Variations
When you already have a voiceover script
I already have a script for a {{total length}} short-form video on the topic: "{{video topic}}". Keeping the voiceover lines in their exact order, break down only the visuals into shots of around 5 seconds each. Format the table as: "Shot # | Duration (sec) | Spoken Line | Visual Prompt (English)", and flag any sections where the spoken line is too long to fit into 5 seconds.
When a script exists, shot lengths must match speech cadence. This variant anchors cuts to the voiceover timing rather than arbitrary visual breaks.
When regenerating only the hook (first shot)
I want to redesign only Shot 1 from the shot list created earlier. Keeping the {{overall atmosphere}} vibe intact, give me 3 alternative hook ideas designed to stop users from scrolling within the first 3 seconds. For each option, provide a one-line description and a full English prompt, maintaining all recurring character/setting elements so it connects seamlessly with the following shots.
Short-form success depends heavily on the opening hook. This variant lets you explore multiple high-retention openings without altering the rest of the sequence.
Model notes
Clip generation limits vary across AI video tools (e.g., Runway, Pika, Kling, Sora). Adjust the shot durations in the table to match your specific tool's native clip length. If your tool supports using the last frame of a clip as the starting image (Image-to-Video initiation) for the next shot, combine that feature with the consistent elements list in this prompt for the highest continuity.
Related prompts
Last updated 2026-09-02 · Found a mistake? Let us know