Why it is written this way
When writing an AI video generation prompt for the first time, people often treat it like a still image description. Prompts like "a runner by a misty river" might generate a clear subject, but the output often ends up as a static frame with subtle, jittery tremors. If motion isn't explicitly defined, the AI guesses randomly, resulting in inconsistent animation.
The Task section defines the essential sequence: Subject & Action → Background Motion → Camera Movement → Lighting → Visual Texture. Weight is distributed from left to right; placing the core motion first ensures the model prioritizes fluid dynamics before styling. The requirement to dedicate "at least two phrases to explicit movement" is the critical difference between a video prompt and a still photo prompt.
Providing a concrete Example is far more effective than abstract formatting rules. The lantern example demonstrates exact phrase length, comma pacing, and kinetic phrasing. A deliberately different subject is used in the example so the AI doesn't accidentally blend the reference content into your output.
Offering 3 variations in the Format ensures that if the first direction doesn't match your vision, you have distinct baselines to compare and refine. The Constraints serve two purposes: ensuring visual rendering quality (stripping non-visual thoughts, avoiding text that video models frequently garble) and filtering out copyrighted names or brand marks.
Unfamiliar terms? See Aha AI: output-format, few-shot
Compared with a bad example
Write an English video prompt of a person running in the morning along a river
A generic request like this usually returns a flat output like "a person running by the river in the morning, cinematic." Without camera direction or specific motion cues, the resulting video will likely just drift or morph the background. When results are poor, you won't know which parameter to tweak and will end up re-rolling the same vague prompt.
Variations
When fine-tuning a single element in a near-perfect cut
I have an existing {{duration}} video prompt that is almost perfect, but I need to tweak one specific detail: "{{scene to depict}}". Do not rewrite the prompt from scratch. Modify only the necessary phrase, and provide a one-line explanation of what was changed and why. Keep all other phrases and their original sequence exactly intact.
Regenerating the entire prompt alters the parts you already liked. Isolating the specific phrase allows controlled, single-variable adjustments.
When generating an extreme close-up shot
I want to create an extreme close-up single shot with an {{overall atmosphere}} mood based on this scene: "{{scene to depict}}". Because prominent facial features easily distort during dynamic motion, keep the face and hands mostly stationary. Focus the motion prompts on subtle secondary elements like drifting hair strands, fluttering collar fabric, or background bokeh. Limit camera movement to a single, very slow crawl.
Close-up shots degrade rapidly with heavy facial movement. Restricting motion to peripheral details prevents facial warping and preserves generation fidelity.
Model notes
Every video generation platform (Runway, Sora, Kling, Pika, Luma) supports different shot durations and aspect ratios, typically managed via platform sliders or UI settings rather than inside the prompt text. Transfer the final shot duration to your tool's dedicated duration setting. Since motion sensitivity varies across models, adjust your motion phrases first whenever porting a prompt from one AI platform to another.
Related prompts
Last updated 2026-09-02 · Found a mistake? Let us know