Precise Camera Movements for AI Video

Command zooms, pans, and tracking shots using exact cinematography terms.

Prompt · 2 variables

I need exact camera movement phrases in English for an AI video generation prompt. The scene is "{{scene}}", and the desired visual tone is {{desired_mood}}.

Translate this visual mood into professional cinematography terms. Instead of vague words like "slowly" or "smoothly," provide English phrases specifying camera height, distance, direction, and magnitude of movement. Restrict each shot to a single distinct movement, clearly stating the start position and end position.

Below is an example of the desired format. The subject is just an illustration, so follow the structure rather than copying the words: Example: static wide shot at knee height, then a slow dolly-in toward the bench, ending on a chest-height medium shot — starts at knee-level wide framing, slowly pushes in toward the bench, and stops at a chest-height medium framing.

Provide 3 distinct movement options. Format each option across three lines: ① English camera direction phrase ② A plain-English explanation of the actual physical movement ③ The emotional or visual impression this movement conveys to the viewer. After listing all 3 options, recommend the one that best matches {{desired_mood}} with a one-line justification.

Do not combine multiple movements in a single shot (no simultaneous zooming and rotating). Do not include impossible physical paths (such as moving through solid walls or clipping through subjects). Do not include real people, celebrities, or brand logos. Do not include tool-specific parameter flags or numerical weight syntax.

Copy, then paste here · ChatGPT and Claude open with the prompt filled in Open in ChatGPT ↗Open in Claude ↗Open in Gemini ↗ Edit in builder Download classroom card

Why it is written this way

Context
I need exact camera movement phrases in English for an AI video generation prompt. The scene is "{{scene}}", and the desired visual tone is {{desired_mood}}.
Task
Translate this visual mood into professional cinematography terms. Instead of vague words like "slowly" or "smoothly," provide English phrases specifying camera height, distance, direction, and magnitude of movement. Restrict each shot to a single distinct movement, clearly stating the start position and end position.
Example
Below is an example of the desired format. The subject is just an illustration, so follow the structure rather than copying the words: Example: static wide shot at knee height, then a slow dolly-in toward the bench, ending on a chest-height medium shot — starts at knee-level wide framing, slowly pushes in toward the bench, and stops at a chest-height medium framing.
Format
Provide 3 distinct movement options. Format each option across three lines: ① English camera direction phrase ② A plain-English explanation of the actual physical movement ③ The emotional or visual impression this movement conveys to the viewer. After listing all 3 options, recommend the one that best matches {{desired_mood}} with a one-line justification.
Constraints
Do not combine multiple movements in a single shot (no simultaneous zooming and rotating). Do not include impossible physical paths (such as moving through solid walls or clipping through subjects). Do not include real people, celebrities, or brand logos. Do not include tool-specific parameter flags or numerical weight syntax.

When directing AI video camera movement, people often write vague prompts like "move the camera slowly." However, video models interpret "slowly" differently every time—one generation might produce subtle jitter, while another drastically warps the entire scene. The larger problem is that when you get a great result, you cannot reliably recreate it.

In the Task paragraph, subjective impressions are converted into technical cinematography terminology. Established terms like dolly, pan, tilt, and tracking have standardized definitions that both AI models and human editors interpret consistently. Specifying camera height, starting positions, and ending positions ensures consistent results across multiple generation passes—a crucial requirement when stitching sequential shots together.

Providing an Example format ensures the output mirrors the structured pairing of English technical phrasing and clear functional explanations. Knowing the exact terminology helps you independently build precise camera movements for future prompts.

Requesting the visual impression within the Format develops an intuition for selecting camera movements. Finally, the Constraints prevent compound camera movements. Combining moves like simultaneous zoom and rotation frequently leads to artifact smearing or melting visual geometry.

Unfamiliar terms? See Aha AI: output-format, few-shot

Compared with a bad example

Common bad example

Cafe window scene, make a video prompt with slow camera movement and cinematic vibe

A generic prompt like this usually generates phrases like "cinematic slow camera movement." Because it lacks direction, height, and depth specifications, every generation yields wildly different motions. When cutting multiple shots together, fluctuating camera levels will break scene continuity.

Variations

Establishing Wide Landscapes

Establishing Wide Landscapes

I want to create a wide aerial shot establishing the scene: "{{scene}}". Style it as a drone or crane shot with specific altitude and flight trajectory in English. Clearly specify the start altitude, end altitude, and camera pointing direction, followed by a brief line explaining the physical trajectory.

High-angle shots become static sky backdrops without explicit altitude and direction constraints. Enforcing starting and ending elevations maintains clear spatial motion.

Standardizing Motion Across Multiple Shots

Standardizing Motion Across Multiple Shots

I am planning to edit together multiple shots within the same location. The reference master shot is "{{scene}}". Create a standardized camera movement rule that maintains visual continuity when cut together. Provide a one-sentence English phrase detailing camera height, direction, and speed that can be appended to prompt descriptions for consecutive shots.

Inconsistent camera framing across sequential cuts breaks the illusion of a continuous scene. Establishing a reusable directional phrase ensures visual consistency.

Model notes

Different AI video engines recognize varying cinematography vocabularies. Common terms (dolly, pan, tilt, tracking, static) work reliably across most platforms, whereas obscure terms may be ignored; if a model misinterprets a term, describe the physical motion in plain English. For tools featuring dedicated motion sliders or parameter values, avoid repeating intensity descriptors in the prompt text to prevent extreme jitter or warping.

Related prompts

Last updated 2026-09-02 · Found a mistake? Let us know