AI video is where image prompting grows a fourth dimension: time. Everything you learned in our AI Image Prompting guide still applies — but now every prompt also needs motion: what the subject does, and what the camera does. Master those two verbs and you can direct any video model like a filmmaker.
This guide is tool-agnostic by design — the same principles drive every major video generator (Sora-class, Runway-class, Veo-class, and the chat-integrated ones), and tools change monthly while the craft doesn’t. Below: the video formula, the camera language bank, the one-action rule that fixes most failures, the image-to-video anchor trick, recipes, and the honest problem list. Guide #25 of the Prompt Engineering roadmap.
Who Uses AI Video for What — Find Your Lane
What Changes When Images Start Moving
- Motion becomes the main slot. A video prompt without motion words produces an awkward semi-still. Two motion layers exist: subject motion (what happens in frame) and camera motion (how we view it) — great prompts specify both.
- Clips are short. Most tools generate seconds, not minutes — so you prompt moments, not stories. Longer videos are chains of prompted moments (storyboarding, below).
- Consistency is the hard part. Faces drift, objects morph, physics wobbles. The craft is choosing shots that hide weaknesses and anchoring with reference frames.
Quick Start — Your First 3 Video Prompts
Prompt 1 moves the world, prompt 2 moves the viewer, prompt 3 does both gently. That’s 80% of video prompting in three generations.
The AI Video Formula
Note the order: camera first. Video models weight early words heavily (same as image tools), and the camera instruction is the one most often ignored when buried at the end.
The Camera Language Bank (Steal These)
The Motion Word Bank (Subject Side)
- Gentle (AI’s sweet spot): drifting · swaying · rising steam · flickering candlelight · hair in a breeze · rippling water · falling leaves/snow · slow rotation
- Human-safe motions: turning toward camera · walking away · hands working (pottery, typing, kneading) · sipping · a slow smile — simple, continuous, no fine finger choreography
- Energy (use carefully): splashing · sparks · billowing smoke · flags whipping — dramatic but artifact-prone; keep clips short
- Speed modifiers: “slow, gentle, continuous” = cinematic and stable · “fast, sudden” = energetic and risky
Pair one subject motion with one camera move and you’ve written a professional video prompt. That pairing discipline is the whole craft.
The One-Action Rule
Clips are moments. One subject, one action, one camera move per generation — then chain clips for sequences. This single rule fixes more failed video prompts than everything else combined.
The Image-to-Video Anchor Trick
Most tools accept a starting image — and pros use it constantly, because it solves the consistency problem at the source:
- Generate the perfect frame first in your image tool (full control, cheap iterations — the 7-slot formula), then animate it: “gentle camera push-in, steam rising, hair moving in the breeze”
- Brand consistency: the same anchored character/product across many clips beats re-describing and praying
- The motion prompt shrinks: with the image carrying the look, your text only needs to describe MOTION — which is exactly what video models are best at following
8 Copy-Paste Video Recipes
Watch One Rewrite Fix Everything
Same idea, different discipline. The second prompt gives the model a director’s shot card instead of a vibe — and the result looks like it cost money.
Storyboarding — From Clips to a Video
- Write the sequence first: ask your text AI — “break this 30-second product video into 6 five-second shots: shot type, camera move, action, and mood for each” (ChatGPT/Claude are excellent shot-list writers)
- Keep a consistency block: reuse the same character/setting/style description verbatim in every shot prompt — copy-paste, don’t re-type
- Vary shot types like an editor: wide → medium → close-up reads as professional; six medium shots reads as AI
- Assemble in any editor (CapCut and friends) with music and cuts — AI generates the shots; editing makes the film
Universal Problems & Fixes
The Tool Landscape (Without the Hype)
- Chat-integrated video (inside AI assistants): easiest entry, conversational iteration — start here if you’re new
- Dedicated generators (Sora-class, Runway-class, Veo-class, and peers): stronger control, longer/cleaner output, more settings — where serious work happens
- The honest truth: leaderboards reshuffle monthly; the formula, camera language, and anchor trick transfer to whichever tool wins this quarter. Learn the craft, rent the tool.
Ethics — The Part That Matters More in Video
- Deepfake line: never generate real, identifiable people saying or doing things they didn’t — in many places this is now illegal, everywhere it’s wrong
- Disclose AI video where viewers could reasonably be deceived — platforms increasingly require labels
- Rights follow the same rules as images: check your tool’s commercial terms; avoid trademarked characters and living artists’ signature styles for paid work
Mistakes That Waste Generations
- Prompting a story into one clip — the one-action rule exists because everyone breaks it first
- No camera language — you’re surrendering the director’s chair; the model picks a random move
- Fast motion requests — speed amplifies every artifact; slow is where AI video looks expensive
- Skipping the image anchor for anything needing consistency — the single biggest pro/amateur divider
- Judging by one generation — video has more randomness than images; 2–3 runs of a good prompt beats 1 run of a perfect one (familiar idea?)
Frequently Asked Questions
Everything from image prompting applies, plus motion: what the subject does and what the camera does. Clips are seconds long, so you prompt single moments — one subject, one action, one camera move — and chain clips for longer videos.
Camera move + shot type first, then subject with a defining detail, ONE clear action, environment, lighting and mood, then style. Early words carry more weight, and camera instructions get ignored when buried.
Too much asked of one clip. Shorten the moment, reduce to one action, slow the motion, and anchor with a starting image — morphing is the model improvising when your prompt demands more change than the clip length allows.
You supply a starting frame (usually AI-generated with full control) and prompt only the motion. It locks the look, solves character/product consistency, and plays to what video models do best — animating, not inventing.
Storyboard: break the video into 5-second shots (your text AI writes great shot lists), generate each with a reused consistency block, and assemble with cuts and music in a normal video editor.
‘Slow.’ Slow camera moves and slow subject motion hide artifacts, read as cinematic, and give the model room to stay coherent.
Yes — camera language, the one-action rule, and image anchoring are universal. Tools differ in duration, resolution, and features, but the prompting craft transfers wholesale.
Depends on your tool and plan — check current terms. Add the video-specific cautions: no real identifiable people, disclose where deception is possible, and follow platform AI-labeling rules.
Direct the camera. One moment at a time. 🎬
Next: System Prompts Explained — the Tool track finale.




