Skip to content
Mango
Launch Mango

How to Write AI Video Prompts That Actually Work (With Templates)

Content creator planning AI video strategy at a desk with notes and laptop

The gap between a great AI video and a forgettable one usually isn't the tool — it's the prompt. Two people using the same model can produce wildly different results based entirely on how they describe what they want. Most creators are leaving enormous quality on the table by writing prompts the way they'd type a search query.

AI video prompts are a skill. A learnable, improvable one. This guide gives you a six-part formula, platform-specific templates, and real before-and-after examples to move from inconsistent outputs to a reliable system that produces scroll-stopping content on demand.

Why Most AI Video Prompts Fall Flat#

Most people write prompts the way they Google things: brief, keyword-forward, stripped of context.

"A person running through a forest." "Coffee being poured." "A futuristic city."

These prompts produce technically correct results that are completely forgettable. The model fills every unstated gap with statistical averages from its training data. You get average lighting, average composition, average pacing — because you didn't specify anything to pull it away from average.

The core problem is specificity debt. Every detail you leave out is a creative decision the model makes for you, defaulting to the most common version of whatever you described. For a tool that can generate genuinely cinematic content, that's an enormous amount of creative control you're surrendering by default.

The other common failure mode is structural chaos — mixing aesthetic qualities, camera directions, and emotional tones in one unstructured sentence. "A dramatic slow-motion shot of a waterfall with beautiful lighting and cinematic vibes at sunset" has several useful elements, but they're all competing for attention without a clear hierarchy. The fix is a structured formula that feeds each type of information to the model in a deliberate order.

The Core Prompt Formula#

Every high-performing AI video prompt breaks into six components:

[Subject] + [Action] + [Setting] + [Lighting/Mood] + [Camera] + [Style]

You don't need all six in every prompt, but understanding what each one does — and what the model defaults to when you omit it — lets you make deliberate choices.

  • Subject: What or who is in the video. Specify appearance, size, color, and distinguishing features.
  • Action: What is happening. Use precise, physical verbs — not "moves" but "rotates," not "falls" but "cascades."
  • Setting: Where the action takes place. Include location, time of day, season, and relevant environmental details.
  • Lighting/Mood: The emotional quality of the light and the resulting atmosphere.
  • Camera: The lens choice, angle, distance, and movement.
  • Style: The aesthetic reference — film stock, color grade, visual era, or a recognized visual style.

When you specify all six, you leave almost nothing to chance. The model executes your vision instead of inventing one.

Before and After#

Weak prompt: "A cup of coffee in a cozy setting"

Strong prompt: "A handmade ceramic white mug of steaming black coffee on a worn oak table, morning condensation forming on the exterior of the mug, warm amber light streaming through gauze curtains from the left, slow rack focus from the rising steam to the mug surface, shallow depth of field, 35mm film aesthetic, muted warm color grade"

Same subject, completely different output. The strong prompt specifies the container, the liquid, the surface, the atmospheric detail, the lighting direction and quality, the camera move, the depth of field, the film stock, and the grade. The model executes a vision rather than invents one.

Subject and Action: How to Describe What Happens#

Your subject and action descriptions are the foundation. Get them vague and nothing else can save the output.

Describing Subjects#

Specificity works at every scale:

  • Not "a woman" — but "a woman in her 30s wearing a cream linen blazer, light brown hair pulled back loosely, standing with her hands in her pockets"
  • Not "a product" — but "a 100ml amber glass serum bottle with a brushed gold dropper cap"
  • Not "a cityscape" — but "a dense block of Manhattan brownstones at street level, fire escapes on every building, a narrow one-way street flanked by parked cars"

For product content, describe the exact colors, materials, and defining visual features. For people, focus on what they're wearing and their body language rather than their expression, which models tend to over-interpret. For environments, include architectural details, surface textures, and what makes the space visually distinctive.

Describing Action#

The most common prompt failure with action is using static or vague verbs.

  • Static: "a person near a waterfall"
  • Active: "a person slowly raising their face toward the mist of a waterfall, eyes closed, water droplets catching late afternoon light on her cheekbones"

Use physical, directional verbs: cascades, drips, unfolds, rotates, spirals, blooms, expands, shatters. Pair them with adverbs that describe pace — slowly, imperceptibly, rapidly — and with specific physical outcomes. For action sequences, include cause and effect: "a glass dropped onto marble, shattering in slow motion, shards scattering outward in a radial pattern and catching overhead light."

Setting, Lighting, and Mood: The Cinematic Layer#

AI video creator reviewing footage on a screen

Setting, lighting, and mood are where most intermediate prompts plateau. Creators get good at subjects and actions, then fall back on "beautiful lighting" and "cinematic feel" — phrases so generic they've lost all signal.

Setting Specificity#

Go beyond naming a location. Include:

  • Time of day: "golden hour" (hard, directional, warm), "blue hour" (soft, cool, diffused), "2 AM city lights" (artificial, high-contrast shadows)
  • Weather and atmosphere: morning fog, humid haze, crisp clear winter air, overcast diffusion
  • Scale indicators: something in frame that establishes how big or small the space feels
  • Texture and material: polished concrete, rough brick, velvet upholstery, burnished copper

The difference between "a forest" and "a Pacific Northwest old-growth forest in late November, moss-covered Douglas firs, diffused overcast light filtering through morning fog, damp forest floor with decomposing leaves and exposed roots" is the difference between a generic stock clip and something that belongs in a prestige nature documentary.

Lighting Direction and Quality#

Light has four properties that matter for prompting: direction (where it comes from), quality (hard or soft), color temperature (warm, cool, or neutral), and intensity (bright, dim, dappled).

Specify all four when you can:

  • "hard directional sunlight from the upper right, casting sharp shadows to the left, warm 5600K"
  • "diffused window light from directly behind the subject, soft fill on the face, neutral 4200K"
  • "single practical lamp in a dark room, pools of amber light, deep shadows filling most of the frame"
  • "three-point studio setup: key from camera left at 45°, fill from the right at half power, rim light directly behind"

The more specific your lighting description, the more controlled the emotional tone of your output. Lighting is mood.

Mood Without Clichés#

Avoid: "cinematic," "beautiful," "dramatic," "epic." These adjectives appear in every type of training image and no longer discriminate between outputs.

Replace them with specific emotional effects:

  • Instead of "dramatic" → "a sense of unease, as if something is about to happen"
  • Instead of "beautiful" → "warm and inviting, the visual equivalent of a Sunday morning with nowhere to be"
  • Instead of "epic" → "scale emphasizing human smallness against an indifferent natural landscape"

These descriptions give the model an emotional target to hit rather than an overloaded aesthetic label.

Camera Movement and Style: Directing the Shot#

Camera direction is the most underused part of AI video prompting. Most creators describe what they want to see without describing how the camera sees it — and the model defaults to static, frontal, medium shots.

Essential Camera Vocabulary#

Distance: extreme close-up (ECU), close-up (CU), medium shot (MS), wide shot (WS), extreme wide shot (EWS)

Angle: eye level, low angle (shooting upward, subjects appear powerful), high angle (shooting downward, subjects appear small or vulnerable), Dutch angle (tilted frame, creates unease), bird's eye view

Movement: dolly in (camera moves toward subject), dolly out, tracking shot (moves with subject), pan (horizontal rotation), tilt (vertical rotation), handheld (subtle shake), Steadicam (smooth through-space movement), rack focus (refocusing within the shot)

Lens characteristics: shallow depth of field (background blur), deep focus (everything sharp), wide angle (distortion, spaciousness), telephoto (compression, intimacy)

For TikTok content, the most effective camera specs are: handheld with subtle shake (feels authentic and native), extreme close-up (maximizes impact in a small vertical frame), low angle upward (makes products feel aspirational), and slow dolly in (builds tension and visual interest).

Style References#

Style references shortcut a huge amount of description. Instead of specifying every visual quality, reference a well-understood aesthetic:

  • "35mm Kodak Portra 400" — warm tones, soft grain, gentle halation, slightly overexposed
  • "iPhone portrait mode vertical" — the native look of modern social content
  • "Wes Anderson symmetry" — centered framing, saturated pastels, precise geometric composition
  • "magic hour handheld" — impressionistic golden-hour light, natural textures, subtle motion
  • "neon night, Wong Kar-wai" — warm artificial light, slight overexposure, motion blur in shadows
  • "clean editorial white" — high-key lighting, no shadows, clinical fashion magazine aesthetic

Pair a style reference with your specifics for the most efficient prompts. The reference handles aesthetic defaults; your specifics handle subject, action, and setting.

Platform-Specific Prompt Templates#

Different platforms reward different visual languages. Here's how to apply the formula for each.

TikTok#

TikTok rewards native, sensory, and rewatchable content. The algorithm optimizes for completion rate, so every second needs to earn the next.

Template: [Extreme close-up of / POV] [specific subject] [specific action with physical texture], [sensory material detail], [time of day with light direction], [natural imperfection — steam, condensation, dust], handheld with subtle shake, [warm vertical mobile aesthetic], 9:16 vertical

Example: "Extreme close-up of a dark chocolate bar snapping slowly in two, steam rising from just-melted ganache between the layers, shards caught mid-scatter in slow motion, afternoon window light catching the surface texture, chocolate dust floating in the air, handheld with subtle shake, iPhone cinematic mode, 9:16 vertical"

Instagram Reels#

Reels rewards aspirational aesthetics with native polish — visually elevated but still feeling like something a real person could have filmed.

Template: [Subject in aspirational lifestyle context] [action conveying the lifestyle], [elevated setting with architectural detail], [golden hour or soft morning light], [smooth Steadicam movement], [editorial lifestyle aesthetic], warm color grade

Example: "A woman in her early 30s in a cream cashmere sweater pouring matcha from a ceramic teapot at a marble kitchen island, early morning golden light through tall windows, herbs growing in white ceramic pots on the windowsill, smooth Steadicam push-in, Kinfolk editorial aesthetic, warm muted color grade"

YouTube Shorts#

Shorts rewards educational clarity and curiosity hooks — content where the visual complements information rather than existing for its own sake.

Template: [Overhead or over-the-shoulder view of] [process or concept demonstration], [clean uncluttered setting], [neutral soft lighting], [medium or close-up distance], [flat lay or POV angle], clean minimal aesthetic, no motion blur

Example: "Overhead flat lay of hands writing in a journal on a white desk, bullet points filling the page in real time, natural overcast window light from the left, medium-distance top-down camera, neutral white-and-warm-wood color palette, clean minimal aesthetic, sharp focus throughout"

How to Iterate: From First Draft to Winning Prompt#

Your first prompt isn't your final prompt. The creators behind consistently strong social media video content treat AI video generation as an iterative loop, not a one-shot output.

The Iteration Framework#

Round 1: Write a complete six-part prompt. Generate 3–5 variations.

Evaluate: What worked? What fell flat? Identify the specific elements that diverged from your intent.

Round 2: Keep what worked. Explicitly correct what didn't:

  • Background too busy → add "simple, uncluttered background, minimal environmental detail"
  • Lighting too flat → add "single directional light source from [angle], creating defined shadow on the opposite side"
  • Motion too fast → add "imperceptibly slow, time appears to nearly stop"
  • Subject looks artificial → add "photorealistic, shot on camera, not 3D rendered, no AI aesthetic artifacts"

Round 3: Introduce micro-variations to stress-test the winning elements:

  • Swap the setting while keeping the lighting and camera
  • Swap the camera angle while keeping the subject and action
  • Swap the time of day while keeping everything else

This diagonal iteration approach separates creators with one-time viral hits from creators who produce strong content consistently.

Saving and Systematizing#

When a prompt produces a strong output, save the full prompt with a note on what worked. Organize your saved prompts by category (product, lifestyle, educational, abstract), by platform, and by performance. Within a few weeks, you'll have a prompt library that reflects your specific aesthetic — and new prompts become fast because you're mixing proven components rather than building from scratch.

A bank of 20 proven base prompts with documented modifiers compounds in value every week. You're not reinventing the wheel; you're assembling it from parts that you already know work.

The quality ceiling for AI video is much higher than most creators are reaching, and it's determined almost entirely by how precisely you describe what you want. Better prompts unlock better outputs from the exact same tools everyone else is using. If you want to put this framework to work at scale — producing dozens of variations across formats and platforms — Mango is built for exactly that kind of high-volume AI video production.

Turn this into output

Make Mango run this playbook for you

Give Mango your site and goals. It builds the strategy, drafts the posts, creates the videos, and keeps the weekly marketing work moving.