why-ai-video-doesnt-look-right

The Question Everyone’s Asking

Earlier today, a user named Evelyn posted this on a creator forum:

“I’ve been diving into AI video lately. I have ideas, but when I write prompts and feed them to the AI, the result feels… off. Not wrong exactly, but not what I envisioned. The depth of field and composition don’t match my expectations. When I try film-inspired sequences, the shots feel unstable. My current workflow is: GPT writes the story → break it into shots → generate prompts + first/last frames → feed into Seedance 2.0.”

Evelyn’s experience is not unique. It’s one of the most common frustrations in AI video generation: the gap between what you imagine and what the model produces. But here’s the thing — that gap isn’t a model limitation. It’s a communication problem. And it’s fixable.

If you’re using Seedance 2.0, VideCraft, or any other AI video tool and feeling the same way, this guide is for you.


Why Your Prompts Are Failing You

Most people approach AI video prompts the same way they approach image prompts: describe the subject, add some style words, hit generate. That works for images. For video, it falls apart.

Video is not a single frame. Video is motion, time, space, and light unfolding over seconds. Your prompt needs to think like a cinematographer, not a photographer.

Here are the three root causes behind Evelyn’s problems — and probably yours too:

1. Depth of Field & Composition: You’re Not Speaking Camera

When Evelyn says “the depth of field doesn’t match my expectations,” she’s describing a prompt that lacks lens and camera language.

AI video models understand cinematography terminology. They’ve been trained on film, commercial, and broadcast content. If your prompt says “a person walking in a city,” the model guesses. If your prompt says “a 35mm lens at f/2.8, shallow depth of field with the subject sharp against a bokeh city background, rule-of-thirds composition with the subject on the left third line,” the model knows exactly what you mean.

The fix: Learn and use these terms in every prompt:

What You Want The Prompt Language
Blurry background, sharp subject Shallow depth of field, f/2.8, bokeh background
Everything in focus Deep focus, f/11, wide-angle establishing shot
Subject on the left Rule-of-thirds composition, subject on left third line
Looking up at subject Low-angle shot, heroic perspective
Looking down at subject High-angle shot, overhead perspective
Smooth, dreamy look Anamorphic lens, horizontal lens flares, soft edge falloff
Sharp, gritty look Spherical lens, deep focus, high contrast, film grain

2. Shot Instability: You’re Not Controlling the Camera

“Camera instability” in AI video almost always comes from one mistake: the model is deciding your camera movement for you.

If you don’t specify camera movement, the model invents it — and it usually invents handheld shake or erratic dolly moves because that’s what looks “dynamic” in its training data.

The fix: Always specify camera motion explicitly. Here’s the vocabulary:

  • Static / Locked-off: “Camera remains stationary on a tripod, no movement whatsoever, locked-off shot”
  • Smooth dolly: “Slow, smooth dolly-in toward the subject, Steadicam-like stability”
  • Tracking: “Camera tracks alongside the subject, lateral movement, gimbal-stabilized”
  • Orbit: “Camera orbits slowly around the subject, 180-degree arc, fluid motion”
  • Drone: “Aerial drone shot, smooth ascent, gimbal-stabilized, no sudden movements”
  • Handheld (intentional): “Slight handheld breathing, documentary-style, natural micro-jitter”

If Evelyn’s film-inspired sequences feel unstable, her prompt probably says something like “cinematic footage” without specifying that she wants “locked-off tripod shot, zero camera movement, static composition.”

3. The GPT Workflow Gap

Evelyn’s workflow is solid conceptually: GPT → shot breakdown → prompts → Seedance 2.0. But there’s a missing step.

GPT writes great stories. It writes mediocre camera directions. When you ask GPT to “break this into shots,” it produces narrative shot descriptions — “The hero walks through the door, looking determined” — not cinematography instructions. Seedance 2.0 needs the latter.

The fix: Add a translation layer. After GPT produces the shot breakdown, either:

  1. Write the camera instructions yourself using the terminology above, or
  2. Give GPT a cinematography system prompt before the shot-breakdown step — something like:

“You are a professional cinematographer. For each shot, specify: focal length, aperture, camera movement (or lack thereof), lighting setup, composition rule, depth of field, and aspect ratio. Use professional film terminology. Do not use vague words like ‘cinematic’ or ‘beautiful’ without technical justification.”


The Complete Cinematography Prompt Template

Here’s a template you can use for every shot. Fill in each section, and your AI video results will improve dramatically:

1
2
3
4
5
6
7
8
9
[SHOT TYPE]: [Establishing wide / medium / close-up / extreme close-up / insert]
[LENS]: [Focal length in mm], [aperture], [lens type — anamorphic / spherical / macro]
[CAMERA MOVEMENT]: [Static locked-off / slow dolly in / tracking left-to-right / orbit / handheld breathing]
[DEPTH OF FIELD]: [Shallow f/2.8 with bokeh / deep focus f/11 everything sharp / rack focus from foreground to background]
[COMPOSITION]: [Rule of thirds subject on left / center-framed symmetrical / leading lines / Dutch angle]
[LIGHTING]: [Key light source and direction / fill light / rim light / color temperature in Kelvin / practical lights in scene]
[COLOR]: [Color grading reference — warm amber and teal / desaturated with cool shadows / vibrant saturated / film stock emulation]
[ASPECT RATIO]: [16:9 / 9:16 vertical / 2.35:1 cinemascope / 1:1 square]
[DURATION]: [N seconds]

Real Example: Before and After

Before (what most people write):

“A product commercial showing a smartwatch, dramatic lighting, cinematic style, 16:9.”

After (what gets results):

“Static locked-off tripod shot (5 seconds). A premium smartwatch centered on a dark gradient surface. 85mm lens at f/2.0, anamorphic, shallow depth of field — the watch face is razor-sharp, the strap falls off into smooth bokeh. A single rim light at 45 degrees camera-left traces the titanium bezel with a warm golden edge (3200K). Subtle particles of light drift across the frame in slow motion. Rule-of-thirds composition with the watch face on the upper-right third intersection. Color graded in warm amber tones with crushed blacks. Zero camera movement. 1080P, 16:9.”

The second prompt is longer, yes. But it also leaves nothing to chance. Every creative decision is on the page.


Mastering First/Last Frame Generation

Evelyn mentions using first/last frames — this is actually one of the most powerful techniques in Seedance 2.0, but it requires careful setup.

First/last frame generation (keyframe interpolation) works by giving the model two images: where the shot starts and where it ends. The AI fills in the motion between them.

Common mistakes:

  1. Frames are too different. If your first frame is a seed in soil and your last frame is a fully bloomed flower, the model has to invent 30 days of growth in 8 seconds. That’s motion blur, morphing artifacts, and uncanny movement. Fix: Narrow the gap. First frame = seedling breaking soil. Last frame = the same seedling with two small leaves. Let each generation handle a small, coherent motion step.

  2. Lighting doesn’t match between frames. If the first frame has golden-hour warmth and the last frame has midday cool light, the interpolation will flicker. Fix: Match lighting conditions across both keyframes. Same time of day, same weather, same light direction.

  3. No motion guidance in the prompt. First/last frames provide the visual anchors, but the prompt still needs to describe what kind of motion connects them. Fix: Add motion direction — “smooth growth animation, stem elongating upward, leaves unfurling outward from center, no morphing artifacts, organic time-lapse feel.”

When to Use First/Last Frames vs. Single-Frame Generation

Scenario Best Approach
Smooth transformation (seed → flower, day → night) First/last frame keyframe interpolation
Action shot with unpredictable motion Single frame, detailed motion prompt
Camera movement around a static subject Single frame + explicit camera motion in prompt
Subject entering and exiting frame First/last frame with subject position markers
Consistent character across multiple shots Single frame, consistent character description across prompts

Building a Pro Workflow

Evelyn’s five-step workflow is a great foundation. Here’s how to level it up:

Step 1: Story Confirmation (GPT)

Don’t just confirm the story — extract the visual language. What’s the color palette? What’s the dominant lighting style? What’s the editing rhythm? Give GPT a cinematographer persona.

Step 2: Shot Breakdown (GPT + Cinematographer Prompt)

For each shot, extract:

  • Shot duration (in seconds)
  • Shot type and lens choice
  • Camera movement (or static)
  • Key light placement and quality
  • Required props, wardrobe, and set dressing
  • Transition into the next shot

Step 3: Prompt Generation (Manual Refinement)

GPT can draft the prompts, but you need to refine them. GPT uses words like “cinematic,” “stunning,” and “beautiful” as filler. Replace every vague adjective with a technical instruction. Every time.

Step 4: Generate First/Last Frames (Seedream or Midjourney)

Generate your keyframes with an image model first. Make sure the lighting, composition, and aspect ratio match between frames. This is where you catch problems before they cost you video generation credits.

Step 5: Generate Video (Seedance 2.0 / VideCraft)

Run your polished prompts through Seedance 2.0. For each generation, note what worked and what didn’t. AI video is iterative — expect 2-3 generations per shot to get it right, and adjust your prompt each time.

Pro tip: Use the random seed parameter to lock in a seed once you get a result that’s 80% there. Then tweak only the parts of the prompt that need fixing. This keeps the overall look consistent while letting you iterate on specific issues.


The VideCraft Advantage

At Seed Audio Prompts, we built VideCraft — powered by Seedance 2.0 — to give creators direct control over these variables. Here’s what makes it different from a raw API call:

  • Resolution control: 480P through 4K, so you’re not wasting credits on test generations
  • FPS and duration settings: Control the temporal quality explicitly
  • Random seed: Reproducibility for iterative refinement
  • Multi-modal input: Start from text, an image, a video clip, or even an audio track
  • Last-frame return: Chain multiple generations seamlessly — the last frame of gen 1 becomes the first frame of gen 2
  • Prompt generator: Upload a reference video and our tool reverse-engineers a cinematography-grade prompt

And if you’re working across media, our platform integrates video (VideCraft), audio (Seedaudio), images (Seedream), and 3D (Seed3D) under one roof — so your video’s soundtrack and visual effects all come from the same workflow.


Quick Checklist: Before You Hit Generate

Run through this list for every shot. If you can answer all eight questions, you’re ready:

  • Shot type? (wide, medium, close-up, insert?)
  • Lens? (focal length, aperture, lens character?)
  • Camera movement? (static or moving? if moving, how exactly?)
  • Depth of field? (shallow with bokeh, deep focus, rack focus?)
  • Composition? (rule of thirds, centered, leading lines, symmetry?)
  • Lighting? (key light direction, quality, color temperature, practicals?)
  • Color grade? (warm/cool, saturation, film stock reference?)
  • Duration? (how many seconds exactly?)

The Bottom Line

AI video isn’t magic. It’s a tool — and like any tool, the quality of the output depends on the precision of the input.

Evelyn’s workflow is 80% of the way there. The missing 20% is the cinematography translation layer: turning narrative shot descriptions into technical camera instructions that Seedance 2.0 can execute with precision. Once you add that layer, the depth of field lands where you want it, the composition holds, and the shots feel stable and intentional.

The best AI video creators aren’t the ones with the most expensive tools. They’re the ones who treat prompt writing like a craft — who learn the language of the camera, who iterate intentionally, and who never blame the model for what the prompt failed to say.

Now go make something beautiful.


Have questions about AI video generation or want to share your own tips? Find us at seedaudioprompts.com or join the conversation.