SunoMV SunoMV
Wan 3.0 Prompt Guide: 8 Tips + Copy-Paste Templates
Methodology

Wan 3.0 Prompt Guide: 8 Tips + Copy-Paste Templates

Published · By SunoMV Team
Add SunoMV as a preferred source on Google See more SunoMV in Top Stories and AI answers.

You type cinematic, 8K, masterpiece into a box and hit generate. Thirty seconds later the clip looks like a stock wallpaper that learned to walk. The chorus still has no camera. The singer’s jacket changes color on the second beat. You did not fail at “being more creative.” You handed the model a mood board and asked it to direct a shot.

People searching Wan 3.0 prompt already picked the model. What they want is a string they can paste: duration and ratio first, then Image 1 / Video 1 / Audio 1 if they uploaded anything, then action, scene, quoted lines, and a closing lock on what must not change.

This is not a methodology lesson. Every tip below is a full block. The anti-example is the same intent written as empty adjectives. Paste the block, pick Wan 3.0 or Wan 3.0 Prime in the SunoMV audio-to-video generator, and export a shot that can sit on a chorus.

Table of Contents

Why “cinematic, 8K” is not a Wan 3.0 prompt

Alibaba Cloud’s Wan 3.0 video generation guide is blunt about the physical facts: a single generation is 2 to 30 seconds, output is 30 fps, and resolution is 480P, 720P, or 1080P. The same page is also blunt about Omni-Reference: up to 20 multimodal assets in one request — images, video, audio, documents, and web pages — addressed in the prompt as Image 1, Video 1, and Audio 1.

That is the whole job. A Wan 3.0 prompt is a shot brief that fills those slots. “Cinematic 8K” fills none of them. I’d rather write one boring, complete sentence per slot than a paragraph of “epic energy.” Empty adjectives are how you get a pretty clip that cannot cut to a downbeat.

On this site, the same-cluster experiment is already measured. In May–July 2026, two Suno prompt pages sat at almost the same Search Console position: the copy-paste title (10 Tips + Copy-Paste Templates) pulled 3.79% CTR; the “7-Step Pro Method” title pulled 0.17% — a 22× gap on the same intent. The variable was the title shape, not “whether we optimized.” This page is the Wan 3.0 version of the winner shape.

Night-city music video still for a Wan 3.0 prompt: singer in rain and neon

Image: SunoMV Team · music-video still written as camera + subject + rain, not as “cinematic 8K”

Practical rule: If the first line is not duration, ratio, or a named reference, you are decorating a shot the model already chose for you.

The official slots Alibaba actually uses

Alibaba’s text-to-video prompt guide gives the reference-to-video formula for Wan 3.0 as reference identifier + action + scene + lines (optional) + background music (optional). The usage order that matches the official samples is: duration and aspect ratio first, then subject and assets, then scene and style, then shot and motion, then dialogue and SFX, then exclusions, then a timeline.

Fill the table. Any empty cell becomes the blandest default.

SlotFill thisAnti-example
Duration / ratio20 seconds, 16:9 or 9:16 (or adaptive)“make it viral”
Reference identifiersImage 1 is the woman. Image 2 is the street. Audio 1 is her voice.”use these references”
SubjectOne person, wardrobe, age band, hair”a beautiful girl”
ActionVerbs + physics (weight, speed, contact)“dancing beautifully”
SceneRoom, weather, time of day, foreground clutter”epic location”
Shot / motionShot size + named move + speed”cinematic camera”
Dialogue / SFXQuoted lines; named sounds”add dramatic music”
Continuity lockWhat must stay identical across the 30 seconds(nothing — the jacket recasts itself)

English identifiers take a space: Image 1, Video 1, Audio 1. Images, videos, and audio are numbered separately by upload order — the first reference image is always Image 1 even if a video was uploaded first. The Wan 3.0 API reference also caps a prompt at 20,000 characters and a clip at 2–30 seconds (default 5; -1 lets the model pick a smart duration).

In SunoMV the picker shows Wan 3.0 and Wan 3.0 Prime. Prime is the same toolkit with a faster turnaround. Single shots go up to 30 seconds; transition clips cap at 15 seconds. If the prompt describes a 30-second oner, pick a single shot, not a transition.

The official Wan 3.0 public-beta clip is useful as a picture of what “one generation” looks like when the brief is specific. A longer walkthrough of the same model — references, native audio, 30-second takes — is Curious Refuge’s Wan 3.0 review (about 21 minutes).

Video: YouTube · Curious Refuge · Wan 3.0 review — references, native audio, 30-second clips

Practical rule: Write the official slots in order, then add audio as its own sentences. Do not bury “she whispers” inside the wardrobe line.

8 copy-paste tips

Each tip is a full block. The anti-example is the same intent as empty adjectives. In SunoMV: pick Wan 3.0 or Wan 3.0 Prime, paste the whole block into the shot description, generate, then cut the clip onto the lyric line.

1. Bind every reference before you describe the shot

One-line rule: uploading is half the job. Wan 3.0 will not guess that picture 3 is the coffee shop.

Image 1 is the woman.
Image 2 is the Tokyo side street at night.
Audio 1 is her speaking voice.
The woman from Image 1 walks toward camera on the street in Image 2
and says, "Don't wait for the chorus."
Keep her face, blunt black bob, oxblood jacket, and silver hoops identical to Image 1.

Anti-example: Use the attached photos. Cinematic music video.

How to use it in SunoMV: if you uploaded stills, this block goes above the rest of the prompt. If you uploaded nothing, delete the Image / Audio lines and keep the rest.

2. Put duration and aspect ratio in the first line

One-line rule: the model will invent a length if you do not write one. Official samples lead with the number of seconds.

Duration: 20 seconds. Aspect ratio: 16:9.
One continuous take, no cut inside the clip.
Start on her hands at the jacket zipper, end on her eyes as she looks up into the rain.

Anti-example: A full music video that follows her all night across the city.

How to use it in SunoMV: one prompt = one generation. A 20-second oner is a chorus or a verse, not the whole MV. For 9:16 Reels, change the first line only.

3. Number the shots and give each one a time range

One-line rule: 30 seconds is a timeline, not a caption. Official long takes use Shot 1 [0s-8s] style ranges.

Shot 1 [0s-8s] Medium close-up, slow dolly push-in, 35mm.
She takes three steps toward camera, left heel striking wet asphalt.
Shot 2 [8s-16s] Tracking from the left, hip height. She turns her head, rain on her lashes.
Shot 3 [16s-20s] Hold on her eyes. She inhales. No cut. No time-lapse.

Anti-example: Then she walks, then she dances, then later that night she is on a rooftop.

How to use it in SunoMV: if the prompt contains “then later,” you are writing a sequence the 30-second lock cannot hold. Split it, or keep the whole story inside the time ranges.

Dolly camera pushing toward a singer on a wet street for a Wan 3.0 prompt

Image: SunoMV Team · the first clause after the timestamp is the camera move, not the adjective “cinematic”

4. Name one subject the way a costume department would

One-line rule: “a girl” is a casting call. Wardrobe is a lock. The same three-to-five anchors should repeat in every shot.

Subject: a woman in her late 20s, blunt black bob, silver hoop earrings,
oversized oxblood leather jacket over a white tank, chipped black nail polish.
Keep the same jacket and hair in every shot. Do not recast. Do not change hair.

Anti-example: A beautiful mysterious girl with good vibes.

How to use it in SunoMV: copy this subject block into every Wan 3.0 shot for the same song so the chorus and the verse do not recast her. Pair it with musician image prompts if you need a still of her before the first video generation.

5. Write action as physics, not as a compliment

One-line rule: “dancing beautifully” has no weight. “Heel strikes wet asphalt” does. The same craft shows up in Runway’s camera-prompt notes: named moves beat mood words.

Action: she takes three slow steps toward camera, left heel striking wet asphalt,
jacket hem swinging, rainwater flicking off the collar.
She does not spin. She does not jump. Weight stays in the front foot.

Anti-example: She dances energetically with amazing choreography.

How to use it in SunoMV: use this on a pre-chorus walk-up. High-burst flips are where long takes fall apart; keep the body slow and let the camera move.

6. Replace “cinematic lighting” with a direction and a texture

One-line rule: light direction is the cheapest quality lever in the official samples.

Style and ambiance: hard key from camera-left neon, magenta, cutting a
narrow rim on her jaw. Soft fill from a shop window on the right.
Low-key contrast, wet highlights, slight grain, no clean beauty lighting.

Anti-example: Cinematic lighting, volumetric god rays, ultra realistic.

How to use it in SunoMV: keep the same key/fill pair across shots so the grade does not reset every generation.

7. Quote dialogue on its own line; name SFX as sounds

One-line rule: if the line is not in quotation marks, it is narration, not a line. Native audio is on by default in Wan 3.0; tell it no score if you already have a Suno track.

She looks just past camera and says, "Don't wait for the chorus."
Keep her mouth in frame. No other speakers.
SFX: a sharp heel click on wet brick; rain ticking on a metal awning.
Ambient: low traffic two streets over, a vending machine hum.
No score. No vocal ad-libs from off-screen.

Anti-example: She talks emotionally. Add epic sound design.

How to use it in SunoMV: you already have a song. Point Wan 3.0 at the picture and keep the track. The Suno prompt tips guide is the pasteable list for the audio half; this page is the pasteable list for the picture half. If you need a lyric block first, use the lyric generator.

Lyric close-up at a vintage microphone for a native-audio Wan 3.0 prompt

Image: SunoMV Team · a lyric close-up is a mouth, a mic, and one quoted line

8. Lock identity with a closing constraint, not a hope

One-line rule: official samples spend a surprising share of their words on “keep / lock / do not change.” Duration is a shot-level lever. Reference capacity is a series-level lever — a serial dies from character drift, not from a weak five-second clip.

Closing constraints:
Keep the woman identical to Image 1: face, blunt black bob, oxblood jacket, silver hoops.
No subtitles. No watermark. No second face. No crowd.
One unbroken take. Do not recast between Shot 1 and Shot 3.

Anti-example: Make her consistent and professional.

How to use it in SunoMV: this block goes last. If you skip it, a 20-second oner will still look expensive — and she will be a different person at 0:16.

Practical rule: If you cannot point to the start frame and the end frame in the prompt, you do not have a shot. You have a vibe.

5 music-video templates

Each template is one paste. Swap the bracketed bits. Keep the slot order.

Pop chorus oner (16:9, 20 seconds)

Duration: 20 seconds. Aspect ratio: 16:9. One continuous take.
Subject: a woman in her 20s, red vinyl jacket, wet hair stuck to her cheek.
Shot 1 [0s-8s] Medium shot, eye-level, slow push-in, 35mm.
She mouths the chorus and steps one pace toward camera, rain on her lashes.
Shot 2 [8s-16s] Hold, then a 2cm drift. City grid behind her, one red aviation light blinking.
Shot 3 [16s-20s] End on her eyes. She smiles on the last beat.
Style: hard backlight, thin fog, teal-and-red grade, light grain.
Audio: she sings, "Stay until the lights go out." Ambient: wind on a metal rail.
No score. No subtitles.

Lyric close-up (mouth + caption room)

Duration: 8 seconds. Aspect ratio: 16:9.
Cinematography: close-up, slight low angle, locked-off with a 2cm drift.
Subject: the same woman, oxblood jacket, silver hoops, a small scar on the left eyebrow.
Action: she inhales, then delivers one line, eyes wet but not crying.
Context: dark studio, a vintage microphone in the lower third, out-of-focus meters behind.
Style: warm key from the right, cool rim, shallow focus on the mouth.
Audio: she says, "I kept the chorus." Ambient: room tone. No score.
Leave headroom at the top for captions.

Vertical short (9:16)

Duration: 10 seconds. Aspect ratio: 9:16.
Cinematography: medium close-up, slow crane-up from chest to eyes.
Subject: the same woman in the red vinyl jacket, rain in her hair.
Action: she looks up, blinks once, then smiles on the last beat.
Context: neon stairwell, one pink tube light, condensation on the rail.
Style: magenta key, green bounce from a sign, crushed blacks.
Audio: she says, "Hit post." SFX: a single heel on metal stairs. No score.
Keep her in the center third. Headroom for captions.

Vertical 9:16 music video frame of a singer in neon rain for a Wan 3.0 prompt

Image: SunoMV Team · 9:16 as a written slot, not a crop of a wide shot

Character lock with Image 1 (Omni-Reference)

Duration: 12 seconds. Aspect ratio: 16:9.
Image 1 is the woman. Image 2 is the convenience-store window at night.
The woman from Image 1 leans on the glass in Image 2 and watches her own reflection.
Preserve her face, blunt black bob, oxblood leather jacket, white tank, silver hoops.
Cinematography: wide-to-medium, slow dolly right, 35mm.
Style: sickly green fluorescents vs magenta street neon, wet glass.
Ambient: compressor hum, distant register beep. No score.
Do not recast. Do not change hair.

Same singer at a convenience-store window, identity locked to Image 1

Image: SunoMV Team · Image 1 holds the face; the prompt only names what is allowed to change

Native-audio two-shot

Duration: 10 seconds. Aspect ratio: 16:9.
Image 1 is the woman. Image 2 is the man in the grey overcoat. Audio 1 is her voice.
Cinematography: two-shot, eye-level, slow pan from him to her, 40mm.
Action: he turns, she does not. She answers without looking at him.
Context: rain under a bus shelter, one fluorescent tube buzzing.
Style: hard overhead fluorescent, greenish, wet pavement sheen.
He says, "You coming?" She says, "After the chorus."
Subject 1 uses the voice timbre of Audio 1.
SFX: rain on the shelter roof; a bus air-brake. Ambient: traffic. No score.
Keep both mouths in frame. No crowd.

Two-shot under a rainy bus shelter for a native-audio Wan 3.0 prompt

Image: SunoMV Team · quoted lines only work if the mouths stay in frame

Empty prompt vs a Wan 3.0 prompt

Same song, same 20 seconds, two inputs. Only one of them is a Wan 3.0 prompt.

EmptyCopy-paste
First line”cinematic, 8K”duration + ratio, or Image 1 is…
Subject”a girl”wardrobe + hair + one distinguishing mark
Action”dancing”three steps, weight, contact with the ground
Time(none)Shot 1 [0s-8s]Shot 3 [16s-20s]
Audio(none, or “add music”)quoted line + named SFX + “no score”
Continuity(hope)“keep the jacket identical to Image 1”
Who it’s fora still that pretends to be a videoa chorus cut you can actually edit

If you already have a Suno track, do not ask Wan 3.0 to write another one. Point it at the picture and keep the song. For a Veo-shaped 8-second lock, see the Veo 3 prompt guide. For Seedance-style second-level timestamps, see the Seedance 2.5 prompt guide. Different slots, same “copy the block” job. The AI music video creation guide is the workflow for stringing those shots.

A sister-site note if you also need a transcript of the finished MV: BibiGPT’s YouTube transcript generator is the paste-a-link path for that, not this page.

Decision filter: If you cannot point to Image 1 (or a wardrobe lock) and a last frame, you do not have a Wan 3.0 prompt. You have a vibe.

FAQ

Do I write “Wan 3.0” inside the prompt?

No. The model name does not belong in the prompt text. People search Wan 3.0 prompt; SunoMV’s picker lists Wan 3.0 and Wan 3.0 Prime. The slots are the same. Prime is the faster option, not a different formula.

Why did my duration get lifted, or why did a 30-second oner get chopped?

Write the duration as a field. Official range is 2–30 seconds. In SunoMV, single shots can use the 30-second lock; transitions cap at 15 seconds. A prompt that describes a minute of story will not get a minute.

First and last frame: do I still write a prompt?

Yes. The frames hold identity. The prompt still has to name the camera, the action, and the audio. A first-last pair with cinematic 8K in the middle is still an empty prompt.

Why is there no dialogue even though I wrote a line?

Usual cause: the line was not in quotation marks, so it was treated as description. Quote it, keep the mouth in frame, and add No score if your Suno track should be the only music.

How many references can I upload?

Official Omni-Reference is up to 20 assets in one request, with a practical split of up to 10 images, 5 videos, and 5 audio clips, plus one document or one public web link. Bind each one. An unbound still is a decoration.

Should I write negative prompts?

Write the positive shot. Official samples spend their words on what is in the frame. If you must exclude something (a score, a crowd, a second face), say No score. No crowd. as a closing constraint — not a paragraph of “don’ts.”

Can one prompt be the whole music video?

No. One prompt is one generation: up to 30 seconds. String the shots. This page is only the pasteable Wan 3.0 prompt.

Copy, pick Wan 3.0, export the MV

A Wan 3.0 prompt is finished when you can paste it, not when you can explain it.

  1. Copy one template above. Fill the brackets. Keep Alibaba’s slot order.
  2. Open the SunoMV audio-to-video generator. Drop in the song (Suno link or your own file).
  3. In the model list, pick Wan 3.0 or Wan 3.0 Prime. Prime if you want the same toolkit faster.
  4. Paste the whole prompt into that shot. Do not split the Image 1 sentence from the action sentence.
  5. Export the clip onto the lyric line. Repeat. The MV is the string of shots, not one heroic paragraph.

The people who get usable Wan 3.0 footage are not the ones with a better adjective list. They are the ones who already have a pasteable prompt before they touch generate.

SunoMV Team

View all 27 articles in Suno Prompts & AI Songwriting →

Try these AI tools