SunoMV SunoMV
FLUX 3 Prompt Guide: 8 Tips + Copy-Paste Templates
Guides

FLUX 3 Prompt Guide: 8 Tips + Copy-Paste Templates

Published · By SunoMV Team
Add SunoMV as a preferred source on Google See more SunoMV in Top Stories and AI answers.

You type cinematic, 8K, masterpiece into a box and hit generate. Ten seconds later the clip looks like a stock wallpaper that learned to walk. The chorus still has no camera. A second music bed fights the song you already made. You did not fail at “being more creative.” You handed FLUX 3 a mood board and asked it to direct a shot and score it.

People searching FLUX 3 prompt already picked the model. What they want is a string they can paste: the official one-liner shape, labeled camera fields when a lever has to move, timestamps as 0.0–5.0s or SHOT ONE / HARD CUT, then lock wardrobe because this picker has no nine-image face-lock slot.

This is not a methodology lesson. Every tip below is a full block. The anti-example is the same intent written as empty adjectives. Paste the block, pick FLUX 3 in the SunoMV audio-to-video generator, and export a shot that can sit on a chorus.

Table of Contents

Why “cinematic, 8K” is not a FLUX 3 prompt

FLUX 3 is Black Forest Labs’ video model. The lab’s own text-to-video prompting guide is blunt: think of the prompt as directing a scene, not listing objects. Official generation is 5 to 20 seconds, 720p or 1080p, with synchronized audio in the same pass. “Cinematic 8K” fills none of those slots.

On SunoMV the picker shows FLUX 3. Clips here run 5–15 seconds — write the length you will actually get. Resolution in the list is 720P or 1080P (no 480P row). The model does generate from text alone. It does take a still as the opening frame, a first-and-last pair, or up to ten storyboard stills. It does not take a vocal file as Audio 1. That last fact is the whole reason this page is not a second launch recap. The product-access writeup — FLUX 3 for AI music videos — covers the five generation modes. This page is only the pasteable prompt.

On this site, the same-cluster experiment is already measured. In May–July 2026, two Suno prompt pages sat at almost the same Search Console position: the copy-paste title (10 Tips + Copy-Paste Templates) pulled 3.79% CTR; the “7-Step Pro Method” title pulled 0.17% — a 22× gap on the same intent. The variable was the title shape, not “whether we optimized.” This page is the FLUX 3 version of the winner shape.

A cinematic AI music-video still: singer, rain, neon, written as a camera shot not as cinematic 8K

Image: SunoMV Team · music-video still written as camera + subject + rain, not as “cinematic 8K”

Practical rule: If the first line is not [camera] shot of [subject] [action] in [environment], or a labeled Camera shot: field, you are decorating a shot the model already chose for you.

The official slots FLUX 3 actually uses

Black Forest Labs’ text-to-video guide does not want MiniMax’s three labeled fields. For a multi-shot look that has to hold, it recommends a six-element schema. For a single chorus cut, three lighter formats are enough:

SlotFill thisAnti-example
One-liner[camera] shot of [subject] [action] in [environment]. then motion and sound“cinematic music video, 8K, masterpiece”
Labeled fieldsCamera shot / Subject + action / Lighting + palette / Motion / Style / Audioa paragraph of adjectives with no labels
Timestep0.0–2.5s — … then 2.5–5.0s — …. Mark a hard cut when the angle changes“then she walks, then later that night…”
Multi-shotSHOT ONE: … HARD CUT. SHOT TWO: … plus one audio bedthree locations with no cut language
Duration / ratioDuration: 10 seconds. Aspect ratio: 16:9. (5–15s here; official window 5–20s)“make it viral”
Continuity lockRepeat wardrobe, hair, and one distinguishing mark in every shot(nothing — the jacket recasts itself)

The six-element schema, when you need it: Core summary (who, where, the arc) → Scene (setting, light, depth of field) → Subject description (held identically across shots) → Dynamic narrative (timecoded camera + action) → Audio (per-shot soundscape) → Style & color. Name each with concrete nouns and verbs the camera can see.

FLUX 3’s camera cheatsheet is equally specific: start with one framing term, one movement term, and one clear subject action. low tracking shot is clear. low aerial handheld orbit push-in usually is not.

This picker has no native negative-prompt field for FLUX 3. Put “don’ts” as closing locks (Do not recast. Do not generate a new score.). Do not invent a second box.

This picker also has no audio-reference slot. MiniMax H3 on this site can take Audio 1. FLUX 3 cannot. Native audio is generated with the picture. For a Suno MV, that means you name rain and heels, and you write Music: none so a second bed does not fight the track you already have. See Black Forest Labs’ audio and speech guide for the four sound layers (speech, ambience, effects, music) — then leave music empty when the song is already the score.

A short 2026 overview of the model family is Black Forest Labs’ own FLUX 3 Sneak Peek. That video also talks about image, audio, and action-prediction in one architecture. Those extras are not what the SunoMV picker runs. Here you paste one shot, 5–15 seconds, onto a song.

Video: YouTube · Black Forest Labs FLUX 3 capability overview. SunoMV FLUX 3 is one clip, 5–15 seconds, 720P or 1080P, with optional first / last / storyboard stills.

Pick Style and batch generate on the SunoMV timeline — one FLUX 3 prompt is one clip

Image: SunoMV Team · FLUX 3 is a row in the model list. The prompt still has to fill camera, subject, and audio.

Practical rule: Write FLUX 3’s one-liner or labeled fields. Put Music: none when the Suno track is already the score. Duration is a shot-level lever; nine face stills are a series-level lever this model does not have.

8 copy-paste tips

Each tip is a full block. The anti-example is the same intent as empty adjectives. In SunoMV: pick FLUX 3, paste the whole block into the shot description, generate, then cut the clip onto the lyric line.

1. Start with the official one-liner, not a mood board

One-line rule: Black Forest Labs’ default shape is [camera] shot of [subject] [action] in [environment]. Skip the shape and the model fills framing for you.

Medium close-up shot of a woman in an oxblood leather jacket standing in rain on a neon Tokyo side street at night. She mouths the chorus, left heel planted, jacket catching magenta neon. Slow push-in, 35mm, eye-level. Rain ticks on the leather. Do not recast. Do not change hair. Music: none.

Anti-example: Epic cinematic banger, 8K, masterpiece, perfect lip sync.

How to use it in SunoMV: paste the whole sentence as the shot description. If you uploaded a still, that still is the first frame — the one-liner still has to name the motion. See Black Forest Labs’ image-to-video guide: the image holds the opening frame; the prompt drives what happens next.

2. When a lever has to move, paste labeled fields

One-line rule: the official labeled format isolates camera, subject, light, motion, style, and audio so you can change one without rewriting the rest.

Camera shot: medium close-up, eye-level, 35mm
Subject + action: a woman in her late 20s, blunt black bob, oxblood leather jacket over a white tank, silver hoops, plants her left heel and mouths the chorus
Depth of field: shallow (sharp on face, neon bokeh)
Lighting + palette: magenta tube light + wet-asphalt bounce — oxblood, silver, rain black
Motion: slow push-in over 8 seconds, rain on the jacket, no cut
Style: live-action, cinematic, fine grain
Audio: rain on asphalt, distant traffic, her breath on the inhales. Music: none
Duration: 8 seconds. Aspect ratio: 16:9.

Anti-example: Use the attached photos and make her sing. Cinematic music video.

How to use it in SunoMV: this is the FLUX 3 equivalent of MiniMax’s three fields — different labels, same job. Do not paste integrated_multimodal_description here; that is a MiniMax slot. For MiniMax H3’s three fields, see the MiniMax H3 prompt guide.

3. Timestamp the FLUX 3 way: 0.0–2.5s or HARD CUT

One-line rule: FLUX 3 does not want MiniMax’s [Shot 2] At 00:05.000. Official timestep prompting uses ranges. A new angle is a hard cut, not “then later.”

Duration: 10 seconds. Aspect ratio: 16:9.

SHOT ONE: 0.0–5.0s — medium-wide, rain-soaked Tokyo side street at night. She walks toward camera, oxblood jacket, blunt black bob. Slow dolly in, 35mm. Heel strikes wet asphalt.
HARD CUT.
SHOT TWO: 5.0–10.0s — tighter close-up, 85mm. Rain on her lashes. She inhales and mouths the last line. Camera holds. Keep face, bob, and jacket identical.
Audio: rain, heel on wet asphalt, one sharp inhale at 8s. One ambient bed across both shots. Music: none.

Anti-example: Then she walks, then she dances, then later that night she is on a rooftop.

How to use it in SunoMV: if the prompt contains “then later,” you are writing a sequence the 15-second lock cannot hold. Split it, or keep every cut inside the timestamps. Official samples also write SHOT ONE: … HARD CUT. SHOT TWO: … as one paragraph — same idea, denser. A 10-second oner is a chorus, not the whole MV.

Character-consistency stills: wardrobe repeated beats one pretty adjective

Image: SunoMV Team · FLUX 3 has no nine-image face-lock; the prompt has to repeat the jacket

4. Put duration and ratio in the first line, then honor the 5-second floor

One-line rule: Black Forest Labs’ public window is 5–20 seconds. This picker note is 5–15 seconds. Write the length you will actually get.

Duration: 8 seconds. Aspect ratio: 16:9.
One continuous take, no cut inside the clip.
Start on her hands at the jacket zipper, end on her eyes as she looks up into the rain.
Camera: slow tilt up, 35mm, locked-off except the tilt.
Audio: zipper, rain, one inhale. Music: none.

Anti-example: A full music video that follows her all night across the city.

How to use it in SunoMV: one prompt = one generation, 5–15 seconds. For 9:16 Reels, change the ratio line only. Pick 1080P in the list when you want Full HD; pick 720P when you want the faster HD pass. Official 20-second clips exist on the lab’s API; they are not what this row currently advertises.

5. Name one subject the way a costume department would, then repeat it

One-line rule: “a girl” is a casting call. Wardrobe is a lock. FLUX 3 on this site has no nine-image face-lock slot — keyframes are a storyboard, not a character library. The same three-to-five anchors have to appear in every shot.

Subject: a woman in her late 20s, blunt black bob, silver hoop earrings,
oversized oxblood leather jacket over a white tank, chipped black nail polish.
Keep the same jacket and hair in every shot. Do not recast. Do not change hair.

Anti-example: A beautiful mysterious girl with good vibes.

How to use it in SunoMV: copy this subject block into every FLUX 3 shot for the same song so the chorus and the verse do not recast her. If you need a still before the first video generation, pair it with musician image prompts — or generate that still on ChatImg with the same wardrobe line. For true cross-shot face-lock (up to nine stills), pick Seedance or MiniMax H3 instead; this page is only FLUX 3.

6. Write action as physics, and write the camera as one move

One-line rule: “dancing beautifully” has no weight. “Heel strikes wet asphalt” does. The cheatsheet wants one framing term + one movement term. Named camera moves beat mood words — the same craft shows up in Runway’s camera-prompt notes.

Camera shot: hip-height tracking left, 35mm, slow
Subject + action: she plants her left heel, weight forward, oxblood jacket swinging half a beat late. She mouths the line without smiling.
Motion: rain on the jacket, one tracking move, no orbit, no push-in stacked on top
Audio: heel strike, jacket leather, rain, no applause. Music: none

Anti-example: She dances beautifully with amazing energy and perfect cinematic audio.

How to use it in SunoMV: if you stacked low aerial handheld orbit push-in, delete two of the four camera words. Keep the heel. Keep Music: none.

7. First and last frame still need a prompt — keyframes are a storyboard, not a face-lock

One-line rule: the frames hold identity. The prompt still has to name the path between them. Official i2v is one keyframes field: one image = opening frame; two images = start and end; three to ten = ordered waypoints, spread evenly or pinned to timestamps.

Duration: 8 seconds. Aspect ratio: 16:9.
Image 1 is the first frame: she stands outside the shop, both hands in her jacket pockets.
Image 2 is the last frame: she is one step closer, right palm on the glass, her reflection sharp in the window.
Take her from the ready stance in Image 1 through one step to Image 2, flowing naturally, no extra cuts.
Camera: slow push-in, 35mm, eye-level. Rain ticks on the awning.
Keep face, bob, and oxblood jacket identical.
Audio: rain, one shoe on wet pavement. Music: none.

Anti-example: Image 1 and Image 2. Make a music video between them.

How to use it in SunoMV: upload the start still and the end still, then paste the path. Do not ask this mode for a three-shot montage — split those into separate generations. Do not dump nine headshots into the keyframe slots hoping for Seedance-style lock; those slots are time, not identity. The character-consistency method is the longer version of this lock on models that actually take nine stills.

8. For a music video, kill the generated score; this picker has no Audio 1

One-line rule: FLUX 3 renders synchronized audio with the picture. A scene that implies rain will invent rain. A scene that implies “epic music” will invent a second bed. There is no Audio 1 slot to copy a vocal from.

Duration: 12 seconds. Aspect ratio: 16:9.
Medium close-up shot of the woman from the still, oxblood jacket, blunt black bob, on the neon side street. She mouths the chorus. Slow push-in, 35mm. Copy the mouth rhythm of a sung chorus — stress on the downbeats, no smiling.
Audio:
  Ambience: rain against a metal awning, distant traffic
  Effects: heel on wet asphalt, leather jacket
  Speech: none
  Music: none
Do not generate a new vocal. Do not generate a new beat. Do not recast.

Anti-example: Perfect lip sync and an epic original soundtrack.

How to use it in SunoMV: drop the song (Suno link or your file) onto the timeline, pick FLUX 3, keep one still as the first frame so the face does not recast, and keep Music: none so the generated clip does not also invent a competing bed. Lip motion will be approximate — this is not MiniMax H3’s Copy phrase timing from Audio 1. If you need a vocal lock, use the MiniMax H3 prompt guide. If you need a picture-only lock with no native audio at all, see Kling O3.

Diagnostic still of an AI music video that drifted off the beat

Image: SunoMV Team · the usual failure is a second score, not a missing adjective

Practical rule: If you already have a Suno track, Music: none is the score lock. “Perfect lip sync” is not a slot on FLUX 3.

5 music-video templates

Copy a block. Fill the brackets. Keep FLUX 3’s one-liner or labeled fields. Each template is one generation, 5–15 seconds.

Duration: 10 seconds. Aspect ratio: 16:9.

SHOT ONE: 0.0–6.0s — live-action, cinematic, medium-wide on a rain-soaked Tokyo side street at night. A woman in her late 20s, blunt black bob, oxblood leather jacket over a white tank, silver hoops, walks toward camera. Magenta neon in the puddles. She mouths the chorus. Camera: slow dolly in, 35mm.
HARD CUT.
SHOT TWO: 6.0–10.0s — medium close-up. Rain on her lashes. Hold through the last line. Keep face, bob, and jacket identical.

Audio: rain, distant traffic, heel on wet asphalt. One ambient bed across both shots. Music: none.
Do not recast. Do not generate a new score.

Lyric close-up (mouth + eyes, no recast)

Duration: 6 seconds. Aspect ratio: 16:9.
Camera shot: extreme close-up, 85mm, eyes and mouth only, tiny handheld drift
Subject + action: the woman with the blunt black bob and silver hoops sings one line, no smiling, catchlight from a pink tube light
Motion: no cut, no recast, keep the bob and hoops identical to the still
Style: live-action, shallow depth of field
Audio: soft room tone only. Music: none

Cinematic oner (one take, no invented cuts)

Duration: 12 seconds. Aspect ratio: 21:9.
Wide-to-medium tracking shot of a woman in an oxblood leather jacket walking a wet riverside promenade at blue hour. One continuous take. Slow dolly right, 35mm, eye-level. She looks once at the water, then forward. Hands in pockets. No cut. No time-lapse. Keep identity identical to the still.
Audio: river, distant train, her footsteps. Music: none.

Vertical short (9:16, caption headroom)

Duration: 8 seconds. Aspect ratio: 9:16.
Medium close-up shot of the woman with the blunt black bob, subject in the center third, headroom for captions. Neon stairwell, one pink tube light. She looks up, blinks once, then mouths the hook. Camera locked off. Keep her identical to the still.
Audio: hum of the tube light. Music: none.

First-last frame transition (storyboard stills, not a face library)

Duration: 8 seconds. Aspect ratio: 16:9.
Image 1 is the first frame: chorus still, she stands at the shop window, both hands in the oxblood jacket.
Image 2 is the last frame: she is one step closer, right palm on the glass, magenta neon in the reflection.
Take her from Image 1 to Image 2 in one step, flowing naturally, no extra cuts.
Camera: slow push-in, 35mm, eye-level. Rain ticks on the awning.
Keep face, bob, and jacket identical.
Audio: rain, one shoe on wet pavement. Music: none.

You do not need ten keyframes every time. Two consistent stills already beat one pretty headshot used as a “reference library.” The AI music video creation guide is the workflow for stringing those shots.

Sample clip: SunoMV Team, 2026-08-05, FLUX 3 text-to-video (5 seconds / 720p). Prompt required: stable facial features, beat-synced motion, golden-hour backlight, one camera move.

Practical rule: A template is finished when duration, ratio, wardrobe, one camera move, and Music: none are all on the page — never “epic music.”

Empty prompt vs a FLUX 3 prompt

Same song, same 10 seconds, two inputs. Only one of them is a FLUX 3 prompt.

EmptyCopy-paste
First lines“cinematic, 8K”[camera] shot of [subject] [action] in [environment], or labeled fields
Subject“a girl”wardrobe + hair + one distinguishing mark, repeated
Action“dancing”heel, weight, contact with the ground
Time(none)0.0–5.0s / HARD CUT summing to ≤15s
Audio“perfect lip sync + epic music”named rain/heels; Music: none
Continuity(hope)“keep the jacket identical” — no nine-still face-lock on this row
Who it’s fora still that pretends to be a videoa chorus cut you can actually edit

If you already have a Suno track, do not ask FLUX 3 to write another one. Point it at the picture, kill the generated score, keep the song. For a MiniMax-shaped vocal lock (Audio 1), see the MiniMax H3 prompt guide. For a Kling-shaped picture-only lock (no audio reference in this picker), see the Kling O3 prompt guide. For a Wan-shaped 30-second lock, see the Wan 3.0 prompt guide. For Seedance-style second-level timestamps, see the Seedance 2.5 prompt guide. Different slots, same “copy the block” job.

Default music-video generation on this site is still storyboard still → image-to-video, not “open every mode on the lab page.” Video edit restyles a clip you already have and keeps its motion — that is not face-lock, and it is not the first button to press. The FLUX 3 music-video recap is the product page; this page is only the pasteable prompt.

A sister-site note if you also need a transcript of the finished MV: BibiGPT’s YouTube transcript generator is the paste-a-link path for that, not this page.

Practical rule: If you cannot point to a first frame (or a first-and-last pair), plus Music: none, you do not have a FLUX 3 music-video prompt. You have a vibe.

FAQ

Do I write “FLUX 3” inside the prompt?

No. The model name does not belong in the prompt text. People search FLUX 3 prompt; SunoMV’s picker lists FLUX 3. Pick in the list, not in the sentence.

Why did my clip stay at 5–15 seconds when the lab page says 20?

Write the duration as a field. Black Forest Labs’ public window is 5–20 seconds; this picker uses 5–15 seconds for FLUX 3. A prompt that describes a minute of story will not get a minute. See the audio-to-video generator.

First and last frame: do I still write a prompt?

Yes. The frames hold identity. The prompt still has to name the camera and the step between them, and it must not ask for extra cuts. See tip 7. Official i2v is one keyframes field — one still, a pair, or up to ten waypoints. That is a storyboard, not a nine-image character library.

Why is there a second music bed under my Suno track?

You left audio empty or wrote “epic score.” Fill Music: none, and name the rain. FLUX 3 generates sound with the picture; it cannot take Audio 1. That is the opposite of MiniMax H3, which can.

Can I upload nine face stills to lock the singer?

Not on this row. FLUX 3’s extra stills are keyframes in time. Face-lock with up to nine images is Seedance / MiniMax H3 on this site. Repeat the wardrobe line instead, or switch model. See the character-consistency method.

Should I write negative prompts?

Not in a separate box — FLUX 3 here has no native negative field. Put exclusions in the description: Do not recast. Do not generate a new score.

Can one prompt be the whole music video?

No. One prompt is one generation: 5–15 seconds. String the shots. This page is only the pasteable FLUX 3 prompt. The AI music video creation guide covers the string.

Is “video edit” the same as reference-to-video / face-lock?

No. On this picker, video edit restyles a source clip and keeps its motion. That is rewrite-an-existing-shot, not lock-this-face-across-the-MV. Default out is still storyboard stills into image-to-video. Specs independent of this site live in Black Forest Labs’ FLUX 3 video page and the generation blog.

How does FLUX 3 relate to the FLUX image models?

Same lab, two product lines. The image models are stills and text rendering. FLUX 3 is the video line: time, camera, and synchronized audio. Do not paste a FLUX.2 still prompt into this box and expect a chorus cut.

Copy, pick FLUX 3, export the MV

A FLUX 3 prompt is finished when you can paste it, not when you can explain it.

  1. Copy one template above. Fill the brackets. Keep the one-liner or labeled fields. Repeat wardrobe. Write Music: none.
  2. Open the SunoMV audio-to-video generator. Drop in the song (Suno link or your own file).
  3. In the model list, pick FLUX 3. Set 5–15 seconds. Use 1080P when you want Full HD.
  4. Paste the whole prompt into that shot. If you uploaded stills, say what Image 1 / Image 2 are for (first frame, last frame, or a waypoint) — do not treat them as a face library.
  5. Export the clip onto the lyric line. Repeat. The MV is the string of shots, not one heroic paragraph.

Every model you can click is a prompt guide that has not been written yet. The people who get usable FLUX 3 footage are not the ones with a better adjective list. They are the ones who already have a pasteable prompt before they touch generate.

SunoMV Team

View all 31 articles in Suno Prompts & AI Songwriting →

Try these AI tools