SunoMV SunoMV
HeyGen Video for Music Videos: Sound in the Shot, 40 Credits to Try
Trending

HeyGen Video for Music Videos: Sound in the Shot, 40 Credits to Try

Published · By SunoMV Team
Add SunoMV as a preferred source on Google See more SunoMV in Top Stories and AI answers.

HeyGen published HeyGen Video on 30 September 2026. The loud line in that post is the October price: about $0.01 per second, called half of list. The catalog is quieter. It is a 5–15 second clip, 480p or 768p, from text, a first frame, or references, and the picture comes with sound.

If you are finishing a music video, the price is not the part that usually breaks the chorus. The break is a silent shot, or a new face on the next line. This note does not retype the catalog. It says what you can pick tonight in SunoMV.

A chorus still is the picture. The model still has to deliver a moving shot with sound, on the length the picker allows.

Cinematic still from an AI music video built from a song

Image: SunoMV Team · a music-video still is not yet a 5–15 second shot with sound.

What the picker actually offers

HeyGen’s announcement covers the model they shipped. The SunoMV video list is shorter, and it is the list that charges you:

CapabilitySunoMV picker (HeyGen Video)
InputsText, first frame, or references (up to 9 images, plus video and audio)
Duration5–15 seconds, whole seconds
Resolution480P or 720P (720P is the 768p picture)
SoundIncluded. There is no off switch
Quota40 credits / 5s at 480P; 60 credits / 5s at 720P
Not in this UIFirst-and-last-frame, 1080P, native 2K

Practical rule: Credits follow the list price, not the October sale. 480p list is $0.02/s (8 credits/s). 768p list is $0.03/s (12 credits/s). A 5-second draft is 40 credits, not the promotional invoice.

The product page with the contrast lives at HeyGen Video.

The list price is not the October sale

HeyGen’s post calls the October rate about $0.01 per second, half of list. The published list we price against is $0.02 per second at 480p and $0.03 per second at 768p for a text or image generation. Reference generation is double that list.

SunoMV turns those list rates into credits with the same rule as the other video rows: round the dollars-per-second by 400. That is a 2× markup on the large credit pack, which is a 50% gross margin. It is about 4× on the small pack. The October half-off does not enter the picker.

Practical rule: A sale you see in an announcement is not a sale on the export button. Read the credit line in the picker.

We ran one 5-second 480P clip

On 1 October 2026 we generated one text-to-video clip: 5 seconds, 480P. It came back 832×480, about 5.3 seconds long, with sound. The invoice was $0.0495. That is the October promotional rate, about $0.01 per second. It is not the number on the picker.

We are not publishing that file as a demo asset, and we are not turning one clip into a speed promise. The only point: the shot is a real picture-with-sound take, and the sale price is temporary.

Practical rule: If a quote says $0.01 per second, ask whether that is October or the list. SunoMV charges the list.

Reference images hold a face; first-and-last-frame does not

Two consecutive lyric lines fail when the singer changes. There are two different fixes, and this model only has one of them.

Reference mode takes up to nine images, and you can also attach video and audio. That is how you tell the shot who the person is. It does not let you pin the last frame of shot A as the first frame of shot B. First-and-last-frame lives on other rows.

The picture below is the failure mode, not a HeyGen sample. A new jacket and a new room between two lines is what reference images are for.

Typical AI music-video failure: the next line recasts the singer

Image: SunoMV Team · the cut fails when the next line is a different person.

Holding one face across shots is a reference job. Walking from a pinned opening picture to a pinned closing picture is a different control, and HeyGen Video does not offer it.

Character consistency across AI music-video shots

Image: SunoMV Team · same person across shots is the job of reference images on this row.

When this row is the right one

Use it when the chorus shot should arrive with sound, or when you already have stills of the singer and need up to nine of them held as references.

Skip it when you are still throwing ideas away and you need a first frame locked to a last frame. That is MiniMax H3 Max: cheaper iteration on a different job, no reference-to-video. Skip it when you need native 2K or a singing mouth matched to a vocal. That is another row.

If you already have a Suno track, start from Suno to Video and switch the video model to HeyGen Video.

The timeline is where separate shots become one video. Each HeyGen clip is only one lyric line.

Aligning picture, lyrics and beat on the timeline

Image: SunoMV Team · stack the 5–15 second shots on the lyric timeline, then export.

Pick it today

  1. Open SunoMV.
  2. Paste the song, or start from a Suno link.
  3. Choose HeyGen Video in the video-model list. The same name is on the homepage model list.
  4. Set 5–15 seconds and 480P (40 credits) or 720P (60 credits).
  5. Add reference images when the face has to hold. Do not look for a first-and-last-frame switch.
  6. Stack on the lyric timeline, preview free, export.

Specs, the credit math, and the announcement-versus-picker contrast are on HeyGen Video. This note only has to do one job: the October cent-per-second is a sale. The shot you export is priced at the list.

View all 41 articles in AI Music & Video Tool Comparisons →

Try these AI tools