Back to blogfr
7 August 2026Maxime Jegat

Seedance 2.5 Guide: Mastering the Art of 30-Second Videos

Crafting a 30-second Seedance 2.5 prompt requires a shooting plan, not just a description. Focus on structure, audio syntax, and pitfalls to avoid.

Prompt Seedance - Guide complet

A Seedance 2.5 prompt that worked for a 5-second clip can't cut it for 30 seconds. The model now creates a seamless 30-second take in one go, complete with audio, transforming how you write. You're not just describing a moving image anymore; you're crafting a narrative.

This is where most prompts fail. The model has 30 seconds to fill, and if you don't guide it at every step, it improvises. Faces change at second twelve, props teleport, and you end up with an unusable rush that still costs you.

Here's how we craft Seedance prompts: what the official docs require and what we've learned from delivering rushes.

What Seedance 2.5 changes, and where to use it today

ByteDance announced the model on June 23, 2026, at its Volcano Engine FORCE conference, opened it to creators on July 31 on Dreamina, Jimeng, and Doubao, and the public API has been available since August 7, 2026, on Volcano Ark and BytePlus ModelArk. API resellers are also showcasing it.

Four key points that will transform your writing approach.

The duration. 30 seconds in a single take, no more stitching together 5-second clips. With video extension, you can stretch it to a minute.

The references. Up to 50 per generation: 30 images, 10 video clips, 10 audio tracks. The number is impressive but misleading, as I'll explain.

Audio, generated alongside the image, not added later. Lip-sync supports 11 languages, excluding French. Optimized languages are Chinese, English, Spanish, Indonesian, and Malay; Thai, Arabic, Portuguese, Vietnamese, Japanese, and Korean are supported. For a French ad, generate a silent take and add your voice in post-production. That's our default too, because once a voice is generated with the image, you can't tweak it without starting over.

Region-level editing, which allows you to change an element within a specific time frame without regenerating the entire clip.

As for resolution, be wary of what you read. The model claims native 4K, but the initial API routes cap at 480p and 720p. That's our case too: Seedance 2.5 outputs at a maximum of 720p, whereas Seedance 2.0 goes up to 1080p. If your deliverable requires native 1080p, the previous version is still your best bet. Check what your provider actually offers before building a workflow around it.

A quick note on benchmarks, as the model is too new to be ranked. Seedance 2.5 hasn't yet appeared in Artificial Analysis's text-to-video arena. Its predecessor ranks third with an Elo score of 1,224, behind Gemini Omni Flash (1,244) and MiniMax H3 (1,240), and ahead of Kling 3.0 Pro (1,111) and Veo 3.1 (1,093). In image-to-video with audio, Seedance 2.0 is top-ranked. This is the starting point for 2.5, not its final score, but it gives you a sense of the level.

The anatomy of a Seedance 2.5 prompt

The official formula consists of six blocks: subject, action, setting, visual style, camera movement, and audio. Only the first two are mandatory. In practice, for a 30-second take, all six are crucial: prompts that falter mid-way almost always lack a defined setting and style, leaving the model to decide them every second.

A working framework:

[GOAL] 30s vertical 9:16 product ad, single continuous take.
[REFERENCES] @Image 1 defines the bottle. Do not use its background.
[SUBJECT] ...
[TIMELINE] [0-6s] ... End state: ...
[CAMERA] ...
[LIGHT & LOOK] ...
[SOUND] ...
[LOCK] ...
[AVOID] ...

A longer prompt isn't a better prompt. We work within a character budget of 2,000 to 2,900, with a hard cap at 3,000 including spaces. Beyond that, you don't gain control; you add noise that the model interprets for you. Every sentence should alter the image. If you can remove it without changing the take, do so.

Final point, write your prompts in English, even for a French ad. The model is much more stable in English. Never translate brand names or text printed on the product.

Breaking down the 30 seconds: the rule of the end state

Divide your 30 seconds into intervals. One main action per interval, never two. And end each interval with its final image, the end state.

[0-6s] Hands lift the bottle from the marble counter, slow rotation to face camera.
End state: bottle upright, label fully readable, centered in frame, hands out of shot.

[6-14s] The dropper is pulled out; one drop falls onto the back of a hand.
End state: dropper held mid-air above the hand, single drop visible on skin.

Why it matters: the end image of an interval is what the next inherits. If you don't write it, the model starts from a fresh interpretation. That's exactly where objects appear out of nowhere and poses reset.

Second rule for breaking down: treat your intervals as time budgets, not frame-precise edit points. A gesture that takes four seconds in real life won't fit into a two-second interval. The model will speed it up, and a hand moving too fast is the first sign of an unusable rush.

Third rule, which we apply consistently in production: one state change per take. Closed to open, off to on, wrapped to unwrapped. Stack two, and you'll get a take where both transitions are half-baked.

Finally, always close your prompt with an identity lock. One line at the end of the prompt:

Same face, same hairstyle, same outfit, same body type for the entire video.

For product shots, lock similarly, but to a maximum of 5 attributes: silhouette, color, logo, label text, material. We've tried listing twelve, and the result is worse. Beyond five, the signal dilutes, and the model decides on its own what to preserve.

Linking references: one image, one subject, one exclusion

50 references is the theoretical max. The useful max is much lower: between 1 and 8 distinct subjects from images, 1 to 5 from video, reference clips of 5 to 10 seconds. Go above that, and coherence degrades instead of improving.

Name each reference by its rank, in the order you send them: @Image 1, @Image 2, @Video 1.

Give each an explicit role, and importantly, state what not to use. The effective phrasing: @Image 1 defines the bottle shape, glass texture and label typography. Do not use its background or lighting.Without the exclusion, you inherit the background, lighting, and sometimes an object next to it that you didn't need.

One reference, one subject. The phrasing "images 1 to 4 define the four characters" doesn't work; you must map them one by one. Useful exception: when multiple images show the same object from different angles, state it explicitly, or you'll end up with multiple copies of the object in the take. All four images define one single lamp. The output must contain exactly one lamp.

A point you'll encounter quickly: reference images containing humans are often blocked by content filters, including perfectly innocuous photos. The workaround is to describe the character in text, in a dedicated block, rather than providing an image. You lose fidelity but gain passage rate.

The audio syntax almost no one uses

Seedance 2.5 distinguishes four types of sound content by their punctuation, and mixing these in plain text is one of the most common causes of audio failure.

(...) for music: (slow acoustic guitar riff)

<... for sound effects: <a bicycle bell rings twice

{...} for dialogue: {Order ready, table six.}

【...】 for embedded subtitles: 【Day 1】

Before a non-English dialogue, specify the language and accent: Dialogue language: natural conversational American English.

The most useful case in ads: cutting the music. Writing "no music" frequently fails, and the model layers a track over it. The directive that holds:

[SOUND] Strictly only naturally occurring sound and foley, no music allowed.

We go further: we request zero speech, zero subtitles, zero on-screen text for nearly all rushes. Overlays and voice are added in post. A text generated in the image is a nuisance to mask out, and it's almost always misspelled.

Camera and emotion: name the movement, describe what's observable

"Cinematic" means nothing to the model. Standard terms, however, pass directly: push-in, pull-out, pan, orbit, high angle, low angle, handheld, whip pan.

For less common terms, translate into visible change rather than using them as-is. Instead of "rack focus," write: shift focus smoothly from the foreground leaves to the person standing behind them.

Same logic for emotions. "She's disappointed" results in a generic expression. What works is describing what you see: eyes lowering, jaw relaxing, hand placing the object down more slowly than necessary. Write the physical consequence, not the label.

And write your transitions instead of leaving them to chance. An undescribed junction is where an object can appear or disappear. → WHIP PAN RIGHT on her turn, smears to white, hard cut.

Fixing without regenerating everything

This is the model's most underused feature and the one that saves the most time when a rush is 90% there.

An edit prompt consists of four elements: which video is the source, what changes, over which time window, and crucially, what stays the same. The last point is the most important; it's the preservation clause.

[EDIT] Edit @Video 1. Only from 4s to 7s, change the cool blue light to warm orange.
[PRESERVE] Keep character identity, clothing, expression, position, motion,
room structure, camera movement, dialogue and ambience exactly as in @Video 1.

To extend a take, describe only what's new and demand continuity:

Extend the video naturally. [new content only]. Smooth motion continuity,
no hard cuts, nothing appears out of thin air.

Three locks to know before launching. In editing, the ratio is locked to the source, and the duration changes only by a few tenths of a second. In first image/last image generation, the ratio is locked to the first image. In extension, the ratio is locked to the source. So if you want to change format, it's a new generation, not an edit.

The 6 most common mistakes

Describing a photo instead of a narrative

The symptom is a 30-second take where nothing happens for 25 seconds. The cause is a prompt written in descriptive present, without intervals. Break it down and set your end states.

Stacking two actions in the same interval

The model picks one or merges them into a strange gesture. One main action per interval. If you have three to convey, extend the take or cut one action.

Misjudging the intervals

The gesture plays sped up, and hands blur. Recalculate the real-time of the gesture and give it its interval back.

Sending too many references

Beyond 8 distinct subjects, characters blend, and objects duplicate. Select your references scene by scene instead of sending them all at once.

Pasting a generic list of don'ts

Twenty prohibitions copied from another prompt dilute the three that mattered. Stick to 3 to 5 don'ts, all plausible for this specific take.

Rewriting the entire prompt after a failed rush

This is the most costly mistake because you lose the information. Identify the main flaw, modify only the concerned section, and relaunch. One variable at a time. And we have a rule to avoid burning the budget: after 2 relaunches on a clip that keeps drifting, we stop and rebrief the take instead of continuing to pay for generations.

A commented ad prompt

30 seconds, vertical format, product-focused, no voice. It's the format we deliver most often because voice and subtitles are added in post.

[GOAL] 30s single continuous take, vertical 9:16, product-focused UGC-style ad.

[REFERENCES] @Image 1 defines the serum bottle: shape, amber glass, white label
typography, gold dropper cap. Do not use its background or lighting.

[SUBJECT] A woman's hands only, no face in frame, on a pale marble bathroom counter,
soft morning window light from the left.

[TIMELINE]
[0-7s] Hands enter frame and lift the bottle from the counter, slow rotation until
the label faces camera. End state: bottle upright and centered, label fully readable,
hands holding it at the base.
[7-16s] The dropper cap is unscrewed and lifted out. End state: dropper held in the
right hand above the open bottle, one drop hanging at the tip.
[16-24s] One drop falls onto the back of the left hand. End state: single visible drop
on the skin, dropper out of frame.
[24-30s] Fingertips spread the serum in one slow circular motion. End state: skin
lightly glossy, hands resting flat on the counter, bottle standing beside them.

[CAMERA] Locked medium close-up, very slow push-in across the whole take.
No cuts, no whip pans.

[LIGHT & LOOK] Natural window light, soft shadows, realistic skin texture,
handheld phone-camera realism, no color grading, no lens flare.

[SOUND] Strictly only naturally occurring sound and foley, no music allowed.
<glass cap unscrewing> <a single drop landing on skin> <quiet room tone

[LOCK] Same bottle throughout: amber glass, gold dropper cap, white label with the
exact printed text from @Image 1, same proportions. The bottle at 30s must be the
same object as at 0s.

[AVOID] No spoken words, no lip movement, no on-screen text or subtitles,
no second bottle appearing, no fingers covering the label.

What holds this prompt together: a one-line goal, a reference with its exclusion, four intervals each with an end state, a single state change for the whole take (the bottle opening), a nominal lock on five attributes, and four don'ts that match the most likely errors here. About 1,800 characters. Under budget, and every line serves a purpose.

Using Seedance 2.5 in Hoox

The model has been available in Hoox since its launch, on Pro and Enterprise plans. You write your prompt, choose your duration between 4 and 30 seconds, second by second, and retrieve the rush.

Three useful details if you're comparing to other models on the platform. All six ratios are available, including the 9:16 you need for ads. You can attach up to 50 reference images, compared to 12 on Seedance 2.0 and 3 on Veo 3.1. And a Seedance 2.5 generation costs about a third less than a Seedance 2.0 in 1080p, for a maximum duration twice as long. The comparison isn't at equal resolution, since 2.5 caps at 720p, but for a rush intended for 9:16 social, that's rarely a deal-breaker.

The rest of the method doesn't change: the prompt in English, the timeline with end states, the identity lock at the end, overlays, and voice in post. If you're new to this, I've explained elsewhere how AI-generated UGC video production works.

FAQ

What's the structure of a Seedance 2.5 prompt?

Six blocks: subject, action, setting, visual style, camera movement, and audio. Only the subject and action are mandatory, but for a long take, defining the setting and style prevents the model from re-deciding them mid-way. Add a timeline broken into intervals with an end state for each, an identity lock at the close, and 3 to 5 don'ts.

How long can a Seedance 2.5 video last?

30 seconds in one pass. With video extension, you can extend up to about a minute in total. Beyond that, you assemble in post.

How many references can you use?

50 at the theoretical maximum: 30 images, 10 videos, and 10 audios. In practice, stay between 1 and 8 distinct subjects from images and 1 to 5 from videos. Beyond that, coherence degrades.

Does Seedance 2.5 support French lip-sync?

No. Lip-sync covers 11 languages, and French isn't one of them. For a French ad, generate the take without speech and add the voice in post. You also retain the ability to correct the script without regenerating the video.

How to remove music in a Seedance generation?

Writing "no music" isn't enough. Use an explicit directive in the sound block: Strictly only naturally occurring sound and foley, no music allowed. Then specify the sound effects you want to hear within angle brackets.

Should prompts be written in English?

Yes. The model is more stable in English, even for a video aimed at a French market. Don't translate brand names or text printed on the product.

Where to use Seedance 2.5?

Directly with ByteDance, through Dreamina, Jimeng, and Doubao for creators, and through the API Volcano Ark or BytePlus ModelArk for developers. The model is also available in Hoox on Pro and Enterprise plans, from 4 to 30 seconds and up to 50 reference images.

The prompt isn't the bottleneck

Writing better prompts increases your usable rush rate, and that's already a lot. But if you're producing ads, your limit is the number of concepts you release each week. A perfect prompt on an account delivering four creatives a month doesn't change your ROAS.

The above method becomes valuable when it's on repeat: one concept, one prompt, one controlled rush, one variant the next day.

That's exactly the work we do for the brands we support, with UGC-type videos generated and delivered in batches, at a cost per video we've detailed elsewhere. Seedance 2.5 is on the platform if you want to prompt yourself, and if you prefer us to handle the volume, this is the way.

Ready to create videos with AI?

Start for free and create your first videos in seconds.

Start for free