Seedance 2.5 in Genosai — 30-second video with sound and prompt-based editing

Seedance 2.5 is ByteDance's new flagship video model, available in Genosai. It generates a coherent clip up to 30 seconds long in a single pass, creates sound and lip sync together with the visuals, accepts up to 50 references per generation, and can edit finished footage by timecode without regenerating the whole video. In Genosai the model runs online — no developer account or API setup required.

Updated: August 10, 2026

Contents

What is Seedance 2.5

Seedance 2.5 is the new flagship video model from ByteDance, the creators of TikTok. It comes out of the Seed lab — the same team behind the Seedream image model and the Seed Audio voice line. The model was announced on June 23, 2026 at the Volcano Engine FORCE conference in Beijing, reached general availability at the end of July, and opened up to third-party platforms in August. It is the freshest flagship in the family: the previous version, Seedance 2.0, topped the independent Artificial Analysis image-to-video leaderboard among models with sound for a stretch, ahead of Veo 3.1 and Kling 3.0 — and version 2.5 builds on that foundation.

The real edge of Seedance 2.5 is not just image quality. The model takes over what used to be an editor's job: a 30-second studio-grade clip is generated in a single pass, with sound and lip sync, while dozens of references keep the character, style and camera movement consistent through the whole scene. On top of that comes an editing toolset — timecode edits, background replacement, re-shot camera moves — that works on finished footage without full regeneration.

In Genosai the model runs online: no developer account, API keys or infrastructure of your own. You open the video studio in the browser, describe the scene, optionally upload references — and get a finished clip with sound.

Capabilities

Seedance 2.5 covers the full production cycle of a short video — from idea to final edit. Here is what the model does in practice.

30 seconds in a single pass

Most competitors are capped at 8–15 seconds per generation: Veo 3.1 at 8 seconds, Kling 3.0 and Seedance 2.0 at up to 15. Seedance 2.5 natively produces a coherent scene up to 30 seconds long with no seams between segments — no stitching required. There is also an extension feature: a finished clip can be expanded up to two more times, adding up to 30 seconds of new footage per pass, so the final video can grow to roughly a minute and a half.

Up to 50 references at once

A single generation accepts up to 30 images, 10 videos and 10 audio files. You can simultaneously define the character's appearance with a photo, the camera trajectory with a video reference, the soundtrack with an audio file and the scene style with another image — and the model keeps all of these consistent through the entire clip, not just the first seconds. This is the key feature for branded content, where the logo, character and product need to stay recognizable for the full 30 seconds.

Native sound and lip sync

Voice, music and sound effects are generated in the same pass as the video, and lip sync works in more than 10 languages. There is no separate audio step in an editing app — a short talking scene comes out complete in one generation. For multilingual campaigns this means one video localized for different markets directly in the frame and in the voiceover.

Editing without reshoots

The second big block of this release is editing finished footage. The model accepts timecode edits ("replace the background from second 12 to second 18"), swaps backgrounds green-screen-style while adapting light and physics on the subject itself, re-shoots the camera movement on an existing clip, and makes localized edits to a region of the frame based on a reference — all without regenerating the whole video.

3D staging and formats

White-model control (informally, clay-render): you upload a simple untextured 3D scene, set the poses and camera path, and the model dresses it in realistic light, materials and motion. This is handy for complex camera choreography that is hard to describe in words. Supported aspect ratios are 16:9, 9:16, 1:1, 4:3, 3:4 and 21:9 — covering YouTube, Reels and Shorts, and social feeds alike.

Examples and use cases

Seedance 2.5 shines in tasks that used to require editing. Product ads and UGC: the product, background and style references go into a single prompt, and A/B variants are generated quickly and cheaply at draft resolution. Brand videos and stings: logo animations and match-cut edits synced to music, with shot changes landing on the beat.

Educational content benefits from the scene length: a process is explained in a 30-second continuous shot instead of a series of static slides. Agencies can show a client a living version of a concept before any shooting happens. And mini-stories with a persistent character — say, a vertical clip where the hero moves through several locations — rely on references: the character's look is pinned by a photo and does not drift from scene to scene.

How to use it in Genosai

Running Seedance 2.5 in Genosai requires no setup: everything happens in the browser. The path from idea to finished clip looks like this.

  1. Sign in to Genosai and open the video studio.
  2. Pick Seedance 2.5 from the model list.
  3. Describe the scene: action, camera movement, light and sound.
  4. Optionally upload references — character and style images, a camera-path video, an audio track.
  5. Set the duration, resolution and aspect ratio, then start the generation.
  6. Download the clip — or refine the prompt, extend the clip, or apply timecode edits.

Prompts

A few rules noticeably improve results. Start with the subject and setting, then describe the action over time: for clips longer than 8 seconds, break the scene into timecodes (0–3s, 3–6s and so on) — the model holds structure much better when you define it yourself. Describe the camera with concrete terms — dolly in, orbit, handheld, static wide shot — rather than a vague "beautifully shot". If you upload several references, state each one's role explicitly: @Image1 is the character's face, @Image2 is the background style, @Video1 is the camera path. Here are real examples: the prompt itself and the clip Seedance 2.5 generated from it in Genosai.

A brand logo video with match-cut editing, 1:1, 8 seconds, no references:

A fast-paced match-cut commercial. A glowing glass sphere stays locked in the exact center of frame while the background changes on every cut: a bustling city street at night, a quiet mountain sunrise, a neon-lit studio, ocean waves at golden hour. The sphere stays razor-sharp throughout; backgrounds carry motion blur and shift on quick rhythmic cuts. On the final cut the sphere shatters into light particles that reform into a clean, minimal abstract logotype shape on a black background. Dynamic rhythmic editing, realistic glass refraction, strong visual pacing, silent, no dialogue, no music, no text, no watermarks.

Match-cut logo, 1:1, 8 seconds — generated in Genosai.io with Seedance 2.5

A sound and lip-sync demo in Russian speech, vertical 9:16, 8 seconds, no references:

A confident young woman (AI-generated, fully fictional character, not a real person) speaks directly to camera in a bright minimal home studio, soft ring-light glow, casual oversized sweater. She leans slightly toward the lens with an enthusiastic expression and says in Russian: "Смотри — тридцать секунд одним кадром, и герой не меняется." Natural hand gesture, warm smile, direct eye contact with the camera, lips synced precisely to the Russian speech. Vertical smartphone-style framing, natural handheld micro-movement, realistic skin texture, authentic soft lighting, upbeat casual tone, no on-screen text, no watermarks, no music.

Sound and lip sync in Russian, 9:16, 8 seconds — generated in Genosai.io with Seedance 2.5

A mini-story with a persistent character, vertical 9:16, 20 seconds, character reference @Image1 and location reference @Image2:

One continuous vertical shot following a young woman in a beige coat (reference @Image1) as she walks through a quiet European old-town street (reference @Image2) at golden hour. 0-5s: she steps out of a wooden door, pauses, looks up with a faint smile; ambient street sounds, distant church bells. 5-12s: she walks forward, camera tracking from the side, passing a flower stall and a street musician playing violin, warm light filtering through the buildings. 12-20s: she stops at a small cafe table, sits down, camera slowly pushes in on her face. Naturalistic handheld camera feel, warm color grade, consistent character appearance throughout, no text overlay.

Generation cost

Genosai bills per second of generated video, and the rate depends on resolution: a 480p draft costs several times less than a final 720p render, so it pays to iterate on composition and pacing with cheap drafts and rebuild the final version in high quality. Reference-video modes and editing operations are billed separately from generation from scratch. The exact cost is shown in the studio before each generation starts, and current credit packages are listed on the Pricing page. Starter credits after sign-up let you try the model for free.

Comparison

Within its own family: Seedance 2.0 remains a solid pick for short clips up to 15 seconds that do not need dozens of references or editing features — it is cheaper and battle-tested. Reach for Seedance 2.5 when you need a long coherent scene, character and product consistency across the whole clip, or edits to finished footage.

Against competitors: Veo 3.1 Quality is still stronger for short close-ups with a spoken line — if you need one striking 8-second clip with expressive facial acting, start there. Kling 3.0 is a predictable option for clips up to 15 seconds. But if the task is a storyboard inside a single shot — a video where the logo, character and product must stay recognizable for the full 30 seconds — that is Seedance 2.5 territory: none of the competitors combine this scene length with this many references and prompt-based editing.

Limitations and tips

Moderation is strict: you cannot upload real faces — selfies, portraits, photos of celebrities — as appearance references, and you cannot request copyrighted characters. Illustrations, anime styles and fully fictional AI faces are fine. At the same time, the model is excellent at creating photorealistic fictional people from scratch — the restriction targets cloning real people's likenesses, not realism as such. Graphic violence and adult content are blocked in every mode.

Practical tips: start with a simple prompt without references to get a feel for the model, and run drafts at 480p. For branded content, bring references from the start — logo, product, character: without them the model invents its own style, with them it keeps consistency through the whole clip. For long scenes, always lay out the action by timecode. And keep in mind the release is very fresh — parameters and limits may still be refined, so trust the actual results you see in the studio.

FAQ

What is Seedance 2.5 and who developed it?

Seedance 2.5 is the flagship video model from ByteDance, announced in summer 2026 as the successor to Seedance 2.0. Its headline features are 30-second clips in a single pass, up to 50 references per generation and prompt-based editing of finished footage. In Genosai it is available online without any API setup.

How is Seedance 2.5 different from Seedance 2.0?

Twice the single-pass duration (30 seconds versus 15), far more references (up to 30 images, 10 videos and 10 audio files versus 9, 3 and 3), and a new editing toolset: timecode edits, background replacement, re-shot camera movement and clip extension.

How long can a Seedance 2.5 video be?

Up to 30 seconds in a single pass with no stitching between segments. The extension feature adds up to 30 seconds per pass and can be applied up to two times, so the final clip can grow to roughly a minute and a half.

Does Seedance 2.5 generate sound and lip sync?

Yes. Voice, music and sound effects are created in the same pass as the video, and lip sync works in more than 10 languages. There is no need to add audio separately in an editor.

How many references does Seedance 2.5 accept?

Up to 50 files per generation: up to 30 images, 10 videos and 10 audio files. You can simultaneously define a character's look with a photo, the camera path with a video reference, the soundtrack with an audio file and the scene style with another image.

Can I edit a finished clip?

Yes, this is one of the release's headline features: timecode edits, background replacement with light adaptation on the subject, new camera movement on finished footage and localized edits of a frame region by reference — all without regenerating the whole video.

What are the content restrictions?

Moderation is strict: you cannot upload photos of real people as appearance references or request copyrighted characters. The model freely creates photorealistic fictional people from scratch — the restriction targets cloning real people's likenesses, not realism itself.

Try Seedance 2.5 on Genosai