Seedance 2.5: Thirty Seconds, One Take, No Splicing
Seedance 2.5 doubles the single take to 30 seconds, takes up to fifty references and lets you edit a clip by the second. What it does, where it beats Kling 3.0 and Wan 3.0, and where it doesn't.

Fifteen seconds was always a strange length for a story. Long enough to set something up, too short to pay it off. So you'd generate the setup, generate the payoff, and spend the next hour making two clips pretend they were one: matching the light, nudging the character's jacket back to the right shade of green, hiding the seam with a cut nobody asked for.
The splice was the real job. The generation was the easy part.
Seedance 2.5 is ByteDance's answer to that, and the answer is blunt: make the take longer and let the model handle the cuts inside it. Thirty seconds of video with synchronized audio in a single pass, references by the dozen, and timestamp control so you can say what happens at second twelve instead of hoping it happens somewhere. It's live on each::labs now, and it changes what a "clip" is supposed to be.
What Seedance 2.5 Actually Is
ByteDance's Seed team launched Seedance 2.5 on July 31, 2026, building on the unified audio-video joint-generation architecture of Seedance 2.0. The team's framing is plain: users stopped wanting a clip and started wanting a finished piece of work. So 2.5 focuses on long-form storytelling, multimodal reference and editing.
The headline numbers: single-pass generation goes from 15 seconds to 30, the reference ceiling jumps to 30 images, 10 video clips and 10 audio clips, and editing gets timestamp-level control. But the numbers aren't the interesting part. What matters is that ByteDance stopped treating thirty seconds as one long moment and started treating it as a sequence. According to the launch post, the model organizes several logically connected shots within a single generation, so a story moves through setup, development, a turn and a resolution rather than one moment stretched thin.

Thirty Seconds Changes the Unit of Work
ByteDance leads with a good example. A singer adjusts her earpiece in a dressing room, a staff member tells her it's time, she walks a backstage corridor greeting dancers, someone hands her a microphone, and she steps onto an arena stage while the camera pulls out to the crowd. One take. Not a montage of four generations.
That's a different unit of work. With a 15-second ceiling you plan in fragments and spend your effort on continuity. With 30 seconds and real transitions inside the take, you plan in scenes.
Extension pushes it further. You can feed an existing output back in and ask the model to continue: same characters, same environment, same pacing, another thirty seconds. ByteDance says this lets users reach videos lasting several minutes while keeping main subjects consistent. Treat that as the ceiling of a careful workflow, not a one-click promise.

Show It Fifty Ways: References That Actually Scale
The Reference to Video variant is where Seedance 2.5 stops being a longer Seedance 2.0 and becomes a different tool. Thirty images, ten videos, ten audio clips. You call them in the prompt by token (@Image1, @Video1, @Audio1) and tell the model what each one is for.
ByteDance's own demo is a classical concert built from eighteen images: one for the venue, one each for the pianist, cello and violin, one the lead vocalist "must strictly follow", five for the orchestra, four for the choir, four for the audience seating. That's not a gimmick. It's how you'd brief a crew.
The quieter upgrade is clay render referencing. You block a shot with textureless 3D models (spatial layout, poses, motion paths, camera angles) and the model renders on top of that structure, using the geometry to place light direction, color temperature and shadows. If you previsualize in a 3D tool, this is the feature: a gray-box animatic becomes a finished-looking shot without describing the camera move in words that never quite land.
Don't confuse ceiling with target, though. On each::labs, reference videos are capped at 30 seconds combined and reference audio at 30 seconds combined, so ten clips means ten short clips. And more references only help when they agree. A set that fights itself on style or lighting gives the model a reconciliation problem, and the consistency you came for is the first thing it gives up. (More on that in our guide to character consistency across AI video shots.)
Direct by the Second, Edit After the Fact
Timestamp control works in two directions. Before generation, you write the prompt in time segments and the model follows the beat sheet. After generation, you can target a specific stretch and change a character, an action or a plot element while the footage on either side stays intact.
A text prompt in the style ByteDance publishes looks like this:
16:9, cinematic, single continuous take, no cuts. 0 to 6s: close-up of a lighthouse keeper winding a brass clock, storm audible outside. 6 to 14s: the camera follows him up a spiral staircase, lantern swinging, rain hammering the glass. 14 to 24s: he reaches the lamp room and lights the wick; the beam sweeps out over a black sea. 24 to 30s: slow pull back through the window to a wide shot of the lighthouse in the storm. He says quietly, "Still here."
Spoken lines go in double quotes for the best speech. Editing works through wording on the reference variant. Phrase the request with add, remove or replace and it switches to an edit of your input video. Say extend or continue and it becomes an extension. ByteDance also lists green screen editing, where the subject's clothes, hair and gait respond to the new background, and camera perspective editing, which keeps the action and re-plans the move.
A take that's mostly right is no longer a take you throw away.

Three Variants, Three Ways In
There are three Seedance 2.5 models on each::labs, and they share one spine: 4 to 30 seconds (or auto), 480p, 720p or 1080p output, and synchronized audio on by default.
Text to Video is the director's-brief mode. Aspect ratios run from 21:9 to 9:16 plus auto, prompts can run to roughly a thousand words, and English sits alongside Spanish, Indonesian, Portuguese, Japanese, Malay, Thai, Arabic, Vietnamese and Korean.
Image to Video anchors the frame. Pass a first frame, or a first and a last, and the model solves the path between them. If the two images disagree on aspect ratio, the first one wins and the last is center-cropped, so design them as a pair.
Reference to Video anchors identity, style and sound, and it's the only one that edits and extends. If your project has a recurring character, a brand world or a previs pass, start here.
Seedance 2.5 vs Kling 3.0: Which Should You Pick?
"Kling vs Seedance" is the comparison everyone searches, and it deserves an answer instead of a scoreboard. Both families are live on each::labs and both generate native audio. They're built around different ideas of control.
Kling 3.0 runs 3 to 15 seconds and gives you explicit structure. Its multi-prompt mode splits a video into storyboard shots, each with its own prompt and duration. Its Elements system lets you register a character from a frontal image plus up to three angle references and call it as @Element1. You can attach up to two voice IDs. And there's a dedicated Kling 3.0 4K variant that outputs 4K in a single step, plus Turbo and Motion Control models in the same family.
Seedance 2.5 bets on length and reference volume instead: twice the take, a far bigger reference budget, timestamp beats, and editing and extension of existing footage. Output tops out at 1080p.
So pick by the shape of the job. If you're cutting short, punchy shots that must be delivered at 4K, or you want discrete storyboard control with a registered character, Kling 3.0 is the cleaner fit. On each::labs it also returns faster, with a median runtime of about two minutes on Kling 3.0 Pro against about four on Seedance 2.5. If the piece is one continuous scene, needs a crowd of consistent characters, starts from a clay-render previs, or needs a fix in the middle of an otherwise good take, Seedance 2.5 is the one. Our Seedance 2.0 vs Kling 3.0 comparison covers the earlier matchup; 2.5 mostly changes the length and reference math.

Seedance 2.5 vs Seedance 2.0 and Wan 3.0
Against its predecessor, the jump is concrete. Seedance 2.0 on each::labs runs 4 to 15 seconds and its reference variant takes 9 images, 3 videos and 3 audio files, twelve files in all, with audio only allowed alongside an image or video. Seedance 2.5 doubles the duration, raises the reference ceiling to fifty files, accepts audio-only input, and roughly doubles the prompt languages. Seedance 2.0 still has reasons to stay in your stack: a seed parameter for repeatable runs (2.5 doesn't expose one), Fast and Mini variants for quick iteration, and a 4K option on its reference model.
Wan 3.0 is the interesting rival on length. Alibaba's model also generates up to 30 seconds, at 480P to 1080P with audio, and its reference variant does something Seedance doesn't: alongside up to 10 images, 5 videos and 5 audio clips, it accepts a file or a public web page as a reference. Where Seedance 2.5 pulls ahead is reference volume (30 seconds of reference video and audio against Wan's 15), clay render and the editing and extension workflow.
Using the Seedance 2.5 API on each::labs
All three variants run through the same endpoint as every other model on each::labs. Switching between them is a change of model string and input shape:
curl -X POST \
-H "Authorization: Bearer $EACHLABS_API_KEY" \
-H "Content-Type: application/json" \
--data '{
"model": "bytedance-seedance-2-5-reference-to-video",
"input": {
"prompt": "@Image1 is the courier, @Image2 the market. 0 to 15s: she cycles through the stalls. 15 to 30s: wide shot as she rides into the evening crowd.",
"image_urls": ["https://.../courier.png", "https://.../street.png"],
"duration": "30",
"resolution": "1080p",
"aspect_ratio": "16:9",
"generate_audio": true
},
"webhook_url": ""
}' \
https://api.eachlabs.ai/v1/prediction/For Image to Video, point the model at bytedance-seedance-2-5-image-to-video and pass image_url, optionally with end_image_url. You get a prediction ID back, not a video, so poll it or use a webhook; with a median runtime near four minutes, design for the wait. If Seedance 2.5 is one step in a longer pipeline (a still generated first, a voice track after), each::labs flows let you chain those steps in one place.
Wrapping Up
Most video models still hand you moments and leave the story to you. Seedance 2.5 hands you scenes: thirty seconds with cuts inside, a reference budget big enough to cast an orchestra, and timestamp edits that rescue the take you nearly threw away. It isn't the fastest model on each::labs and it doesn't output 4K, but for continuous storytelling it's the one that removes the most splicing from your week.
Try Seedance 2.5 on each::labs with one scene you've been building out of fragments. See how much of the editing it eats.
Frequently Asked Questions
Seedance 2.5 vs Kling: which should you pick?
Look at the shape of the shot list. Short, discrete shots, 4K delivery or a registered character with assigned voices point to Kling 3.0, which runs 3 to 15 seconds with storyboard-style multi-prompt control. One continuous scene, many characters to keep consistent, clay-render previs or mid-take edits point to Seedance 2.5, with its 30-second takes and up to fifty references. Many projects will use both.
How long can a Seedance 2.5 video be?
Each generation runs 4 to 30 seconds on each::labs, or auto, where the model picks a length from your prompt. Beyond that, the Reference to Video variant can extend an existing clip when you phrase the prompt as extend or continue, and ByteDance says chaining extensions can reach several minutes while keeping characters and environments consistent.
Does Seedance 2.5 generate audio?
Yes, by default. Voice, sound effects, ambience and music come out in the same pass as the picture, and you can switch audio off for a silent clip. Put dialogue in double quotes in the prompt for the cleanest speech.