Seedance 2.5 is live on Masonry as seedance-2-5, with text-to-video, image-to-video, and reference-to-video, which is new for this version. It is the successor to the route covered in our Seedance 2.0 prompt guide.
Multi-shot generation is the claim every version of Seedance has carried. This article treats it as a claim to be measured rather than repeated: three clips generated, downloaded, watched, and run through a scene detector, with the result reported whether or not it flattered the model.
Evidence boundary: this article rests on six clips generated on September 4, 2026 across two model families, three of them on seedance-2-5. Each is a single run of a single prompt. There were no repeated trials, no held-out prompt set, no standardized queue timing, no reliability rate, and no leaderboard position. File metadata came from ffprobe on the downloaded files, and cut counts came from ffmpeg scene detection at threshold 0.4 — a detector setting, not an absolute truth about what a human would call a cut. Every reference and source still used here was generated on Masonry for this article, so no real individual's likeness was submitted to any route. Provider rates were read September 4, 2026 and change without notice.
Quick answer
- Route:
seedance-2-5, listing text-to-video, image-to-video, and reference-to-video. - Measured multi-shot: the text-to-video clip contains two detected hard cuts, at 2.625 s and 4.125 s, which is three shots inside a 5.041667-second file.
- Measured control: the same detector at the same threshold finds zero cuts in the other five clips from this launch. They are single continuous takes.
- Measured output: all three clips requested at
720preturned 1280×720, 24 fps, H.264 with an AAC stereo track at 32 kHz. - Measured duration: 5.041667 s, 5.041667 s, and 5.056 s against a requested duration of 5.
- Provider pricing, checked September 4, 2026: token-priced at $21.40 per million tokens (480p and 720p) and $23.40 per million (1080p). Audio does not change the cost.
- Reference policy: a generated portrait was refused with HTTP 422. Object-only references were accepted.
- Not established here: durations beyond 5 seconds, 1080p or 480p output, latency, reliability, prompt adherence rate, or a ranking against any other model.
The multi-shot claim, measured
Multi-shot output is easy to assert and easy to check. The text-to-video prompt asked for three shots inside five seconds: a cafe shutter rolling up at dawn, a wide interior, then a hard cut to a close-up of ice dropping into cold brew.
| Reading | Result | Production implication |
|---|---|---|
| Detected cuts in this clip | 2, at 2.625 s and 4.125 s | Three shots arrived inside a five-second request without a video editor. |
| Detected cuts in the other five launch clips | 0 | The detector is not simply firing on motion; the contrast is real. |
| Shots matching the written brief | 2 of 3 | The model cuts on request; it does not necessarily cut to the shot you named. |
| Shot lengths | Roughly 2.6 s, 1.5 s, 0.9 s | Later shots get squeezed; write the important beat first. |
Watching it changes the reading. The first shot rolls the shutter up and reveals the interior behind it in one continuous move, so the brief's first two beats arrive as one. The second shot is the ice-and-cold-brew close-up the brief put third, landing a beat early. The third shot was not in the brief at all: a matte-black bottle standing on the counter with espresso machines behind it. Two of the three written shots landed, the wide interior never got a shot of its own, and the model spent the last beat on an idea it supplied itself. The cutting behaviour is real; the shot-by-shot adherence is partial.
The control matters as much as the measurement. Five other clips from the same session, run through the same detector at the same threshold, return zero cuts. If the detector were noise-triggered, they would not. That is why this article leads with a cut count rather than an adjective.
One clip is still one clip. It establishes that the route can cut; it does not establish how often it cuts where you asked.
Image-to-video from a single frame
The clip stays faithful to the source frame through a subtle orbit with modest motion. There is no dramatic reinvention of the product and no visible drift away from the still. Zero cuts detected, as expected for a single-frame animation.
Modest motion is a result, not a verdict. Whether it is the right amount depends on the placement: it is well suited to a product page loop and underpowered for a scroll-stopping social hook. If exact-product fidelity is the requirement, the discipline in our product-ad model test still applies — approve the source, reject any frame that exceeds it.
Reference-to-video, the new mode
Reference-to-video is what 2.5 adds over the previous route. This run submitted two product references (a bottle and a sneaker) and asked for one scene containing both.
Both product references arrive in one scene, and the camera executes a push-in. Zero cuts were detected. A reference shot that comes back as one continuous take is what you want when the job is to show a product rather than to tell a story.
The reference policy is a routing decision
The first attempt at a reference clip used a portrait, generated on Masonry for this article rather than photographed. It was refused:
| Route | Result on the identical portrait reference |
|---|---|
bytedance/seedance-2.5/reference-to-video | HTTP 422, content_policy_violation / partner_validation_failed — the images or videos provided may contain likenesses of real people or other private information that cannot be processed. |
minimax/h3-max/reference-to-video | Accepted; returned a clip carrying that person into a new scene. |
Re-running the Seedance request with object-only references succeeded, and produced the clip above. Reference mode works here; what failed was the portrait. The MiniMax H3 Max family accepted the same file on the same day.
Best for planning around this: send person-carrying reference shots to H3 Max, keep product-only reference shots on either family, and decide the routing before a deadline rather than during one.
The refused file was itself a generation, so this records what each filter does when handed a face rather than a ruling on any specific real person. It is also one observation per route on one image. Provider policy changes, and rights and consent for any likeness you upload remain yours to hold regardless of which route accepts the file.
Pricing, and the number worth seeing first
Seedance 2.5 is token-priced rather than billed at a flat per-second rate.
| Resolution | Token rate, checked September 4, 2026 | Provider per-second equivalent | Production implication |
|---|---|---|---|
| 480p | $21.40 per million tokens | $0.2205/s | The cheapest place to iterate on framing and action. |
| 720p | $21.40 per million tokens | $0.4730/s | What every clip in this article was generated at. |
| 1080p | $23.40 per million tokens | $1.164/s | Roughly 2.5× the 720p rate; reserve it for the final render. |
Enabling audio does not change the cost, so there is no reason to generate silent drafts to save money.
A 30-second generation at 1080p is roughly $34 of provider cost. Settle the shot at 480p or 720p first, then spend the 1080p budget once. These are provider rates; what a Masonry generation is metered at is set by the platform pricing table, not by this article.
A seed, and no expanded prompt
The responses recorded here returned a seed and no expanded prompt. MiniMax H3 Max, tested the same day, does the reverse: a long structured expanded prompt describing what it built, and no seed on text-to-video.
That is a real workflow difference. Here, you pin a result by holding the seed and changing one variable at a time. There, you diagnose a surprising result by reading back what the model decided you asked for. Neither substitutes for the other, and assuming both fields exist is how a rerun script breaks.
Not established here
- Durations beyond the 5-second requests measured here, including the stated 30-second ceiling.
- 480p or 1080p output. Every clip in this article is 720p.
- How often the model cuts to the shots you named, rather than simply cutting.
- Latency, queue time, throughput, or failure rate.
- Any ranking against Seedance 2.0, MiniMax H3 Max, Kling, or Veo.
- Any Masonry credit price, margin, or per-generation cost to you.
Bottom line
The multi-shot claim survives contact with a measurement: two hard cuts at 2.625 s and 4.125 s in one five-second clip, against zero cuts in five control clips run through the same detector. Shot-by-shot adherence was partial (two of three), and reference-to-video worked on product stills but refused the portrait reference.
Write the shots as shots, put the beat that matters first because later shots get squeezed, hold the seed when you want to iterate, settle the framing at 720p before spending 1080p money, and measure the delivered file rather than trusting the requested duration. For prompt structure that predates this version but still applies, the Seedance 2.0 prompt guide has 16 worked examples. For the other family launched the same day, see MiniMax H3 Max and H3 Max Turbo. Browse the rest of the lineup on Masonry's model hub.


