MiniMax H3 Max and MiniMax H3 Max Turbo are new video routes on Masonry, under the keys minimax-h3-max and minimax-h3-max-turbo.
The useful part of a launch post is not the adjective list. It is three clips that were generated, downloaded, measured, and watched, plus the handful of behaviours that will actually change how you queue work: what 768P returns in pixels, what the response hands back for reproducibility, and which route will take a portrait reference at all.
Evidence boundary: this article rests on six clips generated on September 4, 2026 across two model families, three of them on the MiniMax H3 routes shown below. Each is a single run of a single prompt. There were no repeated trials, no held-out prompt set, no standardized queue timing, no reliability rate, and no leaderboard position. File metadata came from ffprobe on the downloaded files; the cut counts quoted later came from ffmpeg scene detection at threshold 0.4. Every reference and source still used here was generated on Masonry for this article, so no real individual's likeness was submitted to any route. Provider rates were read September 4, 2026 and change without notice. Nothing here establishes how either route behaves on your prompts, your references, or your acceptance bar.
Quick answer
- New routes:
minimax-h3-maxlists text-to-video, image-to-video, and reference-to-video with up to 4 reference images.minimax-h3-max-turbolists text-to-video and image-to-video. - Measured output: all three clips requested at
768Pand 16:9 returned 1344×768, 24 fps, H.264 with an AAC stereo track at 32 kHz. - Measured duration: 5.184 seconds on every clip, against a requested duration of 5.
- Standard provider rates, checked September 4, 2026: H3 Max at $0.05/s (480P) and $0.08/s (768P); Turbo at $0.025/s and $0.04/s. Turbo is exactly half the H3 Max rate at every listed resolution.
- Reference behaviour: H3 Max reference-to-video accepted a portrait reference and carried that person into a new scene. Seedance 2.5 refused the identical file, reporting a possible real-person likeness.
- Reproducibility field: these H3 Max text-to-video responses returned an expanded prompt and no seed.
- Not established here: output quality ranking between the two routes, image-to-video behaviour on either route, latency, reliability, 480P output, durations other than 5, or any Masonry credit price.
The two routes and what they cost
| Item | Checked September 4, 2026 | Production implication |
|---|---|---|
| H3 Max key | minimax-h3-max | Pin the exact key in automation; the Turbo key is a different route, not a flag. |
| H3 Max modes | Text-to-video, image-to-video, reference-to-video | Reference mode is the reason to reach for H3 Max over Turbo. |
| Turbo key | minimax-h3-max-turbo | No reference mode is listed; do not plan a reference shot on Turbo. |
| H3 Max rate | $0.05/s at 480P, $0.08/s at 768P | A five-second 768P clip is about $0.40 of provider cost. |
| Turbo rate | $0.025/s at 480P, $0.04/s at 768P | Exactly half the H3 Max rate at both resolutions; the same five-second clip is about $0.20. |
768P in pixels | 1344×768 on all three clips | Do not size a placement from the label; measure the delivered frame. |
| Delivered duration | 5.184 s against a request of 5 | Do not build a hard-cut edit that assumes exactly 5.000 s. |
| Audio | AAC stereo, 32 kHz, on every clip | Audio arrives whether or not your edit wants it; listen before publishing. |
Provider rates are what fal charges for the generation. What a Masonry generation is metered at is set by the platform pricing table, not by this article.
Three shipped clips
Text-to-video on H3 Max
The brief asked for a modern coffee bar in morning light, a barista sliding a matte-black cold-brew bottle across a steel counter, and a slow dolly-in. The returned clip performs a genuine slow dolly-in, and every named prop is present. Adherence to a written camera move is the specific thing worth checking here, because a model that ignores camera direction turns every shot into a lottery.
Watch it through rather than judging the thumbnail: hands on the bottle, the bottle's contact with the counter, the counter edge as the camera moves, and background geometry under parallax.
The same prompt on H3 Max Turbo
Same prompt, same scene structure, different composition and a noticeably darker grade. Watching both, Turbo reads as comparable rather than obviously worse.
That sentence is the whole claim, and it is deliberately weak. One pair of clips is not a benchmark. Two single runs on one prompt cannot establish which route produces better output, how often either fails, or whether the gap widens on harder briefs. If the decision matters to your budget, run your own prompt set on both keys, keep the prompt, resolution, duration, and acceptance sheet fixed, and count accepted clips rather than impressions.
Recommended: treat the halved rate as the reason to test Turbo on your own material, not as evidence that Turbo is worse.
Reference-to-video on H3 Max, two references
The reference run sent two images: a portrait of a woman and a product still of a specific bottle. Both were generated on Masonry for this article rather than photographed, so no real individual's likeness was submitted to either route. That matters for how you read the refusal further down.
The result carried both references into a new scene. The person arrives with face, bob haircut, and charcoal turtleneck intact, the bottle arrives as the specific bottle rather than a generic one, and the requested lift-and-turn action is executed.
That is one run on one pair of references. It does not measure identity drift across many generations, it does not test how the model handles a reference whose lighting fights the target scene, and it does not tell you whether label text on a real product would survive. For commercial work where the product on screen has to be the product you ship, keep the product-ad model test discipline: approve the source, reject any frame that exceeds it, and log every rejection.
Behaviour worth knowing before you queue a batch
One family refused the portrait reference, the other did not
The same generated portrait was submitted to both families on September 4, 2026:
| Route | Result on the identical image |
|---|---|
minimax/h3-max/reference-to-video | Accepted; returned the clip shown above. |
bytedance/seedance-2.5/reference-to-video | HTTP 422, content_policy_violation / partner_validation_failed — the images or videos provided may contain likenesses of real people or other private information that cannot be processed. |
Re-running the Seedance request with object-only references (a bottle and a sneaker) succeeded. So this is a routing decision, not a defect: build a person-carrying reference shot on H3 Max, and product-only reference shots on either family. Plan the fallback before a deadline, not during one.
The refused file was a generated portrait, so what this records is what each filter does when handed a face, not a ruling on any specific real person. It is also one observation per route on one image. Provider policy changes, and holding rights and consent for a likeness remains your responsibility regardless of which route accepts the upload.
H3 Max rewrites your prompt and shows you the rewrite
The H3 Max response carries an expanded_prompt field. The one returned for the barista shot opens like this:
integrated_multimodal_description: [Shot 1] Cinematic, live-action, medium shot. The scene opens in a modern specialty coffee shop during the early morning…
It goes on to name the shot, the lighting mix, the wardrobe, and the camera move — including "slow dolly-in (push in) with small amplitude at slow speed". That is genuinely useful: when a result surprises you, the expanded prompt shows what the model decided you asked for, which is usually where the disagreement lives.
The trade is that these text-to-video responses returned no seed. Seedance 2.5 does the reverse, returning a seed and no expanded prompt. Practical consequence: on Seedance, reproducibility runs through the seed; on H3 Max, you debug by reading the expansion.
A resolution label is not a resolution
768P on H3 Max renders 1344×768. 720p on Seedance 2.5 renders 1280×720. Two similar labels, two different frames, two different crops for a placement that expects one of them. Read the delivered file.
Requested duration is a target
Every H3 Max clip here measured 5.184 seconds against a requested duration of 5, all three identically. Seedance landed at 5.04 to 5.06 seconds on the same request. If your edit assumes exactly 5.000 seconds, the overshoot lands on your timeline, not on the provider's.
Transient provider errors happen
One h3-max-turbo/text-to-video submission returned HTTP 504 downstream_service_unavailable and succeeded on a plain retry. One 504 in a handful of submissions is neither a reliability problem nor a reliability guarantee; it is a reason to make retry the default in any batch script rather than treating the first error as a verdict.
Choosing between the two keys
Choose this when the shot needs a person or a specific product carried in from a reference image: H3 Max is the only key in this family that lists reference-to-video.
Best for a cost-sensitive draft pass: Turbo, at exactly half the H3 Max rate, for iterating on framing and action before committing the final render to whichever route your own test prefers.
The honest version of the comparison is that the rates are known and the quality gap is not. The rate difference is documented by the provider; the quality difference would need a prompt set, repeated runs, and a fixed acceptance sheet that this article does not have.
Not established here
- Which of the two routes produces better output. One paired run cannot answer that.
- Image-to-video behaviour on either key. Both list the capability; neither was exercised for this article.
- 480P output, durations other than 5, or aspect ratios other than 16:9.
- Latency, queue time, throughput, retry rate, or failure rate under load.
- Identity or product fidelity across repeated reference runs.
- Any Masonry credit price, margin, or per-generation cost to you. Provider rates are quoted; platform metering is set elsewhere.
Bottom line
Two new keys, one of them at half the rate of the other, and one of them able to carry a person from a portrait reference into a new scene. The three clips here show the routes working; they do not rank them.
Start with the live route, write the camera move down so the result can be judged against it, read the expanded prompt when the output surprises you, measure the delivered file before you cut, and compare cost per accepted clip rather than cost per generation. The companion launch, Seedance 2.5, covers the other family in the same test and is the one to reach for when a clip needs to cut between shots on its own. Browse the rest of the lineup on Masonry's model hub.


