Masonry Logo
AI & Technology

MiniMax H3 Max and H3 Max Turbo: 3 Clips and the Rates

MiniMax H3 Max and H3 Max Turbo are new video routes on Masonry. Inspect three first-hand clips, the measured 1344x768 output behind the 768P label, the standard per-second provider rates, and the reference-image behaviour that separates this family from Seedance.

Gaurav BisenGaurav Bisen
10 min read

MiniMax H3 Max and MiniMax H3 Max Turbo are new video routes on Masonry, under the keys minimax-h3-max and minimax-h3-max-turbo.

The useful part of a launch post is not the adjective list. It is three clips that were generated, downloaded, measured, and watched, plus the handful of behaviours that will actually change how you queue work: what 768P returns in pixels, what the response hands back for reproducibility, and which route will take a portrait reference at all.

Evidence boundary: this article rests on six clips generated on September 4, 2026 across two model families, three of them on the MiniMax H3 routes shown below. Each is a single run of a single prompt. There were no repeated trials, no held-out prompt set, no standardized queue timing, no reliability rate, and no leaderboard position. File metadata came from ffprobe on the downloaded files; the cut counts quoted later came from ffmpeg scene detection at threshold 0.4. Every reference and source still used here was generated on Masonry for this article, so no real individual's likeness was submitted to any route. Provider rates were read September 4, 2026 and change without notice. Nothing here establishes how either route behaves on your prompts, your references, or your acceptance bar.

Quick answer

  • New routes: minimax-h3-max lists text-to-video, image-to-video, and reference-to-video with up to 4 reference images. minimax-h3-max-turbo lists text-to-video and image-to-video.
  • Measured output: all three clips requested at 768P and 16:9 returned 1344×768, 24 fps, H.264 with an AAC stereo track at 32 kHz.
  • Measured duration: 5.184 seconds on every clip, against a requested duration of 5.
  • Standard provider rates, checked September 4, 2026: H3 Max at $0.05/s (480P) and $0.08/s (768P); Turbo at $0.025/s and $0.04/s. Turbo is exactly half the H3 Max rate at every listed resolution.
  • Reference behaviour: H3 Max reference-to-video accepted a portrait reference and carried that person into a new scene. Seedance 2.5 refused the identical file, reporting a possible real-person likeness.
  • Reproducibility field: these H3 Max text-to-video responses returned an expanded prompt and no seed.
  • Not established here: output quality ranking between the two routes, image-to-video behaviour on either route, latency, reliability, 480P output, durations other than 5, or any Masonry credit price.

The two routes and what they cost

ItemChecked September 4, 2026Production implication
H3 Max keyminimax-h3-maxPin the exact key in automation; the Turbo key is a different route, not a flag.
H3 Max modesText-to-video, image-to-video, reference-to-videoReference mode is the reason to reach for H3 Max over Turbo.
Turbo keyminimax-h3-max-turboNo reference mode is listed; do not plan a reference shot on Turbo.
H3 Max rate$0.05/s at 480P, $0.08/s at 768PA five-second 768P clip is about $0.40 of provider cost.
Turbo rate$0.025/s at 480P, $0.04/s at 768PExactly half the H3 Max rate at both resolutions; the same five-second clip is about $0.20.
768P in pixels1344×768 on all three clipsDo not size a placement from the label; measure the delivered frame.
Delivered duration5.184 s against a request of 5Do not build a hard-cut edit that assumes exactly 5.000 s.
AudioAAC stereo, 32 kHz, on every clipAudio arrives whether or not your edit wants it; listen before publishing.

Provider rates are what fal charges for the generation. What a Masonry generation is metered at is set by the platform pricing table, not by this article.

Three shipped clips

Text-to-video on H3 Max

MiniMax H3 Max text-to-video. Requested 768P, 16:9, duration 5. Measured file: 1344×768, 5.184 s, 24 fps, H.264 + AAC stereo 32 kHz, 5.64 MB, mean volume −22.2 dB.

The brief asked for a modern coffee bar in morning light, a barista sliding a matte-black cold-brew bottle across a steel counter, and a slow dolly-in. The returned clip performs a genuine slow dolly-in, and every named prop is present. Adherence to a written camera move is the specific thing worth checking here, because a model that ignores camera direction turns every shot into a lottery.

Watch it through rather than judging the thumbnail: hands on the bottle, the bottle's contact with the counter, the counter edge as the camera moves, and background geometry under parallax.

The same prompt on H3 Max Turbo

MiniMax H3 Max Turbo text-to-video, identical prompt to the clip above. Requested 768P, 16:9, duration 5. Measured file: 1344×768, 5.184 s, 24 fps, H.264 + AAC stereo 32 kHz, 5.69 MB, mean volume −18.1 dB.

Same prompt, same scene structure, different composition and a noticeably darker grade. Watching both, Turbo reads as comparable rather than obviously worse.

That sentence is the whole claim, and it is deliberately weak. One pair of clips is not a benchmark. Two single runs on one prompt cannot establish which route produces better output, how often either fails, or whether the gap widens on harder briefs. If the decision matters to your budget, run your own prompt set on both keys, keep the prompt, resolution, duration, and acceptance sheet fixed, and count accepted clips rather than impressions.

Recommended: treat the halved rate as the reason to test Turbo on your own material, not as evidence that Turbo is worse.

Reference-to-video on H3 Max, two references

The reference run sent two images: a portrait of a woman and a product still of a specific bottle. Both were generated on Masonry for this article rather than photographed, so no real individual's likeness was submitted to either route. That matters for how you read the refusal further down.

Reference image 1 of 2 submitted to the H3 Max reference-to-video route.
Reference image 2 of 2 submitted to the same request.
MiniMax H3 Max reference-to-video from both images above. Requested 768P, 16:9, duration 5. Measured file: 1344×768, 5.184 s, 24 fps, H.264 + AAC stereo 32 kHz, 6.67 MB, mean volume −33.0 dB.

The result carried both references into a new scene. The person arrives with face, bob haircut, and charcoal turtleneck intact, the bottle arrives as the specific bottle rather than a generic one, and the requested lift-and-turn action is executed.

That is one run on one pair of references. It does not measure identity drift across many generations, it does not test how the model handles a reference whose lighting fights the target scene, and it does not tell you whether label text on a real product would survive. For commercial work where the product on screen has to be the product you ship, keep the product-ad model test discipline: approve the source, reject any frame that exceeds it, and log every rejection.

Behaviour worth knowing before you queue a batch

One family refused the portrait reference, the other did not

The same generated portrait was submitted to both families on September 4, 2026:

RouteResult on the identical image
minimax/h3-max/reference-to-videoAccepted; returned the clip shown above.
bytedance/seedance-2.5/reference-to-videoHTTP 422, content_policy_violation / partner_validation_failed — the images or videos provided may contain likenesses of real people or other private information that cannot be processed.

Re-running the Seedance request with object-only references (a bottle and a sneaker) succeeded. So this is a routing decision, not a defect: build a person-carrying reference shot on H3 Max, and product-only reference shots on either family. Plan the fallback before a deadline, not during one.

The refused file was a generated portrait, so what this records is what each filter does when handed a face, not a ruling on any specific real person. It is also one observation per route on one image. Provider policy changes, and holding rights and consent for a likeness remains your responsibility regardless of which route accepts the upload.

H3 Max rewrites your prompt and shows you the rewrite

The H3 Max response carries an expanded_prompt field. The one returned for the barista shot opens like this:

Prompt

integrated_multimodal_description: [Shot 1] Cinematic, live-action, medium shot. The scene opens in a modern specialty coffee shop during the early morning…

Opens with the prompt already filled inTry this prompt

It goes on to name the shot, the lighting mix, the wardrobe, and the camera move — including "slow dolly-in (push in) with small amplitude at slow speed". That is genuinely useful: when a result surprises you, the expanded prompt shows what the model decided you asked for, which is usually where the disagreement lives.

The trade is that these text-to-video responses returned no seed. Seedance 2.5 does the reverse, returning a seed and no expanded prompt. Practical consequence: on Seedance, reproducibility runs through the seed; on H3 Max, you debug by reading the expansion.

A resolution label is not a resolution

768P on H3 Max renders 1344×768. 720p on Seedance 2.5 renders 1280×720. Two similar labels, two different frames, two different crops for a placement that expects one of them. Read the delivered file.

Requested duration is a target

Every H3 Max clip here measured 5.184 seconds against a requested duration of 5, all three identically. Seedance landed at 5.04 to 5.06 seconds on the same request. If your edit assumes exactly 5.000 seconds, the overshoot lands on your timeline, not on the provider's.

Transient provider errors happen

One h3-max-turbo/text-to-video submission returned HTTP 504 downstream_service_unavailable and succeeded on a plain retry. One 504 in a handful of submissions is neither a reliability problem nor a reliability guarantee; it is a reason to make retry the default in any batch script rather than treating the first error as a verdict.

Choosing between the two keys

Choose this when the shot needs a person or a specific product carried in from a reference image: H3 Max is the only key in this family that lists reference-to-video.

Best for a cost-sensitive draft pass: Turbo, at exactly half the H3 Max rate, for iterating on framing and action before committing the final render to whichever route your own test prefers.

The honest version of the comparison is that the rates are known and the quality gap is not. The rate difference is documented by the provider; the quality difference would need a prompt set, repeated runs, and a fixed acceptance sheet that this article does not have.

Not established here

  • Which of the two routes produces better output. One paired run cannot answer that.
  • Image-to-video behaviour on either key. Both list the capability; neither was exercised for this article.
  • 480P output, durations other than 5, or aspect ratios other than 16:9.
  • Latency, queue time, throughput, retry rate, or failure rate under load.
  • Identity or product fidelity across repeated reference runs.
  • Any Masonry credit price, margin, or per-generation cost to you. Provider rates are quoted; platform metering is set elsewhere.

Bottom line

Two new keys, one of them at half the rate of the other, and one of them able to carry a person from a portrait reference into a new scene. The three clips here show the routes working; they do not rank them.

Start with the live route, write the camera move down so the result can be judged against it, read the expanded prompt when the output surprises you, measure the delivered file before you cut, and compare cost per accepted clip rather than cost per generation. The companion launch, Seedance 2.5, covers the other family in the same test and is the one to reach for when a clip needs to cut between shots on its own. Browse the rest of the lineup on Masonry's model hub.

Share:
FAQ

Questions from this guide

Concise answers to the questions readers ask after this guide

What separates MiniMax H3 Max from H3 Max Turbo?

They are two routes in the same family. H3 Max lists text-to-video, image-to-video, and reference-to-video; Turbo lists text-to-video and image-to-video, with no reference mode. On provider rates checked September 4, 2026, Turbo costs exactly half of H3 Max at both listed resolutions. One paired text-to-video run in this article is not enough to rank their output quality, and this article does not.

What pixel dimensions does the 768P setting actually return?

All three clips in this article were requested at 768P with a 16:9 aspect ratio and came back as 1344x768, not 1280x720. Seedance 2.5 requested at 720p returns 1280x720. Similar-looking resolution labels mean different things across model families, so read the delivered file rather than the label when a placement needs exact dimensions.

Can MiniMax H3 Max accept a portrait reference?

In the run recorded here on September 4, 2026, the reference-to-video route accepted a portrait reference and returned a clip that carried the face, bob haircut, and charcoal turtleneck into a new scene. The reference was itself generated on Masonry for this article, not a photograph of a real individual. Seedance 2.5 reference-to-video refused the identical file with an HTTP 422 content policy error reporting a possible likeness of a real person, so the two filters read the same image differently. That is one observation per route on one image, not a general permission, and provider policy can change. Confirm you hold rights and consent for any likeness you submit.

Does MiniMax H3 Max return a seed for reproducible reruns?

The text-to-video responses recorded here returned no seed. They returned an expanded prompt instead: a long structured shot description that names framing, lighting, wardrobe, and camera move. Seedance 2.5 is the opposite in this comparison, returning a seed and no expanded prompt. Plan reproducibility around whichever field the chosen route actually returns.

What does a five-second MiniMax H3 Max clip cost at the provider?

At the standard rates checked September 4, 2026, H3 Max is 0.08 US dollars per second at 768P, so a five-second clip is about 0.40 US dollars of provider cost, and the same clip on Turbo is about 0.20 US dollars. Those are provider rates, not what a Masonry generation is metered at. Rates change, so recheck before planning a batch.

Is a requested five-second duration exactly five seconds?

No. Every H3 Max clip measured for this article landed on 5.184 seconds against a requested duration of 5. Seedance 2.5 landed on 5.04 to 5.06 seconds against the same request. Treat requested duration as a target, and measure the delivered file before cutting to a fixed-length slot.