A consistent-character AI video is less about writing a magical identity prompt and more about reducing contradictions before the video model runs. In this worked example, one authorized portrait becomes a controlled establishing frame and close-up frame. Kling 2.6 Pro then generates a restrained camera push between them.
The page includes the actual five workflow images and final clip that were shipped with the original tutorial. That makes it first-hand process evidence, but only for one successful example. It does not establish a success rate, prove that two frames always beat one, or show that a particular duration is universally best.
Quick answer
- Start with a portrait you have permission to transform.
- Create an establishing frame and reject identity or scene errors.
- Derive a compatible close-up without changing identity-critical details.
- Compare the two stills before asking a video model to connect them.
- Select both frames in Masonry's Make Video flow and choose Kling 2.6 Pro.
- Ask for one small camera move, then review the entire clip frame by frame.
Before you generate: consent and provenance
Do not pull an attractive portrait from Pinterest or another public page and assume it is available for synthetic-media work. Use a photo you created, licensed, or received permission to transform. Confirm that the depicted adult understands the intended context, especially for advertising, advocacy, or a realistic first-person presentation.
Keep the original file, license or release, prompts, selected outputs, and final approval together. Avoid public-figure impersonation and misleading depictions. Disclose that the asset is synthetic when the audience could reasonably mistake it for a record of a real event.
Step 1: Choose an identity reference you can audit
The reference should make important identity features easy to inspect:
- face visible at useful resolution, without heavy smoothing or a beauty filter;
- natural, even lighting across the eyes, nose, mouth, jaw, and hairline;
- limited occlusion from hands, sunglasses, hair, or deep shadow;
- enough wardrobe and accessory detail to define what must remain stable.
A neutral portrait is easier to evaluate than a dramatic action frame, but it is not automatically better. The real requirement is an authorized source with observable details.
Step 2: Add the reference to Masonry
Add the portrait to the Masonry canvas and begin a reference-based image edit. This keeps the source attached to the new frame, but it does not guarantee that the generated face, skin, hair, or clothing will be exact.
Before continuing, write a short identity sheet from the source: face shape, eye spacing, eyebrow shape, hairline, hair color and length, distinguishing marks, wardrobe, and accessories. The sheet gives reviewers the same criteria later; otherwise “looks like the same person” becomes a moving target.
Step 3: Create and approve the establishing frame
This example uses Nano Banana Pro to place the subject in a beach-resort scene. Describe the composition and observable constraints instead of repeating vague quality adjectives. The model may still change the person or props, so the output remains a candidate until it passes review.
Establishing-frame prompt
Using the authorized portrait as the identity reference, create a natural vertical vacation photograph of the same adult seated on a sunbed at a beach resort. Preserve the observable facial structure, eye spacing, eyebrows, hairline, hair color and length. Keep the expression relaxed. The subject holds one opened young coconut with a white straw in both hands. Keep the jewelry and neutral swimwear consistent and anatomically plausible. Frame a medium-wide seated portrait. Use soft directional daylight, warm sand, an orange-and-white towel, and a rustic stone wall. Render believable skin, hands, fabric, jewelry reflections, and coconut texture. Do not add text, logos, extra jewelry, extra fingers, duplicated objects, or new people.
Review the face at full size, then check hands, jewelry, straw, coconut, clothing edges, and background geometry. Reject the frame rather than carrying a visible error into every downstream step.
Step 4: Derive a compatible closing frame
Use the approved establishing frame—not the original portrait alone—as the reference for the close-up. The second endpoint should change as little as possible beyond camera distance. Large changes in pose, expression, hand position, prop location, lighting, or background geometry give the video model incompatible states to reconcile.
Closing-frame prompt
Create a closer version of the provided approved beach frame. Preserve the same adult identity, facial proportions, hair, expression, wardrobe, jewelry, seated orientation, coconut, straw, lighting direction, and background layout. Move the camera forward to an upper-chest portrait. Keep both hands and the coconut plausible at the lower edge of the frame. Use a slightly shallower depth of field without moving or replacing background elements. Do not reshape the face, change the hairline, add or remove accessories, move the straw to the mouth, alter the product, introduce text, or invent new objects.
Place the start and end frames side by side. Compare eye and mouth shape, jawline, hairline, skin marks, wardrobe seams, jewelry count, hand placement, prop geometry, light direction, and horizon. If the endpoints disagree, regenerate the closing frame before making video.
Step 5: Configure the image-to-video run
Select both approved frames on the Masonry canvas, choose Make Video, and select Kling 2.6 Pro. Match the vertical composition used by the stills and start with the shortest duration that can express the intended move. In this shipped example, that was five seconds; it is an example setting, not a universal optimum.
The live Masonry CLI route was checked on August 4, 2026 with masonry models params kling-v2-6-pro-i2v. It currently requires one --image, supports 16:9, 9:16, and 1:1 output, and exposes a negative prompt and audio toggle. Because that route does not expose an end-frame flag, this tutorial does not pretend the UI's two-frame process can be reproduced exactly from the current CLI.
Step 6: Request one motion and audit the result
The motion brief should say what moves, what stays stable, and what constitutes a failure. Avoid asking simultaneously for a camera push, head turn, hand gesture, hair motion, expression change, prop interaction, and environmental action.
Video prompt
Create one continuous slow handheld camera push from the approved wider frame to the approved close-up. Keep the adult's facial structure, hair, wardrobe, jewelry, hands, coconut, and straw visually consistent. Preserve the light direction and beach layout. Allow only subtle breathing, one natural blink, and small handheld micro-movement. No speech, lip movement, cuts, sudden zoom, head turn, pose change, added objects, disappearing accessories, duplicated fingers, prop morphing, or background jump. Use realistic parallax as the camera moves forward.
The result below is the single shipped example behind this tutorial.
Frame-by-frame acceptance sheet
Watch once at normal speed for overall motion, once frame by frame for discontinuities, and once without looking at the subject so background jumps are easier to see.
| Check | Accept when | Reject when |
|---|---|---|
| Identity | Face shape, eyes, mouth, jaw, hairline, and distinguishing marks remain recognizably stable | Features drift, swap, smear, or briefly become another person |
| Anatomy | Hands, teeth, eyes, and body proportions remain plausible | Fingers merge, eyes deform, teeth flash, or limbs change length |
| Wardrobe and accessories | Counts, colors, shapes, and placement remain stable | Jewelry, straps, or fabric edges appear, vanish, or move unnaturally |
| Props | Coconut and straw keep their shape, location, and contact with the hands | The prop morphs, duplicates, floats, or intersects the body |
| Camera motion | One gradual push with plausible parallax | A digital jump, cut, reverse move, or background freeze appears |
| Lighting | Direction and exposure remain continuous | Highlights pulse, shadows flip, or skin tone changes abruptly |
Record the number of candidates generated, the number accepted without repair, generation cost, and review time. Without those denominators, a polished example cannot answer how reliable or economical the workflow is.
Troubleshooting identity drift
| Symptom | Likely conflict | Next test |
|---|---|---|
| Face morphs midway | Endpoint faces or expressions differ too much | Regenerate the closing frame with a smaller composition change |
| Hands or prop deform | Hand placement changes or the prop has complex contact | Keep the hands nearly still or crop below the contact point |
| Background jumps | Endpoint perspective or object layout does not match | Rebuild the closing frame from the approved start frame |
| Motion feels frozen | Prompt prohibits nearly all motion | Allow one blink, breathing, and subtle camera micro-movement |
| Clip becomes theatrical | Too many actions compete with the camera move | Remove all but one subject action and one camera action |
What this example does—and does not—show
The five images document a real Masonry workflow, and the embedded video shows that the selected endpoints produced one usable transition. That is original evidence readers can inspect.
It does not compare one frame with two, measure identity against a face-similarity threshold, report failed candidates, or benchmark Kling against another model. A stronger future test would hold the reference, prompt, duration, aspect ratio, and candidate count constant across one-frame and two-frame runs, then score blinded frames against the same acceptance sheet.
The bottom line
Two compatible frames can narrow the job from “invent a character video” to “connect these approved visual states.” They cannot lock identity or eliminate review. The practical workflow is to use an authorized reference, approve each still independently, request one restrained motion, and reject the full clip when any identity-critical detail drifts.
For adjacent model workflows, read the Kling 3.0 prompt guide and Seedance 2.0 prompt guide. Both make the provider-versus-Masonry boundary explicit and include their own evidence limitations.


