Masonry Logo
AI & Technology

Consistent Character AI Video: A 6-Step Two-Frame Workflow

Build a short character-consistent AI video from an approved portrait, two controlled keyframes, and a narrow motion brief—with prompts and an acceptance checklist.

Gaurav BisenGaurav Bisen
7 min read

A consistent-character AI video is less about writing a magical identity prompt and more about reducing contradictions before the video model runs. In this worked example, one authorized portrait becomes a controlled establishing frame and close-up frame. Kling 2.6 Pro then generates a restrained camera push between them.

The page includes the actual five workflow images and final clip that were shipped with the original tutorial. That makes it first-hand process evidence, but only for one successful example. It does not establish a success rate, prove that two frames always beat one, or show that a particular duration is universally best.

Quick answer

  1. Start with a portrait you have permission to transform.
  2. Create an establishing frame and reject identity or scene errors.
  3. Derive a compatible close-up without changing identity-critical details.
  4. Compare the two stills before asking a video model to connect them.
  5. Select both frames in Masonry's Make Video flow and choose Kling 2.6 Pro.
  6. Ask for one small camera move, then review the entire clip frame by frame.

Do not pull an attractive portrait from Pinterest or another public page and assume it is available for synthetic-media work. Use a photo you created, licensed, or received permission to transform. Confirm that the depicted adult understands the intended context, especially for advertising, advocacy, or a realistic first-person presentation.

Keep the original file, license or release, prompts, selected outputs, and final approval together. Avoid public-figure impersonation and misleading depictions. Disclose that the asset is synthetic when the audience could reasonably mistake it for a record of a real event.

Step 1: Choose an identity reference you can audit

The reference should make important identity features easy to inspect:

  • face visible at useful resolution, without heavy smoothing or a beauty filter;
  • natural, even lighting across the eyes, nose, mouth, jaw, and hairline;
  • limited occlusion from hands, sunglasses, hair, or deep shadow;
  • enough wardrobe and accessory detail to define what must remain stable.

A neutral portrait is easier to evaluate than a dramatic action frame, but it is not automatically better. The real requirement is an authorized source with observable details.

Step 1: Choose an authorized reference with identity-critical details that a reviewer can inspect

Step 2: Add the reference to Masonry

Add the portrait to the Masonry canvas and begin a reference-based image edit. This keeps the source attached to the new frame, but it does not guarantee that the generated face, skin, hair, or clothing will be exact.

Before continuing, write a short identity sheet from the source: face shape, eye spacing, eyebrow shape, hairline, hair color and length, distinguishing marks, wardrobe, and accessories. The sheet gives reviewers the same criteria later; otherwise “looks like the same person” becomes a moving target.

Step 2: Attach the approved portrait to a reference-based image edit

Step 3: Create and approve the establishing frame

This example uses Nano Banana Pro to place the subject in a beach-resort scene. Describe the composition and observable constraints instead of repeating vague quality adjectives. The model may still change the person or props, so the output remains a candidate until it passes review.

Establishing-frame prompt

Prompt

Using the authorized portrait as the identity reference, create a natural vertical vacation photograph of the same adult seated on a sunbed at a beach resort. Preserve the observable facial structure, eye spacing, eyebrows, hairline, hair color and length. Keep the expression relaxed. The subject holds one opened young coconut with a white straw in both hands. Keep the jewelry and neutral swimwear consistent and anatomically plausible. Frame a medium-wide seated portrait. Use soft directional daylight, warm sand, an orange-and-white towel, and a rustic stone wall. Render believable skin, hands, fabric, jewelry reflections, and coconut texture. Do not add text, logos, extra jewelry, extra fingers, duplicated objects, or new people.

Opens with the prompt already filled inTry this prompt

Review the face at full size, then check hands, jewelry, straw, coconut, clothing edges, and background geometry. Reject the frame rather than carrying a visible error into every downstream step.

Step 3: Approve identity, hands, wardrobe, prop count, and scene geometry before animating

Step 4: Derive a compatible closing frame

Use the approved establishing frame—not the original portrait alone—as the reference for the close-up. The second endpoint should change as little as possible beyond camera distance. Large changes in pose, expression, hand position, prop location, lighting, or background geometry give the video model incompatible states to reconcile.

Closing-frame prompt

Prompt

Create a closer version of the provided approved beach frame. Preserve the same adult identity, facial proportions, hair, expression, wardrobe, jewelry, seated orientation, coconut, straw, lighting direction, and background layout. Move the camera forward to an upper-chest portrait. Keep both hands and the coconut plausible at the lower edge of the frame. Use a slightly shallower depth of field without moving or replacing background elements. Do not reshape the face, change the hairline, add or remove accessories, move the straw to the mouth, alter the product, introduce text, or invent new objects.

Opens with the prompt already filled inTry this prompt

Place the start and end frames side by side. Compare eye and mouth shape, jawline, hairline, skin marks, wardrobe seams, jewelry count, hand placement, prop geometry, light direction, and horizon. If the endpoints disagree, regenerate the closing frame before making video.

Step 4: The closing frame changes camera distance while keeping identity and scene constraints compatible

Step 5: Configure the image-to-video run

Select both approved frames on the Masonry canvas, choose Make Video, and select Kling 2.6 Pro. Match the vertical composition used by the stills and start with the shortest duration that can express the intended move. In this shipped example, that was five seconds; it is an example setting, not a universal optimum.

Step 5: Send the two approved endpoints to Kling 2.6 Pro in Masonry's Make Video flow

The live Masonry CLI route was checked on August 4, 2026 with masonry models params kling-v2-6-pro-i2v. It currently requires one --image, supports 16:9, 9:16, and 1:1 output, and exposes a negative prompt and audio toggle. Because that route does not expose an end-frame flag, this tutorial does not pretend the UI's two-frame process can be reproduced exactly from the current CLI.

Step 6: Request one motion and audit the result

The motion brief should say what moves, what stays stable, and what constitutes a failure. Avoid asking simultaneously for a camera push, head turn, hand gesture, hair motion, expression change, prop interaction, and environmental action.

Video prompt

Prompt

Create one continuous slow handheld camera push from the approved wider frame to the approved close-up. Keep the adult's facial structure, hair, wardrobe, jewelry, hands, coconut, and straw visually consistent. Preserve the light direction and beach layout. Allow only subtle breathing, one natural blink, and small handheld micro-movement. No speech, lip movement, cuts, sudden zoom, head turn, pose change, added objects, disappearing accessories, duplicated fingers, prop morphing, or background jump. Use realistic parallax as the camera moves forward.

Opens with the prompt already filled inTry this prompt

The result below is the single shipped example behind this tutorial.

First-hand result from the documented two-frame workflow; one clip demonstrates feasibility, not a reliability rate

Frame-by-frame acceptance sheet

Watch once at normal speed for overall motion, once frame by frame for discontinuities, and once without looking at the subject so background jumps are easier to see.

CheckAccept whenReject when
IdentityFace shape, eyes, mouth, jaw, hairline, and distinguishing marks remain recognizably stableFeatures drift, swap, smear, or briefly become another person
AnatomyHands, teeth, eyes, and body proportions remain plausibleFingers merge, eyes deform, teeth flash, or limbs change length
Wardrobe and accessoriesCounts, colors, shapes, and placement remain stableJewelry, straps, or fabric edges appear, vanish, or move unnaturally
PropsCoconut and straw keep their shape, location, and contact with the handsThe prop morphs, duplicates, floats, or intersects the body
Camera motionOne gradual push with plausible parallaxA digital jump, cut, reverse move, or background freeze appears
LightingDirection and exposure remain continuousHighlights pulse, shadows flip, or skin tone changes abruptly

Record the number of candidates generated, the number accepted without repair, generation cost, and review time. Without those denominators, a polished example cannot answer how reliable or economical the workflow is.

Troubleshooting identity drift

SymptomLikely conflictNext test
Face morphs midwayEndpoint faces or expressions differ too muchRegenerate the closing frame with a smaller composition change
Hands or prop deformHand placement changes or the prop has complex contactKeep the hands nearly still or crop below the contact point
Background jumpsEndpoint perspective or object layout does not matchRebuild the closing frame from the approved start frame
Motion feels frozenPrompt prohibits nearly all motionAllow one blink, breathing, and subtle camera micro-movement
Clip becomes theatricalToo many actions compete with the camera moveRemove all but one subject action and one camera action

What this example does—and does not—show

The five images document a real Masonry workflow, and the embedded video shows that the selected endpoints produced one usable transition. That is original evidence readers can inspect.

It does not compare one frame with two, measure identity against a face-similarity threshold, report failed candidates, or benchmark Kling against another model. A stronger future test would hold the reference, prompt, duration, aspect ratio, and candidate count constant across one-frame and two-frame runs, then score blinded frames against the same acceptance sheet.

The bottom line

Two compatible frames can narrow the job from “invent a character video” to “connect these approved visual states.” They cannot lock identity or eliminate review. The practical workflow is to use an authorized reference, approve each still independently, request one restrained motion, and reject the full clip when any identity-critical detail drifts.

For adjacent model workflows, read the Kling 3.0 prompt guide and Seedance 2.0 prompt guide. Both make the provider-versus-Masonry boundary explicit and include their own evidence limitations.

Share:
FAQ

Questions from this guide

Concise answers to the questions readers ask after this guide

How do I keep a character consistent in an AI video?

Use an authorized, high-resolution identity reference; make the start and end frames agree on facial structure, hair, wardrobe, props, lighting, and environment; request one restrained camera move; and reject any clip that changes identity-critical details. Two frames constrain the transition but do not guarantee identity preservation.

Is one image or two images better for consistent-character video?

Two compatible endpoint frames can specify both the opening and intended closing composition, reducing some ambiguity. They can also create morphing when geometry, pose, props, or lighting disagree. Run a controlled test because model and product support for end frames varies.

Which model does this tutorial use?

The shipped example uses Kling 2.6 Pro inside Masonry's Make Video flow. The current Masonry CLI route named kling-v2-6-pro-i2v accepts one required source image, so this article does not claim that the UI's two-frame workflow has command-line parity.

Does a reference image guarantee the same face?

No. A reference image conditions the generation; it is not an identity lock. Inspect the full clip for changes to face shape, eyes, teeth, hairline, hands, clothing, and props, and compare key frames with the authorized source before publishing.

Can I use any portrait I find online?

No. Use a portrait you created or have permission to use, confirm the person's consent for synthetic media, and avoid impersonation or misleading contexts. Keep provenance and approval records and disclose synthetic content where the publishing context calls for it.