Masonry Logo
AI & Technology

AI Video Creative Agent for Ecommerce: Approval Workflow

A merchant workflow lets an AI creative agent propose video briefs, but requires product authority and human approval before generation, release, and spend.

Gaurav BisenGaurav Bisen
9 min read

An AI video creative agent should make production easier. It should not quietly become the authority for what a product is, what a merchant may claim, which offer is live, where a click lands, or whether an asset deserves spend.

The practical design is a split system: let the agent propose and execute, but require merchant-owned evidence at the moments where a wrong decision can misrepresent a SKU or waste budget. We tested that system with one fictional candle source. The agent proposed three distinct briefs. One had enough authority to generate. Two were held before credits were spent.

Evidence boundary: the candle and merchant are fictional. The source file, planning gate, Masonry job, returned clip, and file inspection are first-hand. No live ad, sale, revenue result, or platform account test occurred. Community discussions establish current merchant concerns, not failure rates or performance benchmarks.

This page entered the queue from non-brand evidence rather than brand-navigation traffic. Two exact visible Search Console query variants for an AI video creative agent total 276 impressions, zero clicks, and average positions 7.30 and 10.40. In PostHog's separate mature 90-day read, the broad video comparison has 408 qualified readers, seven assisted signups, and two assisted checkout completions; the product-ad comparison has 87 readers and two signups; the existing agent review has 41 readers and one signup. None recorded a qualified activation in the seven-day window. These are all-source, person-level assisted outcomes—not query-level or causal revenue attribution.

Quick answer: where the agent stops

StageAgent may doMerchant authority required
Intakenormalize the goal and inventory supplied filesexact SKU, variant, source owner, claims, offer, destination
Planningpropose distinct hypotheses and list missing evidencedecide which hypothesis is worth testing and what is forbidden
Generationtranslate one approved row into a bounded prompt and run one jobcandidate cap, spend cap, route permission, likeness or product constraints
Reviewextract frames, compare visible features, flag uncertaintyaccept, reject, request evidence, or authorize a correction
Releasepackage the asset, lineage, copy, destination, and rollback recordfinal placement, budget, disclosure, policy, accessibility, and launch approval
Learningjoin the declared creative change to delivery and commerce recordsdecide with contribution and incrementality, not generation volume

This is intentionally less autonomous than “paste a product URL and receive twenty ads.” It is more useful because each expansion of scope has a named owner and a reviewable record.

Download the three-row creative-agent approval manifest. Edit it before generation; do not reconstruct authority after a polished output arrives. The two blocked rows use HOLD_MISSING_AUTHORITY; the generated row remains pending until its observed frame and file review is written back to the manifest. In this recorded run, H01 advanced to SECONDARY_REVIEW_NOT_RELEASED; it did not become an approved ad.

Why ecommerce creative agents need an authority gate

Current merchant discussions describe automated creative features replacing an intended video with unrelated storefront stills and changing product brightness or color. Another advertiser asked whether a system should turn one product photo into dozens of images and ten to twenty videos. These are individual reports and questions, not universal platform behavior. They expose the same operating risk: output volume can outrun product, offer, and release review.

Amazon Ads' current partner guidance makes a compatible distinction. It recommends defining where AI fits, keeping the human team in the director's seat, and holding AI creative to the same performance standards as other assets. The page includes one partner case, so its reported campaign result should not be generalized into an ecommerce benchmark. The durable lesson is the control structure, not the number: use the tool deliberately and preserve human direction.

Step 1: convert the request into an authority record

The source for this run was one 1254 × 1254 image of a fictional amber candle jar. Visible authority included one closed amber glass vessel, one matte black lid, one cream rectangular label, and the words NORTHLINE, CEDAR, and 8 OZ.

Approved source authority, 1254 × 1254. It shows a closed jar and label; it does not authorize a flame, scent experience, discount, price, review, or customer outcome.

The source did not authorize an open jar, burning wick, flame behavior, smoke, scent diffusion, room context, safe-use claim, price, discount, inventory, review, or destination. Those facts remain outside the model even if a creative brief would benefit from them.

A product URL can help retrieve this record, but retrieval is not approval. Freeze the exact variant, source bytes, visible features, claims, offer, destination, and owner before asking an agent to plan. For the full URL-to-source contract, use the Shopify product-page to video workflow.

Step 2: make the agent propose hypotheses before assets

We required three one-sentence hypotheses before any generation call:

IDProposed jobAuthority statusDecision
H01slow source-frame push-in with a warm background light sweepvisible product source is sufficient; no new product action or copyapprove one generation
H02open and light the candle in a quiet evening ritualno approved open, wick, flame, safety, room, or human-use sourcehold
H03end on a 20% off offer cardno approved discount, dates, terms, inventory, or destinationhold

This gate improved the workflow before the model did anything. H02 might create the most emotionally interesting footage, but the source cannot prove that state. H03 might be easy to render, but the agent cannot invent commercial authority. Only H01 changes motion and lighting around the visible source.

The goal is not to reject ambitious ideas forever. It is to turn missing authority into an explicit request: obtain an approved open-state source for H02 or an approved offer record for H03, then re-evaluate that row.

Step 3: translate only the approved row into a generation brief

The approved brief preserved the exact product count, silhouette, amber glass, black lid, cream label, readable wording, orientation, and closed state. It allowed a slow push-in and soft background light sweep. It prohibited opening, ignition, rotation, people, props, new text, claims, cuts, and audio.

Prompt

masonry video "Create one restrained five-second ecommerce product-ad review candidate from this approved source frame. Keep the single amber glass candle jar stationary, centered, fully visible, and identical in silhouette, proportions, amber glass color, black lid, cream rectangular label, label wording NORTHLINE CEDAR 8 OZ, product count, orientation, and visible condition. Use only a slow camera push-in and a soft warm light sweep across the neutral background. Do not open the jar, ignite a flame, rotate or move the product, add hands or people, add props, change or animate the label, alter any letter or number, invent claims, add text, logos, badges, price, CTA, cuts, transitions, or audio. This is a review candidate, not an approved ad." \ --model kling-o3-pro \ --image ./approved-source.webp \ --aspect 1:1 \ --no-audio

Masonry returned asynchronous job 3830fb6a-be0c-433f-8ce4-0318fe476d8f. Job creation estimated 67.2 credits, and the completed job recorded 67.2 credits used. The delivered H.264 file is 5.041667 seconds, 1440 × 1440, 24 fps, and 121 frames with no audio stream.

Step 4: inspect the returned timeline, not the success flag

Observed first return from approved hypothesis H01. The job completed; release still depends on frame and file review.
Timeline review at five intervals. Check label characters, jar and lid geometry, product count, new objects, and the exact motion before assigning a disposition.

The review must answer observable questions:

  1. Does every sampled frame contain one product with the same silhouette and closed state?
  2. Do amber glass, black lid, cream label, and label boundary stay consistent?
  3. Are NORTHLINE, CEDAR, and 8 OZ still readable and unchanged?
  4. Does the product remain stationary while only the camera and background light move?
  5. Did any hand, flame, prop, badge, claim, new text, transition, or audio appear?
  6. Does the file match the intended duration, frame rate, dimensions, and audio contract?

Any material mismatch rejects the candidate for exact-SKU advertising. A visually appealing clip is not a correction record. Preserve the failed return and update the manifest; do not quietly replace it with the next seed.

Observed disposition: SECONDARY_REVIEW_NOT_RELEASED. Across frames 0, 30, 60, 90, and 120, the clip keeps one closed amber jar, one black lid, one cream label, and the readable strings NORTHLINE, CEDAR, and 8 OZ. No person, flame, new prop, new copy, claim, cut, or audio appears. The allowed push-in and warm background light sweep are visible. However, the route reconstructs pixels through time and we did not compare the result with a physical SKU, packaging specification, or alternate approved view. It is useful evidence for the approval system, not sufficient authority for commercial release.

Step 5: package the evidence for a real release decision

A review packet should travel with the asset:

  • immutable source path or hash and authority owner;
  • selected hypothesis and the held alternatives;
  • exact prompt, model route, job ID, estimated and actual cost;
  • returned file hash, dimensions, duration, frame rate, and audio state;
  • frame review, reviewer, disposition, and unresolved uncertainty;
  • deterministic headline, offer, disclosure, captions, and CTA added later;
  • exact destination and expected SKU, variant, price, offer, and cart behavior;
  • campaign, placement, audience role, budget, test window, metric, stopping rule, and rollback owner.

The agent can assemble that packet. It should not sign its own release. If the job needs several visual hypotheses, use the one-SKU creative testing matrix to keep concept, execution, copy, offer, and destination from changing at once.

Step 6: measure the commercial system, not agent activity

Generation count, accepted-asset rate, and cost per accepted asset are production metrics. They help budget the workflow, but they do not show that the creative generated profit.

For a paid test, preserve the comparison contract and read:

  • spend, impressions, reach, and frequency for delivery context;
  • video starts, early hold, completion, clicks, and landing sessions as diagnostics;
  • add-to-cart, checkout, purchases, purchase value, cancellations, and returns;
  • contribution after media, production, discount, product, fulfillment, service, and return costs;
  • incrementality evidence when the spend and decision justify it.

Do not call the agent successful because it produced twenty videos. Do not call the clip a winner because it has a high click-through rate. The system wins when one truth-preserving creative change improves the merchant's declared business outcome after costs.

The operating rule

An ecommerce creative agent should expand execution capacity, not its own authority. Give it the approved sources, bounded decisions, candidate and spend caps, inspection checklist, and measurement contract. Let it plan, generate, inspect, and package evidence. Keep product truth, claims, offer, release, budget, and commercial judgment with the merchant.

For model selection, use the controlled AI video generator comparison. For one source-to-ad pipeline, use the Shopify product-page workflow. This page owns the missing layer between them: deciding what the agent is allowed to do before the first generation and what a human must verify before the first dollar of spend.

Share:
FAQ

Questions from this guide

Concise answers to the questions readers ask after this guide

What is an AI video creative agent for ecommerce?

It is a workflow that can turn a merchant goal and approved sources into proposed briefs, generation jobs, review evidence, and delivery records. The useful version does not invent product facts, offers, destinations, or release authority; it asks for or holds those decisions.

Which decisions can a creative agent make automatically?

An agent can safely propose hypotheses, organize approved inputs, translate one selected hypothesis into a bounded visual brief, call a generation route, inspect files, and assemble a review packet. A merchant still supplies product authority, approved claims and offers, risk limits, release permission, budget, and the business metric.

Should an AI creative agent generate dozens of video variants?

Not before the brief and source pass an approval gate. Start with a small hypothesis set, approve one distinct variable, generate one candidate, and inspect it. More outputs multiply review work and can hide repeated product or claim errors.

How do you review an AI product video?

Compare several frames with the approved source. Check product count, silhouette, geometry, materials, color, labels, readable characters, parts, motion, added objects, implied use, claims, audio, crop, and destination. A completed generation job is only a returned candidate.

How should merchants measure AI video creative?

Use delivery and attention metrics as diagnostics, then decide with purchases and contribution after media, production, discounts, fulfillment, returns, and other variable costs. Preserve audience, offer, destination, placement, budget, attribution, and test window so the creative is the declared change.