Masonry Logo
AI & Technology

AI Lifestyle Product Photography: A 3-Model Same-SKU Test

We gave three current AI image routes one exact product and one person-with-product brief. See the unselected outputs, grip and SKU-fidelity scorecard, fixed prompt, and source-preserving release workflow.

Gaurav BisenGaurav Bisen
12 min read

An AI lifestyle product photo has to solve two jobs at once: make the scene feel real and keep the product true. A quiet background replacement can redraw an edge; putting the product in a person's hand adds scale, grip, occlusion, contact shadow, perspective, and anatomy to the same decision.

We gave three current image routes one fictional bottle, one fixed person-with-product prompt, and one square output. GPT Image 2 came closest to this single brief. Nano Banana 2 also kept the visible product identity while using a tighter grip. Qwen Image Edit Plus returned an attractive but stylized scene, moved the grip to the cap, and did not preserve the required interaction contract.

That is a practical screen, not a model leaderboard. One result per route cannot establish a reliability rate, and none of the candidates becomes factual primary photography merely by looking convincing.

Evidence boundary: NORTHLINE is a fictional product reused as a controlled source. The three images below are the first unselected returns from August 16, 2026: one Nano Banana 2, one GPT Image 2, and one Qwen Image Edit Plus candidate. A Seedream 5 Pro submission returned an insufficient-credits error before a job was created, so there is no Seedream image or result to review. No real SKU, shopper, listing, ad, conversion, return, or revenue outcome was tested.

Why this is a merchant problem, not a prompt trick

The supplied non-brand Search Console export contains the exact question, “how can AI generate compliant product images and lifestyle mockups for ecommerce listings?” It recorded 26 impressions at an average position of 5.46. That is a small striking-distance signal, not proof of a large market. It does show that searchers are joining two concerns: creating a scene and releasing it safely.

Merchant discussions describe the same operational failure. One Shopify merchant reports that person-with-product attempts changed colors, scale, front-versus-back details, or replaced the item with a similar object. Another ecommerce merchant wants lifestyle backgrounds while explicitly protecting labels, colors, and small details, then concludes that testing multiple models against the same source is more useful than choosing from claims. A third discussion repeatedly separates useful lifestyle variation from text, color, reflection, and hand failures. These are qualitative signals, not verified model win rates or market research. Read the person-with-product question, the same-source lifestyle-tool question, and the broader merchant experience thread.

The useful response is therefore not “add more prompt detail.” It is a test with a fixed source, declared invariants, unselected outputs, visible denominators, and a disposition that can be audited.

The exact product and release gate

Controlled source, 627 × 1254. It authorizes one tall sage bottle, black cylindrical cap, one camera-right cap loop, bottom seam and base, vertical NORTHLINE, and 750 ML. It does not prove real-world scale, use, material, or performance.

We reviewed the product before judging the mountain, clothing, or photographic finish.

GateRequired evidence
IdentityExactly one sage NORTHLINE bottle; no substitution or extra product
GeometrySource-consistent body, rounded shoulders, black cap, camera-right loop, bottom seam, and base
AppearanceSource-consistent sage hue and matte finish; no invented liquid, transparency, damage, or condensation
MarkingsExact vertical NORTHLINE and exact 750 ML, both completely visible
InteractionOne plausible hand contacting the side edges; no fused fingers, impossible grip, product penetration, or hidden required evidence
SceneAdult trail context and soft morning light without an invented performance or use claim
ReleaseSECONDARY_REVIEW, COMPOSITE_PRODUCT, or REJECT; never automatic factual-primary approval

Download the complete source, job, gate, and disposition manifest. It includes the source hash, prompt version, job IDs, route names, seed state, raw-file hashes, dimensions, review results, and the failed Seedream attempt.

The fixed prompt and run contract

Every submitted route received the same source, square output, and prompt:

Use the supplied fictional product source as immutable product-identity evidence. Create one square photorealistic ecommerce secondary lifestyle image of an adult trail walker at a quiet overlook in soft early-morning light, framed from shoulders to hips. The person holds the exact bottle upright at mid-torso with one natural hand: palm behind the bottle and fingers contacting only the outer side edges so the entire front, vertical NORTHLINE text, and 750 ML remain visible. Preserve exactly one tall sage cylindrical bottle, its silhouette and proportions, rounded shoulders, black cylindrical cap, single black circular loop on camera-right, bottom seam and base, sage color and matte finish, exact vertical NORTHLINE text, exact 750 ML text, and quantity one. Match believable hand contact, scale, perspective, lighting, and shadows without changing the product. Change only the person, environment, framing, and scene lighting. Add no second bottle, straw, handle, lid part, sticker, label panel, logo, claim, badge, price, CTA, watermark, extra text, condensation, liquid, damage, or invented product feature. This is a review candidate, not factual primary photography.

The live routes were gemini-3.1-flash-image-preview for Nano Banana 2, gpt-image-2, qwen-image-edit-plus, and seedream-5-pro. The first, third, and fourth routes exposed a seed; GPT Image 2 did not expose one in this run. We did not replace the failed Seedream job with a cheaper route because that would silently change the test.

Prompt

masonry image "<fixed prompt above>" \ --model gemini-3.1-flash-image-preview \ --ref ./approved-sage-source.webp \ --aspect 1:1 \ --seed 62101 masonry job wait <job-id> masonry job download <job-id> --output ./candidate.png

Result 1: GPT Image 2 followed the whole brief most closely

GPT Image 2, first unselected return, 1024 × 1024. The product front, NORTHLINE, 750 ML, cap, and camera-right loop remain visible; the side-edge grip and shoulders-to-hips overlook composition match the prompt most closely.

The candidate kept one recognizable bottle, exact visible words, sage finish, base seam, and cap loop. The hand contacts the right edge while leaving the front readable. The scene reads as a secondary lifestyle photo instead of a product floating in front of a generated backdrop.

The boundary still matters. The bottle is a re-render, not a pixel-preserved composite. Its width, shoulder curve, cap relationship, surface, and hand contact were synthesized. Without a physical product or an approved multi-view source, the image cannot prove exact geometry, material, scale, or grip clearance.

Disposition: SECONDARY_REVIEW. Closest single result to the fixed brief; not factual-primary authorization and not evidence that GPT Image 2 will repeat the result.

Result 2: Nano Banana 2 preserved the product with a tighter grip

Nano Banana 2, first unselected return, 1024 × 1024. The visible product, text, black cap, loop, bottom seam, and quantity survive. The hand wraps farther around the bottle than the declared side-edge grip.

This candidate also clears the visible identity, appearance, marking, interaction-plausibility, and scene gates. It is a strong lifestyle direction: one product, coherent morning light, readable front, and a hand that belongs to the person rather than an isolated stock-photo insert.

The tighter wraparound grip changes the review problem. Fingers cross more of the left product edge, so a merchant has less visible evidence for the source silhouette. The output still shows the required words and loop, but the generated hand-product boundary makes exact source comparison harder than the GPT candidate.

Disposition: SECONDARY_REVIEW. Useful after product review; composite or photograph the real interaction when edge geometry and scale must be authoritative.

Result 3: Qwen preserved the name but missed the image contract

Qwen Image Edit Plus, first unselected return, 1024 × 1024. The model kept recognizable NORTHLINE markings but returned a dark stylized illustration, reduced the product, moved the grip to the cap, and lost required cap-loop evidence.

The output is not unusable art. It is unusable evidence for this declared job. The image became a graphic illustration rather than a photorealistic secondary photo; the bottle moved down beside the body; the hand grips the cap instead of the side edges; and the source loop is not available for review. A polished mountain composition cannot compensate for those contract failures.

This distinction protects the benchmark from aesthetic bias. If the task were a stylized campaign poster, the same file might be a direction. Under this source-preserving person-with-product brief, it is a rejection.

Disposition: REJECT. Do not rescue the file by cropping away the failed grip or product scale.

What the three returns support

RouteCandidatesVisible identity and markingsRequested hand interactionPhotorealistic sceneRelease
GPT Image 21Pass, exact-SKU geometry still unprovenClosest matchPassSECONDARY_REVIEW
Nano Banana 21Pass, exact-SKU geometry still unprovenPlausible but tighter than requestedPassSECONDARY_REVIEW
Qwen Image Edit Plus1Name survives; geometry evidence failsFail, cap gripFail, stylizedREJECT
Seedream 5 Pro0NOT_RUNNOT_RUNNOT_RUNHOLD

The job durations were 125.099 seconds for GPT Image 2, 5.107 seconds for Nano Banana 2, and 10.113 seconds for Qwen Image Edit Plus. These are single queue observations, not latency benchmarks. The Seedream submission failed before job creation because the workspace lacked enough credits for that route. Its zero is an operational constraint, not a model-quality result.

Why hands make product fidelity harder

A background-only edit can keep the product isolated. A human interaction creates new hidden and visible relationships:

  • Occlusion: fingers cover the exact edges a reviewer needs to compare.
  • Contact geometry: the hand must wrap around the actual product depth, not a plausible generic cylinder.
  • Scale: a convincing person can make a bottle look larger or smaller than the source permits.
  • Perspective: a front-on reference does not authorize unseen side, back, cap-top, or underside details.
  • Light transfer: skin, fabric, matte product surfaces, and contact shadows must agree without repainting the SKU.
  • Use implication: a trail, kitchen, bathroom, child, pet, or athlete can imply suitability or performance the source never established.

Prompting can narrow these failures; it cannot manufacture missing authority. Multiple approved views, physical dimensions, color references, packaging artwork, and an interaction reference make the review stronger. They still do not make a generated redraw identical by default.

A source-preserving production workflow

1. Choose the commercial role first

Separate factual primary photography, PDP secondary media, paid creative, social content, and a photographer's scene brief. Each role tolerates different changes and follows different placement rules. Use the AI product-photo platform rules comparison before treating “lifestyle image” as one universal release class.

2. Build a real source packet

Store the approved front, side, back, top, and interaction views that exist; exact dimensions; color-managed references; label artwork; variant and quantity; included parts; condition; and approved claims. Record what remains unknown. A single front view should not silently authorize the model to invent the rear or infer hand scale.

3. Declare the interaction

Specify which hand, where it contacts the product, which evidence must stay visible, and what constitutes an impossible or misleading grip. A drink bottle held at the side, a pump pressed by a finger, a handbag worn on a shoulder, and a ring on a hand need different geometry and safety checks.

4. Generate equal candidate counts

Keep the source packet, prompt, output size, aspect ratio, candidate count, and review rubric fixed. Do not compare one route's chosen hero with another route's first return. This three-image screen is useful for finding failure patterns; a procurement or production decision needs repeat attempts and cost per accepted asset.

5. Review the product before the person

Compare identity, geometry, appearance, markings, variant, quantity, and included parts at full resolution. Then inspect anatomy, grip, occlusion, scale, gaze, scene, and implied use. Reject a beautiful interaction if the product changes.

6. Composite when pixels matter

If generation repeatedly redraws the SKU, use a real person-with-product capture or a controlled composite. One workable fallback is to generate the person and environment around a neutral proxy, then replace the proxy with the approved product, rebuild the hand occlusion with an explicit mask, and match perspective, contact shadows, and color. The approved product layer remains product authority; the generated scene does not.

When a believable grip cannot be reconstructed without inventing hidden product geometry, change the composition. Place the exact product on a nearby surface, keep the person in the scene without touching it, or use the generated candidate as a photographer's reference instead of forcing a false interaction.

7. Measure accepted assets and the downstream decision

Track attempts, accepted candidates, rejection reasons, generation cost, reviewer time, correction time, compositing time, and final placement. Cost per generated image is not cost per accepted asset. When a released secondary image enters a real test, measure the declared commercial outcome with product-truth, support, return, page-performance, and placement guardrails.

The same-SKU fidelity benchmark is the next step when the model choice remains unresolved. The supplier-photo ecommerce set shows how one source becomes a governed deliverable package. If the job is a routine-style ad rather than a still, use the real-product AI UGC workflow for actor rights, disclosure, motion, claims, and contribution measurement. The broader product-photography model comparison maps thirty category-specific failure surfaces, while the AI ecommerce workflow map routes the job by its revenue stage.

Bottom line

GPT Image 2 followed this single person-with-product brief most completely. Nano Banana 2 produced another credible secondary candidate with a tighter grip. Qwen Image Edit Plus returned a stylized direction that failed the declared photo and interaction gates. Seedream 5 Pro was not run because the submission lacked sufficient credits.

The larger result is more durable than that ordering: hands turn a lifestyle edit into an interaction and product-truth test. Lock the source, score the product before the scene, preserve every returned candidate, and composite or photograph the real product when generated contact cannot carry factual authority.

Share:
FAQ

Questions from this guide

Concise answers to the questions readers ask after this guide

Which AI model was best for lifestyle product photography in this test?

GPT Image 2 came closest to this one fixed brief: it kept the bottle recognizable, preserved the visible text and cap loop, used a plausible side-edge grip, and matched the shoulders-to-hips overlook composition. Nano Banana 2 also preserved the core product and created a convincing scene, but chose a tighter wraparound grip. Qwen Image Edit Plus returned a stylized image and missed several declared constraints. One candidate per route is a capability screen, not a universal model ranking.

Can AI put a real product in a person's hand without changing it?

It can produce a recognizable secondary candidate, but the interaction forces the model to redraw hidden edges, scale, contact, fingers, shadows, and often the product itself. Compare every candidate with approved sources at full resolution. When exact pixels matter, photograph the real hand-product interaction or composite the approved product with a controlled occlusion mask.

Should an AI lifestyle image replace the main product photo?

Not by default. Keep a factual primary image controlled by the exact SKU and channel rules. Use an AI lifestyle image as a reviewed secondary candidate, a creative-test input, or a photographer's scene brief. Placement-specific marketplace and advertising requirements still apply.

How should ecommerce teams test AI lifestyle product photos?

Use one approved source, one fixed prompt, the same aspect ratio, and equal candidate counts. Predeclare product identity, geometry, color, finish, markings, quantity, interaction, and scene gates. Record every returned candidate, rejection reason, reviewer time, correction path, and final disposition before comparing cost per accepted asset.

What if every generated person-with-product image changes the SKU?

Stop asking the model to redraw the product. Generate or photograph the person and environment separately, then composite the approved product using controlled scale, perspective, hand occlusion, contact shadow, and color management. If credible contact cannot be built without inventing product evidence, use a product-near-person scene or commission the real interaction.