Masonry Logo
AI & Technology

Best AI Model for Transparent and Reflective Product Images: 6-Image Test

We edited one fictional clear serum bottle with three current AI image models. See six real outputs, the fixed prompt, material-fidelity scorecard, and the approval workflow for glass, liquid, metal, labels, and reflections.

Gaurav BisenGaurav Bisen
9 min read

Transparent and reflective products are a difficult ecommerce image-editing case because the product is not only its outline. The cap edge, liquid boundary, internal pump, metal finish, label, highlights, and surface reflection all contribute to what the shopper believes is being sold.

We tested one fictional clear serum bottle with one fixed reference and scene brief. Nano Banana 2 was the most consistent across the three candidates we could run. Seedream 5 Pro created the strongest dark reflective treatment, but changed product proportions more. GPT Image 2 kept the words and core parts, but missed the requested dark cobalt look. None of the six outputs should silently replace an approved factual packshot.

Evidence boundary: NORTHLINE is a synthetic product created for a controlled Masonry test, not a merchant SKU. The six images below are unselected returned outputs from August 15, 2026: three Nano Banana 2 candidates, two Seedream 5 Pro candidates, and one GPT Image 2 candidate. Generation credits ran out before equal sample sizes were complete, so we report raw counts and never convert the unequal set into a universal win rate. This is not a conversion experiment or a physics benchmark.

The source and the product-truth gate

Controlled source, 1254 × 1254. The fixed visible identity is one clear cylindrical bottle, a clear cap, three-tier brushed-silver pump, circular nozzle, orange liquid, white rectangular label, and exact text: NORTHLINE, VITAMIN C, 30 ML.

We separated recognizability from exact-SKU approval before reviewing the results.

GateRequired evidence
Product structureOne cylindrical clear bottle; clear cap; three visible silver pump tiers; circular front nozzle
ContentsOrange transparent liquid with a visible horizontal fill boundary
LabelOne centered white rectangle; exact NORTHLINE, VITAMIN C, and 30 ML text
Quantity and claimsOne product; no props, added products, logo, claim, or extra text
Requested scenePolished dark cobalt surface; soft blue-gray background; camera-left highlight; contact shadow; one controlled reflection
Exact-SKU approvalNo material change to proportions, cap height, pump relationships, bottle base, label geometry, fill, or visible internal parts

The first four gates tell us whether a candidate still reads as the same fictional product. The last gate is intentionally stricter. A clean image can pass the core checklist and still fail as factual primary photography because a generator has redrawn the object.

Exact prompt and run contract

Every job received this prompt and the same source file:

Edit the supplied fictional product reference into a square ecommerce secondary image on a polished dark cobalt stone surface with a soft blue-gray gradient backdrop. Keep the product front-facing at the same scale and preserve the exact clear cylindrical bottle and clear cap, the three-tier brushed-silver pump, circular nozzle, orange transparent liquid and fill line, white rectangular label, and the exact three text lines NORTHLINE, VITAMIN C, and 30 ML. Add one soft camera-left stripbox highlight, a natural contact shadow, and one controlled mirror reflection directly below the bottle. Do not add props, hands, droplets, bubbles, extra products, extra text, logos, or claims. Do not change bottle geometry, cap height, pump tiers, nozzle, liquid color, fill level, label shape, spelling, spacing, or quantity.

The live Masonry routes were gemini-3.1-flash-image-preview (Nano Banana 2), seedream-5-pro, and gpt-image-2. Nano Banana 2 used seeds 41101, 41102, and 41103; Seedream 5 Pro used 51201 and 51202; GPT Image 2 did not expose a user-set seed in this run. All jobs used one reference and a square output.

Prompt

masonry image "<fixed prompt above>" \ --model gemini-3.1-flash-image-preview \ --ref ./northline-source.webp \ --aspect 1:1 \ --seed 41101 masonry job wait <job-id> masonry job download <job-id> --output ./candidate.png

Results: six real candidates

Nano Banana 2: the most repeatable core identity

Nano Banana 2 candidate 1, 1024 × 1024. All locked words and core parts survive. The newly visible dip tube and changed product rendering still require exact-SKU review.
Nano Banana 2 candidate 2. Text, clear cap, pump tiers, liquid, and quantity remain consistent. Gold-like flecks in the surface and a taller product treatment depart from the tightest reading of the brief.
Nano Banana 2 candidate 3. This was the closest balanced scene in its three-image set, while the visible internal tube and redrawn glass edges still separate it from the factual source.

All three candidates passed the core structure, contents, label, and quantity gates. The consistency matters more than one attractive sample: the exact three label lines and pump stack survived three separate seeds. Yet each image remains a redraw. The dip tube becomes visible in two candidates even though it is not visible in the source, and small bottle-to-label relationships move.

Disposition: strongest model in this limited run for secondary-creative review; zero of three approved as an unquestioned factual-primary replacement.

Seedream 5 Pro: strongest dark reflection, more proportion drift

Seedream 5 Pro candidate 1, 1536 × 1536. The requested dark surface and reflection are strong. The bottle, label, cap, and pump relationships are visibly redrawn.
Seedream 5 Pro candidate 2. The material separation is clean and all label text is exact, but the taller narrow bottle and large label are not the same geometry as the source.

Both candidates passed the core identity checklist and delivered the clearest mirror-style floor reflections. They also made the strongest proportion changes: the body becomes narrower and taller, the label scale moves, and the cap-to-bottle relationship changes.

Disposition: good generated-environment reference; zero of two approved as factual-primary replacements. For a real SKU, keep the scene and composite the approved product.

GPT Image 2: accurate words, weaker scene match

GPT Image 2 candidate, 1024 × 1024. The product remains coherent and every label line is exact. The backdrop is much brighter than requested and the bottle and label proportions shift.

The one GPT Image 2 candidate passed the core identity checklist and produced the cleanest label typography. One candidate is not enough to judge repeatability. It also missed the darker scene treatment and redrew the bottle relationship enough to fail the strict exact-SKU gate.

Disposition: usable direction for a secondary image after review; zero of one approved as a factual-primary replacement.

Scorecard: what the six images actually support

ModelCandidatesCore identity checklistExact visible textDark reflective sceneExact-SKU primary approvalPractical read
Nano Banana 233 / 33 / 33 / 30 / 3Most consistent starting point for secondary creative
Seedream 5 Pro22 / 22 / 22 / 20 / 2Strongest reflection; composite when geometry matters
GPT Image 211 / 11 / 10 / 10 / 1Clean text; insufficient sample and weaker scene adherence

These denominators are deliberately visible. “3 / 3” is not a 100% reliability claim, and the single GPT result cannot be compared as if it had equal statistical weight. The useful result is narrower: all three Nano Banana 2 attempts retained the required core attributes in this exact setup, while all six candidates needed a stricter product-truth review.

Why transparent and reflective products fail differently

Opaque products let reviewers concentrate on silhouette, color, marks, and surface texture. Transparent and mirrored materials add several failure paths:

  • Edge drift: a clear cap or glass base can widen, taper, thicken, or acquire a second rim.
  • Internal-part invention: pumps, tubes, stems, liquid menisci, bubbles, or seams appear because the model completes what it expects to see.
  • Reflection mismatch: a floor reflection can duplicate incorrect label text, imply another quantity, or disagree with the product above it.
  • Highlight-driven geometry: a plausible strip highlight can hide a changed shoulder, wall thickness, facet, bevel, or polished edge.
  • Material substitution: clear acrylic, coated glass, chrome, satin metal, foil, polished ceramic, and translucent liquid can collapse into the same generic glossy treatment.

Merchant discussions about AI product imagery repeatedly call out text, precise colors, reflective surfaces, and transparency as failure areas. Those discussions are qualitative intent evidence rather than performance data, but they identify a concrete job: preserve the SKU while creating more usable placements. See the current ecommerce discussion about AI product visuals and an independent multi-generator product test.

A production protocol that protects the SKU

  1. Keep an approved primary image. Store the unedited packshot, artwork, dimensions, materials, finish, fill, variant data, and rights record with the SKU.
  2. Separate product truth from scene instructions. List what cannot change, then list only the environment, crop, and lighting that may change.
  3. Run equal candidate counts. Use the same source, prompt, size, ratio, and number of attempts. This run shows why a credit interruption must be disclosed rather than hidden.
  4. Review the product before the reflection. Compare silhouette, cap, hardware, liquid, label, text, and quantity at full resolution. Then inspect highlights, shadows, and the reflection.
  5. Choose a disposition. Approve as a secondary candidate, composite the approved product into the generated scene, or reject. Do not rescue identity drift with a crop.
  6. Measure cost per accepted asset. Add generation spend, reviewer time, correction time, and compositing cost; divide by outputs that pass the declared release gate.

For category-specific acceptance sheets, use the AI glassware product-photography guide, AI perfume product-photography guide, and AI jewelry product-photography guide. The broader 30-category image-model comparison maps where transparency, metal, glass, text, and geometry fail across product types. The same-SKU fidelity test compares four models in a neutral stone scene; this test adds repeat attempts and a deliberately reflective surface.

When exact pixels dominate, use Masonry to generate the empty environment and composite an approved color-managed packshot. When secondary creative can tolerate controlled review, run the fixed source through Masonry's AI product photography workflow. Use the Masonry CLI to keep prompts, model routes, seeds, job IDs, and filenames auditable.

Bottom line

Nano Banana 2 gave the most repeatable core-identity result in this six-image test. Seedream 5 Pro made the strongest reflective scenes. GPT Image 2 kept the label exact in its single result but did not match the requested dark treatment. No output earned permission to replace the factual primary image without merchant review. For transparent and reflective SKUs, the revenue-safe workflow is to optimize for accepted secondary assets, keep product truth fixed, and composite the approved product when exact geometry matters.

Share:
FAQ

Questions from this guide

Concise answers to the questions readers ask after this guide

Which AI model was best for this transparent product image test?

Nano Banana 2 was the most consistent in this six-image run: all three candidates kept the clear cap, three-tier silver pump, orange liquid, white label, and exact visible text. Seedream 5 Pro produced stronger dark-surface reflections but changed bottle and label proportions more. GPT Image 2 kept the text but missed the requested darker scene. This is one synthetic SKU and unequal sample sizes, not a universal model ranking.

Can AI preserve an exact glass or transparent product?

A reference edit can stay recognizably close, but it still redraws transparent edges, internal hardware, liquid boundaries, highlights, and reflections. Keep an approved factual packshot as the primary image. Treat generated scenes as secondary candidates, or composite the approved product into a generated environment when exact geometry and material behavior matter.

How should ecommerce teams test AI images for reflective products?

Declare the fixed silhouette, transparent parts, reflective hardware, liquid color and fill, label geometry, exact text, quantity, and allowed environment changes before generation. Run the same source, prompt, ratio, and candidate count across models, then review at full resolution before judging aesthetics.

Should a mirror reflection count as product evidence?

No. A generated floor reflection is a scene treatment, not an additional product view or proof of material, finish, dimensions, or condition. Review the product itself against approved sources and score reflection quality separately.

What metric should merchants use to compare AI image models?

Use cost per accepted asset: total generation spend plus reviewer and correction time divided by candidates that pass a predeclared product-truth gate. A low generation price is not useful if most outputs require rejection or compositing.