Transparent and reflective products are a difficult ecommerce image-editing case because the product is not only its outline. The cap edge, liquid boundary, internal pump, metal finish, label, highlights, and surface reflection all contribute to what the shopper believes is being sold.
We tested one fictional clear serum bottle with one fixed reference and scene brief. Nano Banana 2 was the most consistent across the three candidates we could run. Seedream 5 Pro created the strongest dark reflective treatment, but changed product proportions more. GPT Image 2 kept the words and core parts, but missed the requested dark cobalt look. None of the six outputs should silently replace an approved factual packshot.
Evidence boundary: NORTHLINE is a synthetic product created for a controlled Masonry test, not a merchant SKU. The six images below are unselected returned outputs from August 15, 2026: three Nano Banana 2 candidates, two Seedream 5 Pro candidates, and one GPT Image 2 candidate. Generation credits ran out before equal sample sizes were complete, so we report raw counts and never convert the unequal set into a universal win rate. This is not a conversion experiment or a physics benchmark.
The source and the product-truth gate
We separated recognizability from exact-SKU approval before reviewing the results.
| Gate | Required evidence |
|---|---|
| Product structure | One cylindrical clear bottle; clear cap; three visible silver pump tiers; circular front nozzle |
| Contents | Orange transparent liquid with a visible horizontal fill boundary |
| Label | One centered white rectangle; exact NORTHLINE, VITAMIN C, and 30 ML text |
| Quantity and claims | One product; no props, added products, logo, claim, or extra text |
| Requested scene | Polished dark cobalt surface; soft blue-gray background; camera-left highlight; contact shadow; one controlled reflection |
| Exact-SKU approval | No material change to proportions, cap height, pump relationships, bottle base, label geometry, fill, or visible internal parts |
The first four gates tell us whether a candidate still reads as the same fictional product. The last gate is intentionally stricter. A clean image can pass the core checklist and still fail as factual primary photography because a generator has redrawn the object.
Exact prompt and run contract
Every job received this prompt and the same source file:
Edit the supplied fictional product reference into a square ecommerce secondary image on a polished dark cobalt stone surface with a soft blue-gray gradient backdrop. Keep the product front-facing at the same scale and preserve the exact clear cylindrical bottle and clear cap, the three-tier brushed-silver pump, circular nozzle, orange transparent liquid and fill line, white rectangular label, and the exact three text lines NORTHLINE, VITAMIN C, and 30 ML. Add one soft camera-left stripbox highlight, a natural contact shadow, and one controlled mirror reflection directly below the bottle. Do not add props, hands, droplets, bubbles, extra products, extra text, logos, or claims. Do not change bottle geometry, cap height, pump tiers, nozzle, liquid color, fill level, label shape, spelling, spacing, or quantity.
The live Masonry routes were gemini-3.1-flash-image-preview (Nano Banana 2), seedream-5-pro, and gpt-image-2. Nano Banana 2 used seeds 41101, 41102, and 41103; Seedream 5 Pro used 51201 and 51202; GPT Image 2 did not expose a user-set seed in this run. All jobs used one reference and a square output.
masonry image "<fixed prompt above>" \ --model gemini-3.1-flash-image-preview \ --ref ./northline-source.webp \ --aspect 1:1 \ --seed 41101 masonry job wait <job-id> masonry job download <job-id> --output ./candidate.png
Results: six real candidates
Nano Banana 2: the most repeatable core identity
All three candidates passed the core structure, contents, label, and quantity gates. The consistency matters more than one attractive sample: the exact three label lines and pump stack survived three separate seeds. Yet each image remains a redraw. The dip tube becomes visible in two candidates even though it is not visible in the source, and small bottle-to-label relationships move.
Disposition: strongest model in this limited run for secondary-creative review; zero of three approved as an unquestioned factual-primary replacement.
Seedream 5 Pro: strongest dark reflection, more proportion drift
Both candidates passed the core identity checklist and delivered the clearest mirror-style floor reflections. They also made the strongest proportion changes: the body becomes narrower and taller, the label scale moves, and the cap-to-bottle relationship changes.
Disposition: good generated-environment reference; zero of two approved as factual-primary replacements. For a real SKU, keep the scene and composite the approved product.
GPT Image 2: accurate words, weaker scene match
The one GPT Image 2 candidate passed the core identity checklist and produced the cleanest label typography. One candidate is not enough to judge repeatability. It also missed the darker scene treatment and redrew the bottle relationship enough to fail the strict exact-SKU gate.
Disposition: usable direction for a secondary image after review; zero of one approved as a factual-primary replacement.
Scorecard: what the six images actually support
| Model | Candidates | Core identity checklist | Exact visible text | Dark reflective scene | Exact-SKU primary approval | Practical read |
|---|---|---|---|---|---|---|
| Nano Banana 2 | 3 | 3 / 3 | 3 / 3 | 3 / 3 | 0 / 3 | Most consistent starting point for secondary creative |
| Seedream 5 Pro | 2 | 2 / 2 | 2 / 2 | 2 / 2 | 0 / 2 | Strongest reflection; composite when geometry matters |
| GPT Image 2 | 1 | 1 / 1 | 1 / 1 | 0 / 1 | 0 / 1 | Clean text; insufficient sample and weaker scene adherence |
These denominators are deliberately visible. “3 / 3” is not a 100% reliability claim, and the single GPT result cannot be compared as if it had equal statistical weight. The useful result is narrower: all three Nano Banana 2 attempts retained the required core attributes in this exact setup, while all six candidates needed a stricter product-truth review.
Why transparent and reflective products fail differently
Opaque products let reviewers concentrate on silhouette, color, marks, and surface texture. Transparent and mirrored materials add several failure paths:
- Edge drift: a clear cap or glass base can widen, taper, thicken, or acquire a second rim.
- Internal-part invention: pumps, tubes, stems, liquid menisci, bubbles, or seams appear because the model completes what it expects to see.
- Reflection mismatch: a floor reflection can duplicate incorrect label text, imply another quantity, or disagree with the product above it.
- Highlight-driven geometry: a plausible strip highlight can hide a changed shoulder, wall thickness, facet, bevel, or polished edge.
- Material substitution: clear acrylic, coated glass, chrome, satin metal, foil, polished ceramic, and translucent liquid can collapse into the same generic glossy treatment.
Merchant discussions about AI product imagery repeatedly call out text, precise colors, reflective surfaces, and transparency as failure areas. Those discussions are qualitative intent evidence rather than performance data, but they identify a concrete job: preserve the SKU while creating more usable placements. See the current ecommerce discussion about AI product visuals and an independent multi-generator product test.
A production protocol that protects the SKU
- Keep an approved primary image. Store the unedited packshot, artwork, dimensions, materials, finish, fill, variant data, and rights record with the SKU.
- Separate product truth from scene instructions. List what cannot change, then list only the environment, crop, and lighting that may change.
- Run equal candidate counts. Use the same source, prompt, size, ratio, and number of attempts. This run shows why a credit interruption must be disclosed rather than hidden.
- Review the product before the reflection. Compare silhouette, cap, hardware, liquid, label, text, and quantity at full resolution. Then inspect highlights, shadows, and the reflection.
- Choose a disposition. Approve as a secondary candidate, composite the approved product into the generated scene, or reject. Do not rescue identity drift with a crop.
- Measure cost per accepted asset. Add generation spend, reviewer time, correction time, and compositing cost; divide by outputs that pass the declared release gate.
For category-specific acceptance sheets, use the AI glassware product-photography guide, AI perfume product-photography guide, and AI jewelry product-photography guide. The broader 30-category image-model comparison maps where transparency, metal, glass, text, and geometry fail across product types. The same-SKU fidelity test compares four models in a neutral stone scene; this test adds repeat attempts and a deliberately reflective surface.
When exact pixels dominate, use Masonry to generate the empty environment and composite an approved color-managed packshot. When secondary creative can tolerate controlled review, run the fixed source through Masonry's AI product photography workflow. Use the Masonry CLI to keep prompts, model routes, seeds, job IDs, and filenames auditable.
Bottom line
Nano Banana 2 gave the most repeatable core-identity result in this six-image test. Seedream 5 Pro made the strongest reflective scenes. GPT Image 2 kept the label exact in its single result but did not match the requested dark treatment. No output earned permission to replace the factual primary image without merchant review. For transparent and reflective SKUs, the revenue-safe workflow is to optimize for accepted secondary assets, keep product truth fixed, and composite the approved product when exact geometry matters.


