There is no defensible “best product-photography model” in this thirty-category corpus. The tests used fictional text briefs and one displayed output per model. They reveal what each category makes easy to overlook—an invented screen, extra package, changed prong count, shifted shade, unsupported safety cue, or plausible-looking label—but they do not measure preservation of a real SKU.
That makes this page an evidence map, not a winner board. Use it to identify the rejection sheet your product needs, then run a reference-first benchmark on the actual item.
Methodology boundary: the thirty linked category tests are controlled text-to-image comparisons of fictional product briefs. Each article preserves four original outputs and documents visible evidence. Some original prompts were not preserved and are now represented by controlled reconstructions. The tests are not a reference-image benchmark of exact real SKUs, do not provide repeated-run acceptance rates, and cannot support universal winner, reliability, or cost-per-usable-image claims.
Quick answer
- Choose by task and evidence, not category hype. Current model documentation describes capabilities; the linked images show one result. Neither replaces a benchmark on your source product.
- Score product truth before visual craft. Reject changed geometry, color, finish, labels, marks, quantity, condition, variants, safety cues, and product claims before comparing lighting or mood.
- Measure a workflow, not a single generation. The useful unit is accepted assets divided by total attempts, review time, and cost—including compositing and corrections.
- Prefer bounded generation. Generate the empty scene and composite an approved packshot when exact product identity matters.
- Re-test after change. Model versions, endpoints, prompts, source preparation, and post-production can all change the result.
Download the blank model benchmark scorecard
Download the four-route product-photo model scorecard. It is deliberately unscored: each row starts at NOT_RUN, so the template cannot smuggle invented performance into your decision. Replace the placeholders, keep the approved source set, scene brief, crop, output size, and four-candidate count fixed, then store every output in the matching evidence folder.
Calculate acceptance rate as accepted candidates divided by total candidates. Calculate total route cost as generation cost, external cost, and loaded review plus correction time. Divide that total by accepted candidates for cost per accepted asset. If a route accepts nothing, mark it failed instead of reporting a misleading zero-dollar result. A rejected candidate can trigger several failure classes, so identity, data, geometry, appearance, rights or marking, and claim counts do not have to sum to the number rejected.
The current model shortlist
This is a capability shortlist, not a ranking from the thirty category images.
| Model used in the corpus | Current documented reason to test it | What the corpus can say | What it cannot say |
|---|---|---|---|
| Nano Banana 2 | Google's current guide calls Gemini 3.1 Flash Image its versatile workhorse, documents 512px through 4K output, and supports multi-reference workflows. | One fictional output per category exposes scene, detail, text, and design choices. | Real-SKU preservation rate, “best all-rounder,” or expected acceptance cost. |
| GPT Image 2 | OpenAI documents generation and editing, automatic high-fidelity processing of image inputs, and flexible output sizes up to a 3840px edge. | Some displayed text and structured elements are highly legible, including unsupported product data. | That legible copy is correct, or that the model preserves an uploaded SKU without review. |
| Seedream 4.5 | It is the Seedream version used for every linked category output, so it remains useful as a reproducible corpus baseline in Masonry. | The displayed outputs provide direct evidence of visual composition, material cues, and category-specific drift. | That it “wins” materials, optics, cost, or real-product fidelity; Seedream 5.0 Pro was not tested here. |
| FLUX.2 Pro | Black Forest Labs documents a production endpoint, up to ten references, and a fixed flux-2-pro snapshot for reproducible workflows. | The linked concepts reveal useful geometry, material, marking, and prompt-adherence checks. | That it is the cheapest usable workflow or takes more liberties across repeated SKU tests. |
Primary documentation: Google image generation, OpenAI image generation, and the FLUX.2 overview. Google also says the Imagen 4 standard, Ultra, and Fast endpoints are deprecated and will shut down on August 17, 2026, with Gemini 3.1 Flash Image as the migration path.
Two outputs that explain the benchmark
These images do not prove that Seedream always makes these errors. They show why a single attractive output cannot be accepted without a source comparison.
The 30-category evidence map
| Product test | What the displayed test helps you inspect | What a real listing must source and verify |
|---|---|---|
| Skincare | Frosted glass, small label, cap, liquid and scene cues | Exact container, formula state, label, quantity, ingredients, claims, warnings, and variant |
| Jewelry | Stone presentation, setting contacts, band, reflection, view | Exact piece or disclosed representative design, stones, setting, metal, marks, condition, scale, and grading data |
| Supplements | Bottle, capsule, label and panel-like structure | Approved formula, Supplement Facts, ingredients, warnings, claims, code, lot area, and package |
| Makeup | Relative shade, finish, bullet, case, cap and invented marks | Formula, shade, color-managed capture, component, labels, swatch protocol, claims, variant, and assortment |
| Food & beverage | Serving scene, material, condensation, garnish and package cues | Exact food or beverage, quantity, ingredients, preparation, package, claims, condition, and serving truth |
| Footwear | Silhouette, multi-material construction, sole and marks | Exact upper, sole, stitching, colorway, size cues, condition, branding, variant, and pair consistency |
| Candles | Vessel, wax, wick, flame, glow and scene | Exact vessel, fill, wick count, label, scent variant, burn state, warnings, and safe-use claims |
| Clothing | Garment silhouette, print, fabric and styling | Exact cut, size, pattern or artwork, construction, material, color, fit, marks, and variant |
| Furniture | Form, room scale, materials, assembly and shadows | Exact dimensions, geometry, joins, hardware, finish, color, condition, configuration, and included parts |
| Electronics | Screen content, body, controls, finish and wear | Exact device, UI, controls, ports, sensors, marks, accessories, specs, compatibility, and safety evidence |
| Handbags | Silhouette, grain, stitching, hardware and color | Exact construction, material, lining, closure, hardware, marks, dimensions, variant, and condition |
| Sunglasses | Frame symmetry, lens appearance, hinges and reflection | Exact frame, lens, tint, coatings, marks, dimensions, fit, included case, and UV or safety claims |
| Glassware | Geometry, rim, stem, refraction, reflection and crop | Exact item or disclosed handmade range, set count, capacity, dimensions, marks, condition, and use claims |
| Flowers | Organic form, petal detail, arrangement and freshness cues | Exact arrangement or disclosed representative range, stem and bloom count, species, color, size, and condition |
| Watches | Dial, hands, indices, sub-dials, case, crown and strap | Exact watch, time or state, dial copy, movement-facing claims, dimensions, materials, marks, condition, and set |
| Perfume | Bottle geometry, glass, liquid, cap, label and reflection | Exact bottle, fill, color, cap, atomizer, label, volume, ingredients or warnings when shown, and variant |
| Packaging / CPG | Structure, panels, copy, codes, count and print appearance | Approved dieline, artwork, identity, quantity, ingredients, allergens, facts, claims, identifiers, and print proof |
| Pet products | Product form, animal interaction, fit and scene | Exact product, size, attachment, materials, animal-fit boundaries, warnings, included items, and claims |
| Toys | Part shapes, assembly cues, materials, marks and visual resemblance | Owned design, part inventory, count, assembly, age grading, warnings, safety evidence, package, and set |
| Textiles & bedding | Weave, drape, pattern, seams and pile | Exact pattern repeat, material, dimensions, construction, color-managed reference, variant, and set count |
| Cookware | Geometry, reflections, handles, lid and scene | Exact pan or set, dimensions, materials, coatings, hardware, marks, included items, condition, and use or safety claims |
| Stationery | Paper, foil, deboss, binding, text and layout | Exact artwork, copy, page or item count, dimensions, stock, finish, binding, personalization, and set |
| Drinkware | Shape, finish, lid, handle, reflection and marks | Exact vessel, capacity, dimensions, lid or straw, seals, material, color, marks, set, and use claims |
| Soap & bath | Bar shape, swirl, texture, cut and styled batch | Exact bar or disclosed batch range, dimensions, weight, formula, ingredients, label, variation, and claims |
| Ceramics | Form, handle, foot, glaze and firing variation | Exact piece or disclosed range, dimensions, capacity, weight, condition, glaze variation, marks, and food or heat claims |
| Art prints | Frame, mat, glass, glare, room scale and artwork placement | Exact owned artwork, crop, color, edition, dimensions, paper, frame, glazing, marks, and included items |
| Earbuds | Bud and case geometry, ports, seams, marks and resemblance | Exact buds, charging case, controls, indicators, accessories, specs, compatibility, safety evidence, and set |
| Houseplants | Species-like traits, foliage, pot, scale and condition | Exact stock or disclosed representative plant, species, size range, pot, condition, count, care, and toxicity claims |
| Knives & cutlery | Blade shape, pattern, bevel, handle and reflection | Exact blade and handle, dimensions, pattern, edge, marks, sheath or set, condition, materials, and safety claims |
| Automotive wheels | Spoke topology, rim, finish, marks and brake or tire context | Exact wheel, dimensions, PCD, bore, offset, load data, finish, marks, fitment, accessories, and safety evidence |
The recurring failure classes are now more useful than a model leaderboard:
- Identity drift: a plausible category object replaces the exact SKU.
- Data drift: text, UI, facts panels, barcodes, dials, warnings, ingredients, or specifications look authoritative without an approved source.
- Geometry drift: parts, prongs, spokes, handles, controls, seams, set count, and repeated structures change.
- Appearance drift: color, finish, texture, condition, natural variation, reflection, and transparency become a more idealized product.
- Rights and marking drift: logos, lettering, familiar configurations, certifications, or decorative marks appear without approval.
- Claim drift: pixels imply fitment, performance, safety, health, compatibility, material, capacity, or included features they cannot prove.
A benchmark that produces a business answer
1. Define the product promise
State whether the image represents an exact SKU, one-of-one item, disclosed representative range, fixed set, variant, assortment, or fictional concept. The acceptance rule changes with that promise.
2. Assemble approved sources
Collect the views, dimensions, artwork, product data, marks, variants, condition notes, color references, and safety or compliance records needed for that product. Label versions and markets.
3. Predeclare invariants and allowed changes
Write pass conditions before generation. Keep product identity, data, marks, count, and claims fixed. List the environment, props, crop, light, or background that may change.
4. Run multiple candidates per model
Use the same source set, prompt, aspect ratio, output size, and candidate count. Four or more candidates per model begins to expose variance; one attractive sample does not.
masonry image "Change only the environment to polished warm stone with soft camera-left light. Keep the supplied product's geometry, colors, materials, finish, marks, label, condition, quantity, crop, and perspective unchanged. Add no text, logo, feature, claim, prop, or extra product." \ --model gemini-3.1-flash-image-preview \ --ref ./approved-product-three-quarter.png \ --aspect 1:1 masonry job wait <job-id> masonry job download <job-id> --output ./product-scene-candidate.png
The instruction is a constraint, not a product lock.
5. Review blind and record failure reasons
Check product truth before identifying the model or judging style. Record accept or reject plus the exact failure class. Review full resolution, listing size, alternate views, variants, thumbnails, and the rendered page.
6. Compare workflow economics
Track accepted assets, total attempts, generation cost, reviewer time, retouching or compositing time, and rework. The useful metric is cost per accepted asset under a defined standard, not nominal cost per generation.
7. Prefer compositing when exactness dominates
Generate the empty environment, then composite a color-managed approved packshot. Rebuild shadows and reflections deliberately. Use generation for the part allowed to vary.
The AI product-background route fidelity test shows why this is a production choice rather than a prompt trick: both a general edit and a structured placement returned polished but changed packages, while a deterministic source-card route preserved the approved product evidence without pretending to be a seamless composite.
8. Test the accepted image against the commercial control
Fidelity approval makes an image eligible to test; it does not show that the image helps sales. Compare the accepted candidate with the current image on the same SKU, offer, audience, placement, and decision window. Choose the primary commercial metric before exposure—completed-order contribution is stronger than clicks alone—and keep return rate, “not as described” contacts, refunds, and support complaints as guardrails.
The AI product-photo trust test provides the control-versus-candidate plan and explains why an attractive click lift can still be a business loss. Keep the model, generation, and acceptance records joined to the exposed asset so a result can be traced back to the exact source, candidate, and review decision.
Model notes without winner claims
Nano Banana 2
Google's current guide positions Gemini 3.1 Flash Image as a versatile workhorse with text-and-image editing, 1K through 4K output plus 512px, and multi-reference support. Those capabilities make it a sensible pilot candidate. They do not guarantee that ten supplied product views will all be preserved.
GPT Image 2
OpenAI documents generation, image edits, masks, automatic high-fidelity processing of image inputs, and flexible sizes. The category images show why legibility needs a second score: clean text, UI, or panels can still be entirely unsupported by product data.
Seedream 4.5
Seedream 4.5 is the fixed version represented across all thirty linked tests. It often provides useful scene and material directions in the displayed outputs, but the same corpus includes extra products, changed geometry, invented marks, and unsupported copy. Seedream 5.0 Pro exists but is outside this benchmark.
FLUX.2 Pro
Black Forest Labs documents FLUX.2 Pro for production workflows, multi-reference editing, and a fixed endpoint for reproducibility. Reproducibility helps operations; it does not turn a fictional category result into evidence of SKU fidelity.
Bottom line
The thirty tests answer “what should I inspect?” much better than “which model wins?” Their first-hand images expose the category-specific places where visual plausibility outruns product truth. Use the map to build a rejection sheet, benchmark current models on approved sources, measure acceptance rather than showcase quality, and composite the real product when exactness dominates. The same-SKU product-fidelity test applies one exact source and fixed prompt across four current models. For glass, liquid, metal, and mirror-surface failure modes, the transparent and reflective product image model test compares six current same-source edits and keeps its unequal sample sizes visible. For the separate problem of putting a product in a person's hand, the same-SKU lifestyle product-photo test compares three first returns against grip, occlusion, scale, marking, and scene gates. The complete ecommerce image-set test shows the production handoff on one controlled source, including the accepted files and the geometry drift that kept them out of the factual-primary slot.
To turn a selected model into a repeatable production brief, use the 12 ecommerce product-photography prompts; each prompt separates immutable product truth from the scene and includes a merchant acceptance gate.
Before generating a gallery, turn the unanswered purchase questions into an ecommerce product photography shot list; its seven-slot manifest separates factual capture, controlled compositing, reviewed AI scenes, and rows that must stay blocked until product evidence exists.
Once one SKU passes that gate, use the three-SKU batch product-photography workflow to carry the approved recipe into a manifest, asynchronous jobs, separate product-truth and set-consistency review, and cost-per-accepted-image accounting before scaling the catalog.
Run the same controlled source through Masonry's AI product photography studio, or automate the candidate set and file naming with the Masonry CLI. The decision should come from your accepted assets, not this page's favorite model.
Once a pilot supplies a real acceptance rate, move the attempts, reviewer time, correction time, credits, and external spend into the AI product-photography cost calculator. It compares unlike production routes on the same accepted deliverable and translates project cost into a conservative break-even order threshold.
If the model comparison is only one part of a larger merchant job, use the AI for ecommerce workflow map to route catalog, PDP, acquisition, lifecycle, evidence, and wholesale work by its authority record and commercial finish line.


