When I review an AI product photography workflow, I no longer start with a prompt like “make this look professional.”
That sentence sounds useful. It is not. It says nothing about which product details must survive, which parts may change, where the image will appear, or what would make the result unsafe to publish.
I start with a product truth sheet and one small calibration instead.
Direct answer
Do not ask an image model for a full campaign on the first run.
First, choose one image job. Record the visible and verified product facts that must not change. List what the source image does not show and what the output must never imply. Confirm rights. Then make one or two calibration images.
I only expand the set after geometry, color, logo, label text, scale, scene, and crop pass review.
This is less exciting than writing a giant cinematic prompt. It is also much easier to audit.
1. I define the image’s job before choosing a tool
“Product photo” can mean several different things:
- a clean marketplace main image;
- a background replacement;
- a product-page lifestyle image;
- an organic social post;
- a paid-ad draft;
- an on-model apparel image;
- an early concept that will never be used as a listing.
Those jobs should not share one acceptance test.
A marketplace image may need the product to remain almost untouched. A lifestyle image can add a room, surface, prop, and atmosphere, but those additions can imply scale or usage. An apparel image must keep the garment’s construction while also handling the model, pose, anatomy, and body contact.
Before I open a generator, I write one sentence:
This image will be used for [destination] to show [one product fact or use context]. It will not be treated as [listing evidence, performance proof, or another excluded use].
If I cannot finish that sentence, I am not ready to generate.
2. I build a one-page product truth sheet
The product truth sheet is the part I wish more prompt tutorials included.
It has four sections.
Visible facts
These are details a reviewer can confirm from the approved source images:
- silhouette and proportions;
- visible color and finish;
- logo position;
- label layout;
- seams, pockets, buttons, closures, ports, handles, and included parts;
- the visible front, side, and back structure.
Verified facts that are not visible
These come from a current product specification or another approved source:
- dimensions;
- material;
- capacity;
- compatibility;
- included accessories;
- official variant and color names.
I do not add these because they “seem likely.” A plausible detail is still an invented detail.
Unknown details
This is where I record what the model is not allowed to guess:
- an unseen back label;
- a hidden closure;
- the inside of the package;
- the underside of a device;
- the way a fabric behaves when worn;
- a product variant with no source photo.
The word unknown does real work. It gives the reviewer permission to reject an attractive image.
Prohibited implications
I list anything the image must not quietly claim:
- an unverified medical, safety, environmental, or performance benefit;
- a certification that is not documented;
- a before-and-after result that was never tested;
- a celebrity or real person without rights;
- competitor branding;
- unsafe product use.
Without this sheet, “consistent” often means only “the outputs look similar to each other.” That is not the same as being faithful to the product.
3. I decide what must be preserved and what may be generated
The most useful question is not always “Which image model should I use?”
It is:
Which pixels must remain evidence, and which pixels are allowed to become creative material?
If the label, shape, or finish is important, I prefer a workflow that preserves the original product layer and changes the background, surface, shadow, or composition around it.
Regenerating the whole frame gives the model more freedom. That may be fine for a concept image. It is risky when the image will help a buyer decide what arrives in the box.
I use three rough modes:
- Preserve — keep the product pixels and edit the surrounding scene.
- Constrain — allow limited product edits but require a strict detail review.
- Concept — allow invention, label the result clearly, and keep it away from listing evidence.
The mode belongs in the brief. It should not be guessed after the output looks good.
4. I check rights before I describe the scene
Every input has a rights question:
- Who owns the product photo?
- Can the brand mark be used in this channel?
- Is the model authorized for this use?
- Is the location or reference image licensed?
- Does the intended output use fit the relevant terms?
If I create an original adult model for apparel, I do not ask the system to imitate a real person. “Make her look like this celebrity” is not a harmless style instruction.
A beautiful output does not repair a missing permission record.
5. I separate preservation rules from scene direction
I write the prompt in two blocks.
The first block protects product truth:
PRESERVE
- exact silhouette and proportions
- current navy color and matte finish
- logo position and readable label layout
- included cap and visible connector
DO NOT INVENT
- back label
- extra accessories
- certification marks
- liquid, smoke, sparks, or performance effects
The second block describes the scene:
SCENE
- clean editorial product photography
- product centered on a light stone shelf
- one folded neutral towel behind the product
- soft morning bathroom background
- realistic contact shadow
- no people, hands, text overlay, or additional products
Google’s current Product Studio guidance for its own background workflow asks merchants to describe the product, placement, surroundings, and background. I use those fields because they force the scene to become specific. I do not treat them as a universal contract for every image tool.
Vague praise—“premium,” “viral,” “luxury,” “amazing”—does not tell a reviewer what should pass.
6. I create one or two calibration images
My default is a small calibration, not a ten-image batch.
For each calibration, I record:
- source-image revision;
- prompt revision;
- intended channel;
- requested output count;
- task or request count;
- approved cost ceiling;
- returned output;
- pass, repair, or reject decision.
Then I compare the result with the product truth sheet.
The purpose is not to find one lucky image. It is to decide whether the direction is controlled enough to continue.
7. I run product QA and channel QA separately
I review the product first:
| Check | What I look for |
|---|---|
| Geometry | shape, proportions, openings, handles, closures |
| Color | variant, tone, finish, unexpected gradients |
| Text | logo, label, spelling, layout, invented marks |
| Scale | product size relative to hands, furniture, or props |
| Contact | believable shadow, grip, surface contact, garment fit |
| Completeness | included parts present, no invented accessories |
Then I review the destination:
| Check | What I look for |
|---|---|
| Claim safety | scene does not imply an unverified result |
| Crop | important product details survive the placement |
| Platform fit | current format and content rules are checked |
| Rights | intended use matches the recorded permissions |
| Traceability | source, prompt, output, reviewer, and decision are stored |
An API success and a publication approval are different events. I keep them separate.
8. I repair the failed layer instead of rerolling everything
If the product is right but the shadow floats, I repair the shadow.
If the scene works but the label changes, I return to the preserved product layer.
If the model invents a hidden detail, I add another source view or change the camera angle so the output does not need that detail.
Blind regeneration hides the cause of failure. It also makes it harder to tell whether a workflow is improving or just producing more attempts.
The apparel branch I verified in XPLA
I work on XPLA, so I checked its current apparel workflow before mentioning it here.
The live Fashion Street Shoot Skill is intentionally narrower than this general tutorial. It accepts one garment product photo and either an authorized adult model or an original adult model. The workflow starts with two calibration images. After approval, it expands to a 5–10 image, 3:4 street-photo set and creates a QA review page.
That is an apparel workflow. I would not use its existence to claim that XPLA supports every product category or that every output will preserve every detail. Current account access, model usage, price, and output quality still need to be checked at the time of use.
The useful idea is the gate: lock the person, garment, and visual direction with two images before paying to expand the set.
A compact checklist
Before generation:
- [ ] one destination and one image job;
- [ ] rights-cleared source;
- [ ] visible and verified product facts;
- [ ] unknown and prohibited details;
- [ ] preserve, constrain, or concept mode;
- [ ] one prompt revision;
- [ ] task and cost ceiling.
Before expansion:
- [ ] geometry passed;
- [ ] color passed;
- [ ] logo and text passed or have an approved repair;
- [ ] scale and contact are plausible;
- [ ] scene makes no unsupported claim;
- [ ] channel crop and rules checked;
- [ ] source, output, and decision archived.
Limitations
- One source angle cannot define hidden geometry.
- Generative tools can alter logos, labels, patterns, seams, reflections, hands, packaging, and scale.
- A realistic scene can imply a false claim without adding any text.
- Marketplace and advertising rules change.
- A method that works for a bottle may fail for apparel, jewelry, reflective metal, furniture, or food.
- One approved image does not prove a tool is universally “best.”
- More images do not prove more clicks, orders, or revenue.
FAQ
What is a product truth sheet?
It is a short review document that separates visible facts, verified facts, unknown details, and prohibited implications. It becomes the acceptance standard for prompts and outputs.
Should I replace the background or regenerate the whole product?
If exact product details matter, preserve the product layer when possible and change the surrounding scene. Full-frame regeneration belongs in a stricter QA workflow or a clearly labeled concept stage.
How many calibration images should I create?
I usually start with one or two. The goal is to approve a direction before creating a larger set, not to search through a large batch for one lucky result.
How do I keep a logo or label accurate?
Use the clearest approved source, preserve original pixels when possible, add detail views, and compare every output with the truth sheet. Reject or repair altered text rather than explaining it away.
Can one source photo support a full lifestyle set?
Sometimes, but only for what the source actually shows. If a new angle reveals hidden geometry, provide another view or avoid that angle.
Does a technically successful output count as approved?
No. Technical completion only proves the request finished. Product truth, rights, claims, crop, and channel fit still need review.
Sources and disclosure
- Google Merchant Center: About Product Studio, rechecked September 1, 2026. I use its scene fields as a bounded example, not a universal tool contract.
- XPLA: Fashion Street Shoot Skill, rechecked September 1, 2026.
Disclosure: I prepared this article for XPLA. I verified the live page and current repository wording, but I did not invent a paid run, customer result, sales lift, traffic result, or tool ranking.
If your product is apparel, the live XPLA page shows the narrower one-garment, two-calibration, approve-then-expand workflow:
Top comments (0)