I used to treat AI image generation like a text-to-image lottery. Describe a product, roll the dice, and hope the label text came out readable. After a few hundred wasted generations, I switched to a different workflow: reference-led editing.
What that means in practice
Instead of describing the product from scratch, I upload the actual product photo as a reference. Then I only describe what I want to change — the background, the lighting, the scene. The model keeps the product shape, label, and key details intact.
On a recent skincare campaign brief, this made the difference between a mood board and a usable asset:
- From scratch (Flare, 1K): Four rolls, all with garbled label text. Good for direction, not for delivery.
- Reference edit (Sunburst, 2K): Uploaded the real bottle, asked for a warm kitchen scene. Label text stayed crisp across three variations in one pass.
The two modes are not the same tool
GPT Image 2.5 ships with two modes, and I stopped treating them as interchangeable:
- Flare — fast, cheap (10 credits at 1K), designed for iterating on a visual direction. I use this for the first few rounds.
- Sunburst — slower, 20 credits at 2K, built for detail-focused generation and reference edits. Label text and fine typography stay legible here.
The credit math that works for me
I do all my exploration at 1K in Flare. Once the composition and scene are locked, I render the final at 2K in Sunburst. I haven't needed 4K for social or ad placements yet.
If you're doing product marketing work and still generating every shot from a blank prompt, the reference-edit route saves more time than any prompt tweak. You can test the workflow on the GPT Image 2.5 page by uploading your own product image and asking for a scene change.
Top comments (0)