I’d start with Sunburst for edits that must preserve a product or person, Flare for rapid iteration, and Nano Banana 2 for generation that needs search context, unusual canvas shapes, or predictable batch costs.
That split is more useful than declaring an overall winner. The reported Arena results favor GPT-Image-2.5, especially for editing, but a preference leaderboard doesn’t measure retrieval integration, output constraints, or the cost of producing thousands of acceptable assets.
This comparison uses the September 2026 release and leaderboard snapshot discussed below. The OpenAI benchmark entries are still preliminary; I’d treat them as a reason to prioritize an evaluation, rather than a substitute for one.
Start with the API contract
GPT-Image-2.5 has two API variants. OpenAI launched both on September 8, 2026:
- Flare targets everyday generation, fast iteration, and volume.
- Sunburst spends more generation time on precision editing and tighter creative control.
Nano Banana 2 is Google’s Gemini 3.1 Flash Image, introduced on February 26, 2026. Google’s release notes place general availability on May 28, 2026, with video-to-image context support added. The older preview model ID was deprecated, although some integration URLs still contain the preview slug.
Here’s the contract I’d compare before looking at sample images:
| Capability | GPT-Image-2.5 Flare | GPT-Image-2.5 Sunburst | Nano Banana 2 |
|---|---|---|---|
| API model ID | gpt-image-2.5-flare |
gpt-image-2.5-sunburst |
gemini-3.1-flash-image |
| Input | Text, image | Text, image | Text, image, video context |
| Output | Image | Image | Image, text |
| Main role | Fast generation and controlled edits | Precision editing and premium assets | Grounded multimodal generation |
| Quality settings |
low, medium, high, xhigh, max, auto
|
Same | Primarily resolution-driven |
| Common output sizes | 1024×1024, 1536×1024, 1024×1536 | Same | 0.5K, 1K, 2K, 4K |
| Maximum edge | 3840 px | 3840 px | Up to 4K |
| Aspect ratios | Approximately 1:3 through 3:1 | Same | Includes 1:4, 4:1, 1:8, 8:1 |
| Transparent backgrounds | PNG and WebP | PNG and WebP | Not a headline API capability |
| Native search grounding | Not documented at image-model level | Not documented at image-model level | Google Web and Image Search |
| Thinking | Not exposed as a core capability | Not exposed as a core capability | Supported |
| Batch API | Not a headline 2.5 feature | Not a headline 2.5 feature | Supported |
OpenAI’s model documentation describes Flare as the faster everyday option and Sunburst as the variant optimized for editing precision. Both expose the six quality settings above.
For Google, I’d check the current Gemini 3.1 Flash Image documentation, particularly when migrating an integration that still references the preview model.
Editing is where I’d test Sunburst first
The useful editing question is how much unrelated content changes after each revision.
A product workflow might replace a background, adjust lighting, add seasonal styling, revise copy, and then change the aspect ratio. Each individual render can look good while the bottle, label, or person slowly drifts away from the reference.
OpenAI’s Images 2.5 announcement emphasizes preservation of reference subjects, natural lighting, richer textures, and more consistent multi-turn edits. Sunburst specifically targets the workflows where preserving those details matters more than minimum latency.
That makes it my first evaluation candidate for:
- Product photographs where packaging and geometry must survive revisions.
- Campaign assets with repeated changes to copy, placement, and styling.
- Character or subject transformations driven by reference images.
- Structured posters, branded layouts, and transparent design elements.
Flare occupies the faster lane in the same family. OpenAI claims higher-quality output with up to 50% lower generation latency than GPT Image 2.
That figure is an OpenAI generation-to-generation comparison. It says nothing conclusive about Flare’s speed relative to Nano Banana 2. I’d benchmark both with the same prompts, output dimensions, and concurrency before making a latency commitment.
Google’s model also supports conversational editing and subject consistency. Its stronger differentiator is the surrounding Gemini workflow: reasoning, search, multilingual text, and multiple references.
What the Arena numbers actually support
Arena aggregates blind side-by-side human preferences. The reported snapshot looks like this:
| Arena metric | Sunburst | Flare | Nano Banana 2 |
|---|---|---|---|
| Text-to-Image score | 1421±13 | 1399±13 | 1261±5 |
| Text-to-Image rank | #1 | #2 | #9 |
| Text-to-Image votes | 3,149 | 2,856 | 41,957 |
| Single Image Edit score | 1520±9 | 1491±9 | 1387±4 |
| Single Image Edit rank | #1 | #2 | #12 |
| Single Image Edit votes | 6,704 | 5,676 | 157,693 |
Sunburst leads Nano Banana 2 by 160 points for text-to-image generation and 133 points for single-image editing. Flare’s corresponding leads are 138 and 104 points.
That is a strong early preference signal, and the editing result is consistent with OpenAI’s stated priorities. It makes Sunburst a sensible first test for controlled revisions.
The sample sizes matter, though. Both OpenAI entries are marked Preliminary, with only a few thousand votes. Google’s entry has 41,957 generation votes and 157,693 editing votes. The new entries have more room to move as comparisons accumulate.
Arena points also aren’t percentages. A 160-point lead doesn’t translate into a fixed percentage improvement in image quality, and the ranking reflects Arena’s prompt distribution. It doesn’t establish which model handles your packaging, typography, references, or acceptance criteria best.
I’d use the leaderboard to choose evaluation order, then measure performance on the application’s actual workload.
Where Nano Banana 2 changes the implementation
Search can be part of generation
Google documents both Web Search and Image Search grounding for Gemini 3.1 Flash Image. Retrieved text and images can inform generation using current web information.
This applies when the Google Search tool is enabled and its attribution requirements are followed. Merely selecting the model doesn’t make every output grounded.
For a travel application, educational diagram, visual search interface, or information graphic, that integration can reduce the work needed to assemble context. A landmark illustration can use retrieved references; a localized graphic can combine factual context with rendered text.
GPT-Image-2.5 doesn’t document equivalent native grounding at the image-model level. An application can supply its own retrieved context and reference images, which narrows the practical difference. If the application already uses Gemini, keeping retrieval and generation in that workflow may be simpler.
Reference scale is explicitly documented
Google highlights identity preservation for up to five characters and fidelity for up to 14 objects in a workflow.
Those figures are useful when scoping compositions with several recurring subjects. They are documented capability claims, rather than a guarantee that every arrangement will preserve every detail.
OpenAI emphasizes strong reference preservation, especially with Sunburst, but the comparison doesn’t provide equivalent character and object counts for its models. I’d separate documented reference scale from observed fidelity during testing.
Localization has a clear place in the feature set
Nano Banana 2 emphasizes international text rendering and translation of text inside an existing visual. That is useful for adapting one campaign across markets.
OpenAI emphasizes infographic accuracy, hierarchy, and layout. For a static advertisement or UI concept where exact spatial structure dominates, I’d try GPT-Image-2.5 first. For multilingual graphics or visuals whose copy depends on retrieved information, I’d start with Google.
Both still need output verification when numerical or legal copy matters. A convincing layout can contain incorrect text.
Canvas requirements can decide the model before quality does
Nano Banana 2 supports 0.5K, 1K, 2K, and 4K output, spanning the documented 512px-to-4K range. Its extreme aspect ratios include 1:4, 4:1, 1:8, and 8:1.
That gives it a straightforward advantage for panoramic backgrounds, narrow mobile creatives, tall product displays, and campaigns needing many canvas shapes.
GPT-Image-2.5 supports custom dimensions within an approximately 1:3 to 3:1 aspect-ratio range. Its documented limits are:
- Maximum edge: 3840 pixels.
- Maximum output area: 8,294,400 pixels.
Those constraints should stay explicit in application validation. “Supports 4K” would obscure the difference between OpenAI’s limits and Nano Banana 2’s 4096×4096 4K tier.
OpenAI has a separate practical advantage: transparent PNG and WebP backgrounds are explicitly supported. That matters for product cutouts, composited design elements, and assets that need an alpha channel.
My routing decision here would be mechanical: check canvas dimensions and transparency requirements before spending time comparing creative quality.
Photorealism needs its own evaluation
I wouldn’t infer a universal photorealism winner from the Arena ranking.
OpenAI describes improvements to natural lighting, texture, and recognizable reference subjects. Google makes similar claims about detail, texture, lighting, and photographic quality.
Some early subjective comparisons have favored Nano Banana 2 for camera feel, materials, and natural product scenes, while favoring GPT-Image-2.5 for controlled edits and text-heavy layouts. Those observations are useful leads, but they aren’t standardized cross-provider measurements.
For a production evaluation, I’d include skin and hair, glossy and matte products, transparent materials, fabric, food, architecture, indoor lighting, shallow depth of field, and identity preservation from reference photos.
The metric I care about is cost per accepted image. A model that produces one exceptional sample can still be expensive if most outputs need another render or manual correction.
One shared product prompt
The supplied comparison uses this exact prompt:
Photorealistic product photograph of a matte black water bottle standing on a pale concrete ledge. Soft morning light from the left, gentle reflections, shallow depth of field. Centered composition, Include ONLY this text (verbatim): headline "YOURS TO CREATE" in bold sans-serif across the top, subhead "Limited Edition" smaller at the bottom. No other text or logos.
GPT-Image-2.5 Sunburst
Nano Banana 2
I’d score this prompt separately for exact copy, text placement, bottle geometry, material appearance, lighting direction, and unwanted logos. It exercises several requirements at once, so a single overall preference can hide the reason an image fails acceptance.
Compare billing at the image level
The pricing models encourage different budgeting approaches.
GPT-Image-2.5 uses token billing, with consumption affected by quality settings. Nano Banana 2 publishes image-output costs by resolution, making that part of the budget easier to estimate.
| Official billing unit | Price |
|---|---|
| GPT-Image-2.5 text input | $5 / 1M tokens |
| GPT-Image-2.5 image input | $8 / 1M tokens |
| GPT-Image-2.5 image output | $30 / 1M image tokens |
| Nano Banana 2, 0.5K | $0.045 / image |
| Nano Banana 2, 1K | $0.067 / image |
| Nano Banana 2, 2K | $0.101 / image |
| Nano Banana 2, 4K | $0.151 / image |
Comparing OpenAI’s $30 with Google’s $60 per million image tokens doesn’t establish that OpenAI costs half as much. The providers tokenize images differently, and OpenAI’s quality ladder changes consumption.
Google’s published pricing also includes Batch API image-output equivalents approximately 50% below standard pricing:
| Resolution | Standard | Approximate Batch equivalent |
|---|---|---|
| 0.5K | $0.045 | $0.022 |
| 1K | $0.067 | $0.034 |
| 2K | $0.101 | $0.050 |
| 4K | $0.151 | $0.076 |
For asynchronous asset production, that is a meaningful feature. For interactive generation, I’d evaluate the standard path separately.
OpenAI’s quality controls support another useful pattern: inexpensive drafts followed by more expensive final renders. Both variants expose low, medium, high, xhigh, max, and auto; the selected setting belongs in both cost and latency measurements.
A unified API can simplify routing
If I wanted one integration for both families, CometAPI is one option. The source’s quoted rates for both Flare and Sunburst are $4 per million text-input tokens and $24 per million output tokens, versus the corresponding official $5/$30 rates; image-input pricing requires a current provider quote.
The same quoted gateway pricing for Nano Banana 2 is $0.0360, $0.0536, $0.0808, and $0.1208 per image at 0.5K, 1K, 2K, and 4K respectively. Batch availability needs checking with the provider.
The engineering value is a common integration through which the application can select a model per request. I’d still keep provider-specific constraints visible in the routing logic.
The routing policy I’d put into an application
I’d make model selection follow the request’s constraints and the cost of rework.
| Request characteristic | First model I’d evaluate |
|---|---|
| Precise edits with minimal unrelated change | Sunburst |
| Final campaign or product asset | Sunburst |
| Rapid OpenAI variants and everyday generation | Flare |
| Transparent PNG or WebP asset | GPT-Image-2.5 |
| Current web or visual context needed | Nano Banana 2 |
| Extreme vertical or horizontal canvas | Nano Banana 2 |
| Native resolution tiers through 4K | Nano Banana 2 |
| Multilingual text or translation inside a visual | Nano Banana 2 |
| Asynchronous volume with published Batch pricing | Nano Banana 2 |
| Several recurring characters or objects | Nano Banana 2 for its documented reference scale |
| Photorealistic scenes without other constraints | Evaluate both on representative prompts |
For an editing-heavy product, I’d begin with Sunburst and test whether Flare meets the same acceptance criteria at lower latency. For a Gemini application generating grounded diagrams or localized assets in many shapes, I’d begin with Nano Banana 2.
Mixed workloads justify using all three. One possible flow is grounded concept generation with Google followed by a controlled Sunburst revision. Another is Flare for initial variants, Sunburst for difficult final edits, and Google for canvas shapes outside OpenAI’s range.
Those are architectures I’d evaluate, not measured guarantees that switching models improves a result. Any handoff belongs in the same reference-preservation tests as a single-model editing sequence.
Before committing, I’d run identical prompt sets at comparable dimensions and concurrency, record actual latency and billed cost, and judge the final assets against explicit acceptance criteria. The preliminary leaderboard makes GPT-Image-2.5 a strong starting point for quality and editing; the production requirements determine whether it earns the request.
Originally published at cometapi.com


Top comments (0)