Eight Rounds, Two Models, One Surprise: What GPT Image 2.5 Sunburst Actually Changed

OpenAI shipped GPT Image 2.5 in September 2026 in two variants — Flare, the fast one, and Sunburst, positioned as the quality tier. The launch language promised sharper detail, stronger style adherence, and more control over edits.I wanted to know whether that held up in actual production work, not in a demo reel.So I built a controlled test and ran it across eight rounds inside ImagineArt Workflows, generating both chains in parallel on the same canvas.
The Setup
Every variable locked except the model:
- Same prompts, word for word
- Same quality setting (High)
- Same resolution (2K)
- Same reference images fed into the same node positions

What I Tested
Rounds 1–4: Multi-turn edit drift. A fashion editorial chain — one model, four sequential garment swaps, each turn also changing pose, camera angle, or adding accessories. Each edit built on the previous output, never on the original. This is the failure mode everyone complains about: by edit four or five, the subject has quietly become a different person.
Round 1

Round 2


Round 3


Round 4


Round 5: Legible multi-character text. A three-line slogan on a sweatshirt plus a small alphanumeric label on a shopping bag.

Round 6: Real-photo subject fidelity. A genuine photograph — not AI-generated — transformed into a completely different setting, lighting, and framing.

Round 7: Complex layout. A six-element flat lay with specified positions, counts, and orientations.


Round 8: Raw photorealism under hard light. A single portrait, chiaroscuro lighting, explicit instruction for visible pores and no retouching.

What I Found
The drift problem didn't appear. In either model.
Four compounded edits deep — through silk, cream wool, leather, and pleated satin, through a 90-degree angle change and a full back-turn — both chains held the subject's face, bone structure, hair, and skin tone against the original anchor image. GPT Image 2 held identity just as well as Sunburst did.
That's worth sitting with. The headline improvement in 2.5 is supposed to be editing chains that don't degrade. But if GPT Image 2 doesn't degrade either, the fix solves a problem that wasn't there at this chain length.
On image quality, the wins split evenly.
Sunburst took the round on leather grain and anatomical proportion. It took the real-photo fidelity round on accessory accuracy and how it handled light falling across skin. It took the photorealism round on atmospheric detail and environmental specificity.
GPT Image 2 took the round on pleat structure and hardware detail. It took the text round — Sunburst produced a small artifact on a secondary label and ignored a framing instruction. It took the layout round decisively, because Sunburst added an entire element nobody asked for.
That last one matters for commercial work. Sunburst made the prettier flat lay. GPT Image 2 made the one I specified. For client deliverables built to a brief, obedience beats invention.
Three rounds each. Parity in the middle. No systematic quality winner.
The Number That Actually Decided It
One measurement was large, consistent, and repeatable across every single round:
GPT Image 2: 135 credits per generation. GPT Image 2.5 Sunburst: 37.5 credits per generation.
Sunburst also rendered noticeably faster.
That's a 3.6× cost difference for output I could not reliably separate on quality.
A Methodology Note Worth Stealing
Partway through, I noticed the two failure types I was catching needed different viewing distances.
Get close to judge material — fabric grain, specular behavior on silk, whether leather reads as grained or patent.
Step back to judge anatomy — limb proportion, weight distribution, whether a heel is plausibly shaped.
Material errors vanish at distance. Anatomy errors vanish up close. If you evaluate AI output at one fixed distance, you're systematically blind to half of what's wrong with it.
The Honest Limitations
n=1 per round. These were single generations, not averaged across seeds. Round-level "wins" could partly reflect generation variance rather than model capability. Two data points pointing the same direction is suggestive, not conclusive.
One genre. Fashion and portrait photography on controlled backdrops. Text-heavy design layouts, illustration styles, multi-subject scenes, or non-photoreal work could separate these models very differently.
One platform. Run through ImagineArt Workflows, not direct API.See the whole images and prompts here:
https://www.imagine.art/enterprise/flow/7c6d3b6b-d4ca-4e70-95c3-198a82b54131
Anyone telling you they've definitively ranked two image models off a handful of generations is overselling. I'm not doing that.
What I'm Actually Doing with This
Routing to Sunburst by default for photorealistic fashion and product work. Same quality, a quarter of the credits, faster iteration.Keeping GPT Image 2 in reserve for spec-driven layout work where I need the model to follow instructions exactly rather than improve on them.
The Bigger Read
Here's what I think this test actually shows.
The generational gain from GPT Image 2 to 2.5 is efficiency, not capability. OpenAI compressed roughly the same output quality into a dramatically cheaper, faster model. That's a genuine engineering achievement, and for anyone running high-volume production it changes the economics substantially.
But it isn't a creative capability jump. The marketing language around sharper detail and stronger style adherence oversells what a controlled test finds.
Which is fine — as long as you know that going in, and don't rebuild your pipeline expecting a quality leap that isn't there.
Your Turn
Which image model are you routing to by default right now — and what made you choose it? Not what the benchmarks say. What you've actually seen in your own work.
Drop it in the comments. I'd rather learn from people who've put the hours in than from another launch post.