Analysis
Analysis

GPT-Image-2.5 Reality Check: 50% Lower Latency, but API Token Rates Are 2× GPT-Image-2; No Independent 2.5 Leaderboard Yet

Published Sep 11, 2026 Sources checked Sep 9, 2026

OpenAI’s new GPT-Image-2.5 Flare and Sunburst promise faster, more precise image generation and editing. We audit the 50% latency claim, doubled API token rates, safety-eval caveats and missing independent 2.5 leaderboard evidence.

GPT-Image-2.5 Reality Check

OpenAI released ChatGPT Images 2.5 on September 8, 2026, alongside two API models: gpt-image-2.5-flare and gpt-image-2.5-sunburst. Flare is positioned as the default choice for most applications and high-volume generation; Sunburst is the slower, precision-oriented option for detailed creative and editing workflows.

The launch is material because it combines a new image model family with product-level editing tools in ChatGPT: Sketch lets users draw a rough visual guide, templates provide starting structures for formats such as posters and merchandise, and comments can target specific parts of an image for editing. OpenAI says the 2.5 family improves subject fidelity, lighting, textures, precise edits and multi-turn consistency.

Primary launch source: OpenAI — Introducing ChatGPT Images 2.5, September 8, 2026.

The most important reality check is that the launch combines strong vendor claims with very limited independent 2.5 benchmark evidence on day one. OpenAI says Flare delivers higher-quality images than GPT-Image-2 at 50% lower latency, and launch partner Manus says it observed roughly 2–4× faster generation in its own evaluations. Those claims are useful, but they are not yet a neutral, reproducible cross-provider benchmark with a published prompt set, hardware, quality tier, sample size and distribution of generation times.

Flare and Sunburst are two different deployment choices

OpenAI describes the pair by workload rather than by a single “best model” ranking:

  • GPT-Image-2.5 Flare: the default for most API applications, with an emphasis on speed, quality and high-volume generation.
  • GPT-Image-2.5 Sunburst: the premium precision option, intended for campaign creative, polished product imagery and workflows where tighter control across edits matters more than turnaround time.

The consumer-facing ChatGPT Images 2.5 experience is rolling out across ChatGPT, ChatGPT Work and Codex on desktop, mobile and web. The API model IDs are separate and explicit, so production teams should log the exact model rather than writing only “Images 2.5” in evaluation reports.

OpenAI’s launch page says generation latency is reduced by up to 50% compared with Images 2.0, and specifically says Flare provides higher-quality images than GPT-Image-2 at 50% lower latency. The phrase “up to” matters: it is not a p50 or p95 service-level guarantee, and the public launch material reviewed here does not disclose the prompt mix, resolution/quality settings, exact sample count or tail-latency distribution behind that number.

Manus is a launch partner and says its evaluation saw Flare run at about two to four times the speed of GPT-Image-2 while retaining high quality. That is useful corroboration from a customer, but it is still a partner evaluation quoted by OpenAI rather than an independently published reproducible benchmark.

API pricing: the 2.5 token rates are double GPT-Image-2

The current OpenAI API pricing page lists both 2.5 models at:

Model Image input Cached image input Image output Text input Cached text input
gpt-image-2.5-flare $8/M tokens $2/M $30/M $5/M $1.25/M
gpt-image-2.5-sunburst $8/M tokens $2/M $30/M $5/M $1.25/M
gpt-image-2 $4/M tokens $1/M $15/M $2.50/M $0.625/M

Source: OpenAI API pricing.

On a per-token basis, every listed 2.5 rate in that table is therefore 2× GPT-Image-2. That does not prove that every finished 2.5 image costs twice as much. Per-image cost depends on the actual image/text tokens consumed by the requested size, quality and edit workflow. The launch materials reviewed here do not publish a matched per-image token-consumption study for Flare, Sunburst and GPT-Image-2.

This creates a practical tradeoff that the “50% lower latency” headline alone does not answer: Flare can be much faster while still carrying higher token rates. Teams should measure cost per accepted image, not just price per token or seconds per request. If 2.5 needs fewer retries because editing is more reliable, the effective workflow cost could be better even at higher token rates; if token consumption and retry rates are similar, the API bill can rise.

Independent image-quality evidence: 2.5 is not on the current Artificial Analysis boards yet

As of this verification, GPT-Image-2.5 Flare and Sunburst do not appear on the current Artificial Analysis text-to-image or image-editing leaderboards.

The incumbent GPT Image 2 has strong independent evidence:

  • Artificial Analysis text-to-image: GPT Image 2 (high) is #1 at Elo 1178, with a ±10 95% confidence interval across 14,585 samples.
  • Artificial Analysis image editing: GPT Image 2 (high) is #2 at Elo 1117, with a ±10 interval across 11,046 samples, behind MAI-Image-2.6 at 1122.

Current independent boards: Artificial Analysis text-to-image and image editing.

That does not mean GPT-Image-2 is better than 2.5. It means GPT-Image-2 currently has the more mature independent measurement record. Claims that Flare or Sunburst are already “#1 on independent leaderboards” should be treated cautiously unless the exact 2.5 model ID is visible in the board and the result can be reproduced from the board’s current data.

The correct launch-day evidence status is:

Claim Evidence status
Flare 50% lower latency than GPT-Image-2 OpenAI vendor claim; no public matched independent latency distribution accepted yet
Flare 2–4× faster than GPT-Image-2 Manus partner evaluation quoted by OpenAI
Better subject fidelity / editing consistency OpenAI launch claim with demonstrations; no neutral 2.5 leaderboard score accepted yet
GPT Image 2 text-to-image Elo 1178 Independent Artificial Analysis current leaderboard
GPT Image 2 editing Elo 1117 Independent Artificial Analysis current leaderboard
Flare/Sunburst independent Elo Not yet present on the current boards reviewed

Safety: lower overall unsafe-presented rates, but no unsafe-shown difference is statistically significant

OpenAI’s ChatGPT Images 2.5 System Card provides more quantitative evidence than the launch marketing page. The company ran an automated adversarial evaluation designed to elicit policy-violating images and measured whether the final output was safe, blocked, or unsafe and still presented.

OpenAI reports:

  • Sunburst: 77.0% safe generated, 21.9% unsafe blocked, 1.09% unsafe presented
  • Flare: 79.4% safe generated, 19.2% unsafe blocked, 1.41% unsafe presented
  • ChatGPT Images 2.0 baseline: 75.2% safe generated, 23.1% unsafe blocked, 1.64% unsafe presented

Source: OpenAI — ChatGPT Images 2.5 System Card.

Those headline rates are directionally encouraging, but the system card adds an important statistical limitation: no unsafe-presented difference in the policy tables meets the stated p < 0.05 threshold. The significance tests are two-sided exact McNemar tests and are unadjusted for multiple comparisons. OpenAI also says automated labels can contain errors, policy-specific sample sizes vary, and the findings apply to the fixed adversarial test set and the evaluated model/safeguard configurations.

So the defensible conclusion is not “2.5 is proven safer.” It is that OpenAI’s internal adversarial evaluation reports lower aggregate unsafe-presented rates for both 2.5 variants than the 2.0 baseline, while the system card itself warns that the unsafe-shown differences are not statistically significant under the reported test and that the result may not generalize to production traffic.

The safety stack includes upstream policy checks, multimodal input/output monitoring and final-output blocking. OpenAI also continues to attach C2PA provenance metadata and says it is adding Google DeepMind SynthID invisible watermarking across ChatGPT, Codex and the OpenAI API.

SWE-bench, Terminal-Bench and coding benchmarks are not applicable here

GPT-Image-2.5 Flare and Sunburst are image-generation/editing models, not software-engineering agents. SWE-bench Verified, SWE-bench Pro, Terminal-Bench, DeepSWE and agentic coding scores are therefore not applicable to this release.

A common benchmark mistake is to attach the scores of GPT-6 Astra or another text/coding model to the 2.5 image models because they are from the same vendor. This article does not do that. For image models, relevant evidence includes blind image preference, editing preference, prompt adherence, reference fidelity, latency, output cost and safety—not repository issue resolution.

Early hands-on and public feedback: useful, but still anecdotal

Axios tested the new image engine before launch on tasks including logo creation, photo transformation and tattoo design and reported that editing and likeness preservation were noticeably improved. That is a named publication’s hands-on experience, but it is still a small editorial test rather than a controlled benchmark with a published prompt suite and statistical analysis.

Source: Axios, September 8, 2026.

Launch-day Reddit discussion is similarly anecdotal. In one r/codex thread, users were excited about the release but quickly debated output resolution and whether asking ChatGPT to “upscale to 4K” actually changed the underlying image dimensions; at least one user reported that it only appeared to sharpen the image without changing the resolution. That is a useful warning against assuming a chat instruction changes native output resolution, but it is self-selected user feedback, not a controlled test of Flare or Sunburst.

Source: r/codex, September 8, 2026.

I did not find a stable, independently reproducible X post with a pinned 2.5 API model ID, prompt set, quality tier and timing traces that was strong enough to treat as benchmark evidence in this bounded review. OpenAI and partner promotion should be cited as launch provenance, not relabeled as independent community consensus.

Practical verdict

ChatGPT Images 2.5 is a real, same-day product and API release with two distinct production models. Flare is the speed/default path; Sunburst is the precision/editing path. The strongest practical claim is the latency improvement, but it remains a vendor/partner measurement rather than an independent matched benchmark.

The pricing story is more nuanced than “faster and cheaper.” OpenAI’s current pricing page makes the 2.5 image/text token rates twice GPT-Image-2’s token rates. Whether a completed 2.5 workflow is actually more expensive depends on token consumption, retries and edit success, none of which should be invented from the rate card.

Independent quality evidence is also not yet mature. GPT Image 2 still occupies the current Artificial Analysis leaderboards with large sample counts, while the new 2.5 model IDs have not appeared there at this verification time. For buyers, the most valuable next evidence is a matched independent Flare/Sunburst evaluation with the exact API model IDs, quality/resolution settings, identical prompts, p50/p95 latency, token consumption, per-image cost, edit success rate and blind human preference.

Until that arrives, the best production approach is to test both 2.5 siblings on the exact workload you ship, log model IDs and settings, and compare cost per accepted output rather than extrapolating from launch claims alone.

Sources

This article is built from the source material below. Open the originals for full context and the latest updates.

More ways to save

Discover deals, coupons and free courses on our sister site.

Explore DealVorio
Save more with DealVorio: deals, coupons, free courses, apps and books