ChatGPT Images 2.0 (gpt-image-2)

OpenAI's April 2026 image model — production-grade visuals with reliable instruction-following, dense multilingual text, aspect ratios from 3:1 to 1:3, and a first-of-its-kind thinking mode that researches, self-checks, and renders functional QR codes. Sam Altman framed it as "going from GPT-3 to GPT-5 all at once." Source: Gizmodo, https://gizmodo.com/openai-unveils-new-image-generator-to-usher-in-an-ai-slop-renaissance-2000749159, 2026-04-21

Launch tweet: 25,198 likes, 6,985 bookmarks. Source: X/@OpenAI, 2026-04-21

What it is

ChatGPT Images 2.0 is the consumer surface; gpt-image-2 is the underlying model, available the same day in the API and Codex and to all ChatGPT tiers. OpenAI positions it for production workflows — images that must be "accurate, readable, on-brand, localized, formatted for the destination surface, and usable without heavy cleanup" — rather than pure experimentation. It is OpenAI's first image model with thinking (reasoning) capabilities and ships in two modes: instant (default) and thinking (toggle). It deprecates GPT-Image-1.5 as the default while keeping it in the API for legacy support. Source: OpenAI Developer Community, https://community.openai.com/t/introducing-gpt-image-2-available-today-in-the-api-and-codex/1379479, 2026-04-21 Source: VentureBeat, https://venturebeat.com/technology/openais-chatgpt-images-2-0-is-here-and-it-does-multilingual-text-full-infographics-slides-maps-even-manga-seemingly-flawlessly, 2026-04-21

Capabilities

From the launch thread: Source: X/@OpenAI thread, 2026-04-21

  • Precision and instruction-following — preserves requested details and renders the elements that historically broke image models: small text, iconography, UI elements, dense compositions, subtle stylistic constraints — up to 2K in ChatGPT.
  • Dense, structured text — diagrams, infographics, charts, posters, comics, slides, floor plans, image grids, multi-angle character sheets.
  • Stronger across languages — non-English text rendered correctly and coherently, not just transliterated.
  • Photo realism + style range — cinematic stills, pixel art, manga, with consistency in texture, lighting, and composition.
  • Flexible aspect ratios — as wide as 3:1 and as tall as 1:3; API supports resolutions up to 4K (beta).
  • Real-world intelligence — December 2025 knowledge cutoff; can run copywriting → analysis → composition end-to-end.

Artifact and Video Review

The downloaded launch image is a title card ("ChatGPT Images 2.0") rather than a capability proof, so the video frames carry the useful evidence. The launch video samples a chameleon across several generated scenes: a house/topiary image, clock text, cereal package text, a pool float, a newspaper headline, and a ping-pong scene. That artifact supports the page's existing claim that the model targets dense text, object consistency, and multi-scene instruction following. Source: X bookmark media artifact 2046670977145372771, reviewed 2026-06-30

No executable Kevin skill was created from this row. The reusable pattern belongs in design/media-generation routing: use image models when the output needs legible text, controlled aspect ratio, and self-checked visual deliverables, then verify the generated asset like any other artifact. Source: compiled from X launch thread and media review, 2026-06-30

Thinking mode

With a reasoning model selected, Images 2.0 can search the web for real-time information, generate multiple distinct images from one prompt, double-check its own outputs, and create functional QR codes — taking on more of the work between idea and final asset. Thinking-tier features are limited to Plus, Pro, and Business (Enterprise/Edu later). Source: X/@OpenAI thread, 2026-04-21 Source: PetaPixel, https://petapixel.com/2026/04/21/openai-claims-chatgpt-images-2-0-can-think/, 2026-04-21

API and pricing

gpt-image-2 pricing (per 1M tokens): Source: OpenAI Developer Community, https://community.openai.com/t/introducing-gpt-image-2-available-today-in-the-api-and-codex/1379479, 2026-04-21

Modality Input Cached input Output
Image $8.00 $2.00 $30.00
Text $5.00 $1.25 $10.00

Before release the model was tested on LM Arena under codenames "duct tape" / "maskingtape-alpha" / "gaffertape-alpha." Source: Gizmodo, https://gizmodo.com/openai-unveils-new-image-generator-to-usher-in-an-ai-slop-renaissance-2000749159, 2026-04-21

Where it fits Kevin's stack

The Codex integration matters most: generating on-brand assets directly inside an agentic coding loop (GPT-5-Codex) closes the gap between building UI and producing its imagery. Thinking-mode self-check + web search make it a candidate generator inside structured-output pipelines rather than a one-shot toy. Competes with Google's Nano Banana 2 (Gemini image line) and is the OpenAI counterpart to Claude Fable 5's design/motion one-shots; for hand-built layout-to-image work see Illustrated Manuscript (Pretext). Source: compiled, 2026-06-12


Timeline

  • 2026-04-21 | OpenAI announced ChatGPT Images 2.0 / gpt-image-2: instant + thinking modes, dense multilingual text, 3:1–1:3 aspect ratios, 2K (ChatGPT) / 4K beta (API), Dec 2025 cutoff; live in ChatGPT, Codex, and API. 25,198 likes, 6,985 bookmarks. Source: X/@OpenAI thread, 2026-04-21
  • 2026-06-12 | Page created from AI-models bookmark absorption (record 2046670977145372771); enriched with the OpenAI developer announcement, VentureBeat, PetaPixel, and Gizmodo coverage. Source: x-bookmark-absorb, 2026-06-12
  • 2026-06-30 | Deep-reviewed the local launch thumbnail and sampled video frames. Kept the row as a tool/model capability update and capsule input, with no new local skill. Source: X bookmark artifact audit, 2026-06-30