- calendar_today August 10, 2025
OpenAI introduced “Images in ChatGPT,” which embeds image generation capabilities within the ChatGPT interface. The introduction of the GPT-4o model provides users with new capabilities to generate images during their chat dialogues which represents a transformative advancement for AI-generated content creation.
The “Images in ChatGPT” feature offers sophisticated image generation capabilities to all ChatGPT users, regardless of their subscription level, including Plus, Pro, Team, and free. OpenAI spokesperson Taya Christianson confirmed that free tier users face similar usage restrictions as DALL-E 3, which allow for three images per day, although these limits might change depending on the level of demand. DALL-E fans continue to have access through a specialized GPT system.
OpenAI’s research head Gabriel Goh emphasized GPT-4o’s transformative capacities as an “omnimodal” system able to process multiple data formats like text, images, audio, and video. The model’s upgraded “binding” capability marks a significant improvement which solves a persistent problem in AI image generation. Previous models frequently fail to keep object attributes distinct, but GPT-4o successfully controls 15 to 20 objects without confusing their colors or shapes.
The system presents its best advancement through improved text rendering capabilities. AI-generated images have historically displayed issues with text becoming distorted or meaningless. Goh explained that developing the model required repeated iterations, which took several months to perfect. Despite understanding that flawless text presentation continues to be difficult, particularly with small text components, the team has reached a level of reliability that makes text in images dependable for users.
The system utilizes an autoregressive architectural approach, which differentiates it from the diffusion models common in image generation. The technique, which builds images from left to right and top to bottom following a text generation pattern, is thought to enhance its text rendering and binding features.
OpenAI demonstrated the system’s multiple use cases during a briefing, including creating detailed scientific diagrams like Newton’s prism experiment with precise labels, generating multi-panel comics that maintain consistent characters and dialogue throughout, and designing informational posters with exact text information. The team demonstrated practical applications, which included creating transparent background images suitable for stickers and restaurant menus, along with logos.
Multimodal product lead Jackie Shannon for ChatGPT pointed out how the system utilizes extensive world knowledge. She explained that while she draws an image based on her personal capabilities, she also utilizes the complete world knowledge she has acquired. The system incorporates world knowledge, which enables users to request an image of Newton’s prism experiment without needing to provide an explanation to retrieve the correct image.
OpenAI suggests that users should wait longer for image generation because the improved quality and enhanced capabilities provide sufficient value for the delay. Shannon acknowledged the need to reduce latency but emphasized that advanced image quality and extensive world knowledge compensate for waiting time.
Key Technological Advancements: Binding, Text Rendering, and Architectural Shifts
The GPT-4o model delivers major technological improvements through its “binding” function which enables precise depiction of intricate scenes containing multiple objects. The new text rendering feature, developed through extensive iterative development, eliminates a major problem seen in earlier AI image generators. The move toward autoregressive image generation methods represents a departure from conventional diffusion models, which might be responsible for these observed improvements.
Safeguards and User Empowerment: Addressing Misuse and Ensuring Responsible AI
OpenAI emphasized its robust safeguard measures after receiving concerns about potential misuse. The system functions to block watermark removal while also stopping the creation of sexual deepfakes and rejecting requests for CSAM. Generated images will contain standard C2PA metadata to identify them as OpenAI creations despite lacking visual watermarks. The company operates its own internal systems to verify images.
Shannon explained that while no system achieves perfection in this area, they remain dedicated to improving safeguards and consider their approach as an initial framework. Users who generate images through ChatGPT retain ownership rights to those images and may use them freely according to our usage guidelines.
OpenAI’s “Images in ChatGPT” feature improves its primary product while establishing a benchmark for powerful AI image creation and simultaneously handles potential technology risks.






