With the ability to “think with images” in their chain of thought, OpenAI’s newest AI reasoning models, o3 and o4-mini, represent a significant advancement in visual perception. These models can manipulate and analyze visual content as part of their reasoning process while independently using all ChatGPT tools, including web browsing, Python, and image generation.
Capabilities For Image Manipulation
The o3 and o4-mini models from OpenAI revolutionize the way AI engages with visual content by introducing previously unheard-of picture editing capabilities. As part of their reasoning process, these models are able to autonomously process photos by cropping, zooming, rotating, and using other basic image processing techniques. The AI can extract insights from low-quality, hazy, or defective photographs thanks to this intrinsic functionality, which eliminates the need for separate specialized models.
The useful applications are impressive: users may submit images of build issues, diagrams from PDFs, or whiteboard sketches to get thorough examination. When given a picture of an economics issue set, for instance, the models can manipulate the image to better grasp its content and then provide detailed answers. This smooth transition between textual and visual thinking is a big step toward real multimodal AI, which can actively work with images to solve challenging problems rather than just seeing them.
Chain-Of-Thought Visual Reasoning
The ability of OpenAI’s o3 and o4-mini to directly include visuals into their reasoning process, in contrast to conventional models that solely process text, is revolutionizing the way AI tackles challenging issues. As part of their chain-of-thought reasoning, these models are able to change visual content by cropping, zooming, rotating, and using other image processing techniques. This feature greatly enhances performance on tasks requiring spatial awareness by enabling the models to glean insights even from faulty photographs.
According to research, giving AI the ability to think visually significantly increases its success rates on challenging tasks. Success rates increased from 61% to 85% in one experiment when AI was able to depict its ideas instead of depending just on verbal descriptions. These models are especially useful for evaluating business charts, diagrams, sketches, and other visual data that previously presented difficulties for AI systems because of their multimodal chain-of-thought method, which lessens hallucinations and improves cross-modality thinking. The capacity to “think with images” is a significant step toward artificial intelligence that can reason more like a human.
Benchmark Performance For Multimodal
The o3 and o4-mini models from OpenAI set new benchmarks in visual reasoning tasks with their outstanding performance across multimodal benchmarks. With high compute settings, o3’s accuracy on the ARC-AGI benchmark was 87.5%, which is equivalent to human performance of 85%. In the meantime, when provided with a Python interpreter, o4-mini successfully saturated the AIME 2025 benchmark with an astounding 99.5% score.
The models perform exceptionally well in a number of visual tasks, such as zero-shot object detection, document OCR, visual question answering (VQA), and item counting. Both models outperformed their predecessors in expert evaluations, demonstrating better instruction following and more practical, verifiable responses. With o3 outperforming earlier models on benchmarks like Codeforces, SWE-bench, and MMMU and committing 20% fewer significant errors on challenging real-world tasks than OpenAI o1, this performance improvement is cross-domain.

