OpenAI
GPT-4o
Announced May 13, 2024
Text, vision and audio come together
What could it do?
Combine text and image assistance, with integrated audio interaction demonstrated as a major part of the model’s direction.
What changed?
The model was designed around multiple input and output types rather than a voice pipeline of separate models.
Multimodal interaction
OpenAI began extending GPT-4o and additional tools to free ChatGPT users, subject to limits.
Where it fell short
A capability shown in a launch demonstration was not necessarily available in every product on that day.
May 2024 launch. Text and image capabilities arrived before the new voice experience. Check sources
- Name means
- Omni
- Launch date
- May 13, 2024
- Initial rollout
- Text and vision first
What this meant in practice
Access and capability can improve separately. More people receiving a feature matters even when the underlying task already existed. A timeline should show both the research announcement and when users could actually try it.
Discuss an image
Upload a photographed menu and ask about the options. Confirm names and prices against the photo before relying on them. This illustrates an interaction, not a measured success rate.
Common question
Were all the voice demos immediately available?
No. OpenAI described a staged rollout of the new audio and video capabilities.
THE USEFUL CONTEXT
Understanding the GPT-4o milestone
Leapscope explanation · Reviewed October 7, 2026. Examples and practical interpretations below are editorial, not independent test results.
Why different input types matter
OpenAI described GPT-4o as a model trained across text, vision and audio. The launch also involved staged access, so the model’s described abilities and the features available in a particular product were not identical.
For a user, the value of an input type is the work it removes. A photograph can save someone from manually typing a table. Spoken interaction can make a conversation easier while their hands are occupied. These benefits need their own checks: transcription errors, image quality and interruptions can matter more than a broad reasoning score in the workflow being evaluated.
From an image to something usable
Use a fictional café menu as an example. Ask a model to extract item names and prices into a table. Check every cell against the image before asking it to calculate a total. If a price is unreadable, the useful behavior is to mark it as uncertain, not fill the gap with a plausible amount.
A second task might ask for a concise description of the menu’s layout for someone who cannot see it clearly. That requires different scoring from extracting prices. A description can omit an irrelevant decorative flourish and still be useful, while a missing decimal point may ruin the table. Both examples illustrate potential workflows rather than measured performance on this site.
A more useful comparison than one overall winner
For this release, compare input support, response quality, availability and the time needed to finish a task separately. A fast first answer may require several corrections. An apparently slower answer may need less editing. If you measure speed, include the human work required to reach an acceptable result.
Keep historical prices and current prices separate too. This page explains a 2024 announcement; it should not be used as a live purchasing guide. The current model catalogue can help locate later releases, but a buying decision still needs today’s product access and pricing checked at the source. The same distinction applies when a model name remains familiar while the product around it changes.
FOLLOW THE EVIDENCE
What to watch next
Changes that would make this story worth revisiting:
- A feature moving from a demonstration to documented general availability.
- Independent tests that separate audio, image reading and reasoning instead of folding them into a single claim.
Questions about this milestone
Did GPT-4o make every kind of task better by the same amount?
That cannot be inferred from the launch. A model may change speed, access and particular capabilities differently. Use results for the task you actually care about.
What should I compare with GPT-4?
Choose an identical input and a defined output, then compare correctness, corrections and time to completion. Record whether either product supplied extra tools.
Release facts were checked against the sources below. Performance claims belong to the developers; we have not independently tested these models.
OpenAI: Hello GPT-4o OpenAI: GPT-4o rollout to ChatGPT usersBenchmark results
A benchmark estimate, not a percentage or capability multiplier. Reasoning settings are not specified in this source table. Historical estimates can change in later snapshots.
Epoch AI methodology ↗Download the source snapshotChecked Oct 7, 2026 · CC BY 4.0One explicitly named variant per release. Scores come from the same Overall snapshot; preliminary entries and reported intervals are preserved. These are current ratings of earlier variants, not their launch day ratings.
Source: Text Arena ↗Download selected resultsChecked Oct 8, 2026