Early preview
O

OpenAI

GPT-4o

Announced May 13, 2024

Text, vision and audio come together

01

What could it do?

Combine text and image assistance, with integrated audio interaction demonstrated as a major part of the model’s direction.

02

What changed?

The model was designed around multiple input and output types rather than a voice pipeline of separate models.

WHY IT MATTERED

Multimodal interaction

OpenAI began extending GPT-4o and additional tools to free ChatGPT users, subject to limits.

03

Where it fell short

A capability shown in a launch demonstration was not necessarily available in every product on that day.

Which release does this page cover?

May 2024 launch. Text and image capabilities arrived before the new voice experience. Check sources

Name means
Omni
Launch date
May 13, 2024
Initial rollout
Text and vision first

What this meant in practice

Access and capability can improve separately. More people receiving a feature matters even when the underlying task already existed. A timeline should show both the research announcement and when users could actually try it.

ILLUSTRATIVE TASK · NOT A TEST RESULT

Discuss an image

Upload a photographed menu and ask about the options. Confirm names and prices against the photo before relying on them. This illustrates an interaction, not a measured success rate.

Common question

Were all the voice demos immediately available?

No. OpenAI described a staged rollout of the new audio and video capabilities.

THE USEFUL CONTEXT

Understanding the GPT-4o milestone

Leapscope explanation · Reviewed October 7, 2026. Examples and practical interpretations below are editorial, not independent test results.

Why different input types matter

OpenAI described GPT-4o as a model trained across text, vision and audio. The launch also involved staged access, so the model’s described abilities and the features available in a particular product were not identical.

For a user, the value of an input type is the work it removes. A photograph can save someone from manually typing a table. Spoken interaction can make a conversation easier while their hands are occupied. These benefits need their own checks: transcription errors, image quality and interruptions can matter more than a broad reasoning score in the workflow being evaluated.

GPT-4o launch announcement

From an image to something usable

Use a fictional café menu as an example. Ask a model to extract item names and prices into a table. Check every cell against the image before asking it to calculate a total. If a price is unreadable, the useful behavior is to mark it as uncertain, not fill the gap with a plausible amount.

A second task might ask for a concise description of the menu’s layout for someone who cannot see it clearly. That requires different scoring from extracting prices. A description can omit an irrelevant decorative flourish and still be useful, while a missing decimal point may ruin the table. Both examples illustrate potential workflows rather than measured performance on this site.

A more useful comparison than one overall winner

For this release, compare input support, response quality, availability and the time needed to finish a task separately. A fast first answer may require several corrections. An apparently slower answer may need less editing. If you measure speed, include the human work required to reach an acceptable result.

Keep historical prices and current prices separate too. This page explains a 2024 announcement; it should not be used as a live purchasing guide. The current model catalogue can help locate later releases, but a buying decision still needs today’s product access and pricing checked at the source. The same distinction applies when a model name remains familiar while the product around it changes.

FOLLOW THE EVIDENCE

What to watch next

Changes that would make this story worth revisiting:

  • A feature moving from a demonstration to documented general availability.
  • Independent tests that separate audio, image reading and reasoning instead of folding them into a single claim.

Questions about this milestone

Did GPT-4o make every kind of task better by the same amount?

That cannot be inferred from the launch. A model may change speed, access and particular capabilities differently. Use results for the task you actually care about.

What should I compare with GPT-4?

Choose an identical input and a defined output, then compare correctness, corrections and time to completion. Record whether either product supplied extra tools.

Sources checked Oct 7, 2026

Release facts were checked against the sources below. Performance claims belong to the developers; we have not independently tested these models.

OpenAI: Hello GPT-4o OpenAI: GPT-4o rollout to ChatGPT users

Benchmark results

EPOCH AI CAPABILITY ESTIMATE
129.0index points
Tested variant: GPT-4o (May 2024)Source interval: 123.8 to 131.7Variant date in source: May 13, 2024

A benchmark estimate, not a percentage or capability multiplier. Reasoning settings are not specified in this source table. Historical estimates can change in later snapshots.

Epoch AI methodology ↗Download the source snapshotChecked Oct 7, 2026 · CC BY 4.0
PUBLISHED BENCHMARK RESULT
1346rating points
Tested: gpt-4o-2024-05-13Text Arena Overall · October 8, 2026Reported interval: 1343 to 1349112,881 votes

One explicitly named variant per release. Scores come from the same Overall snapshot; preliminary entries and reported intervals are preserved. These are current ratings of earlier variants, not their launch day ratings.

Source: Text Arena ↗Download selected resultsChecked Oct 8, 2026
FOLLOW WHAT HAPPENS NEXT

Breakthroughs, with the followup.

A weekly brief on new discoveries, meaningful checks and what you can actually use.