
One network across four modalities, announced 13 May 2024
OpenAI's Hello GPT-4o announcement describes a single model trained end-to-end across text, vision and audio, rather than the pipeline used for the earlier ChatGPT Voice Mode, which chained a transcription model, a text model and a text-to-speech model together. That pipeline, the post states, lost information about tone, overlapping speakers and background noise, and averaged 2.8 seconds (GPT-3.5) to 5.4 seconds (GPT-4) of latency. GPT-4o responds to audio in as little as 232 milliseconds, averaging 320 milliseconds — the announcement compares this to typical human conversational response time. On text and code the post states GPT-4o matches GPT-4 Turbo, while running faster and 50% cheaper through the API.
A staged release, not a single switch
The announcement is explicit that GPT-4o did not ship complete: at launch only text and image inputs with text output were made broadly available, with audio and video capability following 'over the upcoming weeks and months,' and audio outputs initially limited to a set of preset voices under existing safety policy. OpenAI states it ran a Preparedness Framework evaluation covering cybersecurity, biological threats, persuasion and model autonomy, reporting no category scoring above Medium risk, with more than 70 external red-teamers engaged across social psychology, bias and misinformation. The detail behind those evaluations was deferred: the announcement says fuller findings would appear in a forthcoming system card.
What the system card added three months later
The GPT-4o System Card, published 8 August 2024, names five specific risk categories tied to the voice modality: unauthorized voice generation, speaker identification, ungrounded inference and sensitive trait attribution from voice, generating disallowed audio content, and generating erotic or violent speech. Its Preparedness Framework scorecard records cybersecurity, biological threats and model autonomy each as Low, with persuasion as the sole Medium-rated category. The gap between the May announcement's general risk-level claim and the August card's named categories illustrates why a launch post and a system card do different jobs: one states a conclusion, the other documents how it was reached and for which specific capability.
- Which of the disclosed voice-mode risk categories are relevant to a specific deployment, and were they re-tested after later fine-tuning?
- Is a 'matches GPT-4 Turbo' claim being applied to a task type the announcement actually measured, or extrapolated to a different one?
- Which modalities were available at the announcement date, and which followed later — and does the documentation distinguish the two?
An omni-modal launch announcement and its system card were published almost three months apart, covering claims of different specificity. Treating the first as marketing framing and the second as the evidentiary record is a reasonable default when the two disagree on detail.
Sources & reading trail
Announces GPT-4o's omni-modal design, latency figures, pricing/speed versus GPT-4 Turbo, and the staged rollout of audio and video output.
Source published: 13 May 2024 · Retrieved: 16 September 2026
Names specific voice-mode risk categories and reports the Preparedness Framework scorecard ratings.
Source published: 8 August 2024 · Retrieved: 16 September 2026
Papers and official documents establish the record; the reading and the questions are Model Field Guide editorial analysis. This retrospective draft does not imply the site published on the event date.