RETROSPECTIVE RECORD · PREPARED 16 SEPTEMBER 2026The record · 100 retrospective records ↗

The record / Models

Models / From the record · 13 May 2024 event · prepared 16 September 2026

GPT-4o shipped as one model in stages, not all at once

OpenAI's GPT-4o announcement and its later system card disclose different levels of detail about rollout and voice-mode risk.

Visual for this record: GPT-4o shipped as one model in stages, not all at once
Visual published by nextian.com, shown for identification of the record. Credit: nextian.com · source page ↗ Rights: owner-review-pending.

One network across four modalities, announced 13 May 2024

OpenAI's Hello GPT-4o announcement describes a single model trained end-to-end across text, vision and audio, rather than the pipeline used for the earlier ChatGPT Voice Mode, which chained a transcription model, a text model and a text-to-speech model together. That pipeline, the post states, lost information about tone, overlapping speakers and background noise, and averaged 2.8 seconds (GPT-3.5) to 5.4 seconds (GPT-4) of latency. GPT-4o responds to audio in as little as 232 milliseconds, averaging 320 milliseconds — the announcement compares this to typical human conversational response time. On text and code the post states GPT-4o matches GPT-4 Turbo, while running faster and 50% cheaper through the API.

A staged release, not a single switch

The announcement is explicit that GPT-4o did not ship complete: at launch only text and image inputs with text output were made broadly available, with audio and video capability following 'over the upcoming weeks and months,' and audio outputs initially limited to a set of preset voices under existing safety policy. OpenAI states it ran a Preparedness Framework evaluation covering cybersecurity, biological threats, persuasion and model autonomy, reporting no category scoring above Medium risk, with more than 70 external red-teamers engaged across social psychology, bias and misinformation. The detail behind those evaluations was deferred: the announcement says fuller findings would appear in a forthcoming system card.

What the system card added three months later

The GPT-4o System Card, published 8 August 2024, names five specific risk categories tied to the voice modality: unauthorized voice generation, speaker identification, ungrounded inference and sensitive trait attribution from voice, generating disallowed audio content, and generating erotic or violent speech. Its Preparedness Framework scorecard records cybersecurity, biological threats and model autonomy each as Low, with persuasion as the sole Medium-rated category. The gap between the May announcement's general risk-level claim and the August card's named categories illustrates why a launch post and a system card do different jobs: one states a conclusion, the other documents how it was reached and for which specific capability.

  • Which of the disclosed voice-mode risk categories are relevant to a specific deployment, and were they re-tested after later fine-tuning?
  • Is a 'matches GPT-4 Turbo' claim being applied to a task type the announcement actually measured, or extrapolated to a different one?
  • Which modalities were available at the announcement date, and which followed later — and does the documentation distinguish the two?

An omni-modal launch announcement and its system card were published almost three months apart, covering claims of different specificity. Treating the first as marketing framing and the second as the evidentiary record is a reasonable default when the two disagree on detail.

Sources & reading trail

Hello GPT-4o ↗

Announces GPT-4o's omni-modal design, latency figures, pricing/speed versus GPT-4 Turbo, and the staged rollout of audio and video output.

Source published: 13 May 2024 · Retrieved: 16 September 2026

GPT-4o System Card ↗

Names specific voice-mode risk categories and reports the Preparedness Framework scorecard ratings.

Source published: 8 August 2024 · Retrieved: 16 September 2026

Papers and official documents establish the record; the reading and the questions are Model Field Guide editorial analysis. This retrospective draft does not imply the site published on the event date.