
Testing disclosed alongside a capability release
Alongside GPT-4's release in mid-March 2023, OpenAI published the GPT-4 System Card, a document distinct from the companion GPT-4 Technical Report posted to arXiv on 15 March 2023, which covers capability benchmarks rather than safety testing. The card states that OpenAI engaged more than 50 external experts across domains including cybersecurity, biorisk, misinformation and law to red-team the model, and that internal adversarial testing of the launch version ran through 10 March 2023. The document compares an early, lightly mitigated version of the model, 'GPT-4-early,' against the version actually shipped, 'GPT-4-launch,' throughout, which lets a reader see what the mitigation work changed rather than only what the finished product does.
What was measured, and what mitigation did to it
The card reports that GPT-4-launch reduced responses to disallowed content by 82% relative to GPT-3.5 on the company's internal test sets, and that GPT-4 produced toxic completions on the RealToxicityPrompts benchmark 0.73% of the time against 6.48% for GPT-3.5, figures the card itself states, drawn from OpenAI's own evaluation harness rather than an independent audit. The card also documents specific residual risks under GPT-4-early conditions: the model could narrow the search time for weapons-proliferation information, draft targeted phishing content when supplied background on a target, and exhibit uneven refusal behaviour, including a documented inconsistency where Morse-code-encoded prompts bypassed refusals that applied to the same request in plain English. A separate evaluation by the Alignment Research Center tested whether an early version could autonomously replicate and acquire resources, and concluded it was not effective at doing so under the conditions tested.
Why a system card is not an audit
The document is authored, scoped and released by OpenAI itself; the experts it lists as red-teamers contributed testing, but the card is not reviewed or attested to by an independent body, and OpenAI states outright that 'this system card is not comprehensive.' The distinction from a technical audit matters: an audit typically implies an external party checking a claim against evidence under some agreed standard, where a system card is the developer's own disclosure of what it tested and found, in whatever form and depth it chose. Meta's earlier description of the system card format, from February 2022, frames the genre as documentation meant to be understood by experts and nonexperts alike, a transparency goal distinct from a verification one.
- Which of the card's quantitative claims were measured by OpenAI's own classifiers, and which by an outside party?
- Does the card's list of tested domains match the actual deployment context this model is being used in?
- Has anything changed in the model since the card's testing window closed on 10 March 2023, and does a later system card exist covering it?
Read as a disclosure document rather than a certificate, the card is informative about what OpenAI looked for and found brittle at launch, and silent about anything it did not think to test.
Sources & reading trail
Primary document: red-teaming scope, internal testing date, quantitative mitigation figures, and residual risk examples.
Source published: Not established · Retrieved: 16 September 2026
Companion capability report, dating the arXiv submission and distinguishing capability reporting from the safety-focused system card.
Source published: 15 March 2023 · Retrieved: 16 September 2026
Defines the system card documentation genre and its transparency purpose, used to contrast a system card with an audit.
Source published: 23 February 2022 · Retrieved: 16 September 2026
Papers and official documents establish the record; the reading and the questions are Model Field Guide editorial analysis. This retrospective draft does not imply the site published on the event date.