RETROSPECTIVE RECORD · PREPARED 16 SEPTEMBER 2026The record · 100 retrospective records ↗

The record / Models

Models / From the record · 14 March 2023 event · prepared 16 September 2026

GPT-4's report withheld architecture, size and training data

OpenAI disclosed exam and benchmark results while declining to state parameter count, architecture or compute.

Visual for this record: GPT-4's report withheld architecture, size and training data
Visual published by figures.semanticscholar.org, shown for identification of the record. Credit: figures.semanticscholar.org · source page ↗ Rights: owner-review-pending.

A report that names what it will not say

OpenAI's announcement, dated 14 March 2023, describes GPT-4 as 'a large multimodal model (accepting image and text inputs, emitting text outputs)' and states that the training run finished roughly six months earlier, since the accompanying system card notes the model 'finished training in August of 2022'. The technical report, submitted the same day, is explicit about its own limits: it states the document 'contains no further details about the architecture (including model size), hardware, training compute, dataset construction, training method' citing 'the competitive landscape and the safety implications of large-scale models'. That sentence is itself a primary fact worth citing: OpenAI did not omit these details by accident, and named the omission in the report.

What was measured, and under what conditions

What the report does disclose is a set of exam and benchmark results. OpenAI's post states GPT-4 'passes a simulated bar exam with a score around the top 10% of test takers', against GPT-3.5's 'bottom 10%', and that the exams used 'the most recent publicly available tests' or purchased 2022-2023 practice editions, with 'no specific training for these exams' though the post concedes 'a minority of the problems in the exams were seen by the model during training'. A translated-MMLU experiment is also reported, with GPT-4 said to outperform GPT-3.5, Chinchilla and PaLM in 24 of 26 tested languages. These are OpenAI's own selected benchmarks and exam settings; the report does not claim an independent audit of the exam results, and states plainly that some contamination is possible.

What the system card adds, and what it does not resolve

The system card separates two internal versions, 'GPT-4-early' and 'GPT-4-launch', to show how safety mitigations changed observed behaviour, and records that OpenAI 'engaged more than 50 experts' in adversarial testing. It also states that an evaluation by the Alignment Research Center concluded the model was 'probably not yet capable' of autonomously replicating and acquiring resources. None of this substitutes for the missing architecture and scale figures: a reader cannot compare GPT-4's exam score to another model's on a per-parameter or per-FLOP basis, only on the exam score itself, because the report withholds the terms that would make such a comparison meaningful.

Questions to carry into your own evaluation

  • Is a cited GPT-4 benchmark number attributable to the technical report, the system card, or a third-party reproduction, and do they use the same prompting setting?
  • Does an exam or benchmark comparison account for the stated possibility that some problems were present in training data?
  • What claims become impossible to verify once architecture, size and training data are withheld, and does that change how much weight the exam scores alone should carry?

The GPT-4 report is unusual for naming its own confidentiality rather than staying silent about it. That is more transparent than no disclosure at all, but it does not convert the exam and benchmark figures into an independently verifiable capability claim; they remain OpenAI's measurements, under OpenAI's stated conditions, of a model whose basic construction the report does not describe.

Sources & reading trail

GPT-4 Technical Report ↗

The report's own statement that it withholds architecture, size, hardware, training compute and dataset details, plus its stated exam and translated-MMLU results.

Source published: 15 March 2023 · Retrieved: 16 September 2026

GPT-4 ↗

The launch announcement's bar-exam and benchmark claims and its description of the exam methodology and contamination caveat.

Source published: 14 March 2023 · Retrieved: 16 September 2026

GPT-4 System Card ↗

Confirms training finished in August 2022, describes the GPT-4-early/GPT-4-launch comparison, and records the ARC autonomous-replication evaluation conclusion.

Source published: Not established · Retrieved: 16 September 2026

Papers and official documents establish the record; the reading and the questions are Model Field Guide editorial analysis. This retrospective draft does not imply the site published on the event date.