
A report that names what it will not say
OpenAI's announcement, dated 14 March 2023, describes GPT-4 as 'a large multimodal model (accepting image and text inputs, emitting text outputs)' and states that the training run finished roughly six months earlier, since the accompanying system card notes the model 'finished training in August of 2022'. The technical report, submitted the same day, is explicit about its own limits: it states the document 'contains no further details about the architecture (including model size), hardware, training compute, dataset construction, training method' citing 'the competitive landscape and the safety implications of large-scale models'. That sentence is itself a primary fact worth citing: OpenAI did not omit these details by accident, and named the omission in the report.
What was measured, and under what conditions
What the report does disclose is a set of exam and benchmark results. OpenAI's post states GPT-4 'passes a simulated bar exam with a score around the top 10% of test takers', against GPT-3.5's 'bottom 10%', and that the exams used 'the most recent publicly available tests' or purchased 2022-2023 practice editions, with 'no specific training for these exams' though the post concedes 'a minority of the problems in the exams were seen by the model during training'. A translated-MMLU experiment is also reported, with GPT-4 said to outperform GPT-3.5, Chinchilla and PaLM in 24 of 26 tested languages. These are OpenAI's own selected benchmarks and exam settings; the report does not claim an independent audit of the exam results, and states plainly that some contamination is possible.
What the system card adds, and what it does not resolve
The system card separates two internal versions, 'GPT-4-early' and 'GPT-4-launch', to show how safety mitigations changed observed behaviour, and records that OpenAI 'engaged more than 50 experts' in adversarial testing. It also states that an evaluation by the Alignment Research Center concluded the model was 'probably not yet capable' of autonomously replicating and acquiring resources. None of this substitutes for the missing architecture and scale figures: a reader cannot compare GPT-4's exam score to another model's on a per-parameter or per-FLOP basis, only on the exam score itself, because the report withholds the terms that would make such a comparison meaningful.
Questions to carry into your own evaluation
- Is a cited GPT-4 benchmark number attributable to the technical report, the system card, or a third-party reproduction, and do they use the same prompting setting?
- Does an exam or benchmark comparison account for the stated possibility that some problems were present in training data?
- What claims become impossible to verify once architecture, size and training data are withheld, and does that change how much weight the exam scores alone should carry?
The GPT-4 report is unusual for naming its own confidentiality rather than staying silent about it. That is more transparent than no disclosure at all, but it does not convert the exam and benchmark figures into an independently verifiable capability claim; they remain OpenAI's measurements, under OpenAI's stated conditions, of a model whose basic construction the report does not describe.
Sources & reading trail
The report's own statement that it withholds architecture, size, hardware, training compute and dataset details, plus its stated exam and translated-MMLU results.
Source published: 15 March 2023 · Retrieved: 16 September 2026
The launch announcement's bar-exam and benchmark claims and its description of the exam methodology and contamination caveat.
Source published: 14 March 2023 · Retrieved: 16 September 2026
Confirms training finished in August 2022, describes the GPT-4-early/GPT-4-launch comparison, and records the ARC autonomous-replication evaluation conclusion.
Source published: Not established · Retrieved: 16 September 2026
Papers and official documents establish the record; the reading and the questions are Model Field Guide editorial analysis. This retrospective draft does not imply the site published on the event date.