RETROSPECTIVE RECORD · PREPARED 16 SEPTEMBER 2026The record · 100 retrospective records ↗

The record / Models

Models / From the record · 30 November 2022 event · prepared 16 September 2026

OpenAI called ChatGPT's 2022 debut a research preview

The launch post names its RLHF method and lists specific failure modes, but carries no benchmark table of its own.

Visual for this record: OpenAI called ChatGPT's 2022 debut a research preview
Visual published by comercioynegocios.org, shown for identification of the record. Credit: comercioynegocios.org · source page ↗ Rights: owner-review-pending.

A dialogue interface on a sibling model

OpenAI's launch post, dated 30 November 2022, introduces ChatGPT as 'a sibling model to InstructGPT', fine-tuned from a model in the GPT-3.5 series that the post says 'finished training in early 2022'. The post calls the release a 'research preview', free to use while OpenAI gathered feedback, and frames it as one step in what it calls 'iterative deployment' rather than a finished product. That framing is worth keeping: a research preview invites a different reading than a capability launch, and the post repeats the word 'preview' rather than presenting comparative results.

The training method the post describes

OpenAI states that ChatGPT was trained with reinforcement learning from human feedback, 'using the same methods as InstructGPT, but with slight differences in the data collection setup'. Human trainers wrote conversations playing both the user and an assistant; that dialogue set was mixed with the existing InstructGPT data; a reward model was then trained on ranked comparisons of sampled responses, and the policy was fine-tuned against that reward model with Proximal Policy Optimization. The only quantitative evidence behind this general method sits in OpenAI's earlier instruction-following announcement and the underlying paper: labellers preferred outputs from a 1.3-billion-parameter InstructGPT model over a 175-billion-parameter GPT-3 model it was compared against. No equivalent preference study accompanied the ChatGPT post itself.

What was disclosed, and what was not measured

OpenAI's limitations section is unusually specific for a launch document. It states the model 'sometimes writes plausible-sounding but incorrect or nonsensical answers', that it is 'sensitive to tweaks to the input phrasing', that it is 'often excessively verbose', that it tends to guess at an ambiguous query rather than ask a clarifying question, and that it will 'sometimes respond to harmful instructions or exhibit biased behavior' despite a moderation filter the post admits will produce 'false negatives and positives for now'. These are vendor-disclosed limitations, named without a published rate at which any of them occurs. The InstructGPT paper remains the more rigorous document of the two because it reports a measured human-preference comparison rather than a prose description of behavior. Reading the ChatGPT post as a capability announcement, rather than as a list of known failure modes attached to a free preview, overstates what OpenAI itself claimed on the day.

Questions to carry into your own evaluation

  • Does a claimed improvement come with a measured comparison, or only a prose description of behavior?
  • Which of the five listed 2022 limitations are still present in a current product, and which have since been addressed by later fine-tuning?
  • Is the task you are testing close to the labeller-written demonstrations behind the training data, or closer to the ambiguous query the post says the model tends to guess at?

ChatGPT's research-preview post reads as two documents in one: a short account of an RLHF recipe with no new quantitative results of its own, and an unusually candid list of failure modes that OpenAI chose to publish rather than omit. Both are useful only if the second is not mistaken for the first.

Sources & reading trail

Introducing ChatGPT ↗

The launch post's own description of ChatGPT as a research preview, its RLHF training method, and its listed limitations.

Source published: 30 November 2022 · Retrieved: 16 September 2026

Aligning language models to follow instructions ↗

The InstructGPT announcement that ChatGPT is described as a sibling of, including the 1.3B-vs-175B labeller preference result.

Source published: 27 January 2022 · Retrieved: 16 September 2026

Training language models to follow instructions with human feedback ↗

The peer-reviewable record of the InstructGPT method and its stated limitation that the model 'still makes simple mistakes'.

Source published: 4 March 2022 · Retrieved: 16 September 2026

Papers and official documents establish the record; the reading and the questions are Model Field Guide editorial analysis. This retrospective draft does not imply the site published on the event date.