RETROSPECTIVE RECORD · PREPARED 16 SEPTEMBER 2026The record · 100 retrospective records ↗

The record / Safety & policy

Safety & policy / From the record · 27 December 2023 event · prepared 16 September 2026

The Times sued OpenAI in 2023 over training and outputs

A filed complaint alleges unauthorised training and near-verbatim reproduction; OpenAI's public reply calls the claims without merit.

storage.courtlistener.comprimary record

The New York Times Company v. Microsoft Corporation, 1:23-cv-11195 (Complaint)

Document
27 December 2023
Event
27 December 2023
Retrieved
16 September 2026
No visual was published with this record, so its primary document stands in its place.

What the filed complaint records

The New York Times Company filed suit against Microsoft and several OpenAI entities on 27 December 2023, docketed as case 1:23-cv-11195 in the US District Court for the Southern District of New York. According to the case docket, the action is filed as a copyright infringement claim, before Judge Sidney Stein. The complaint, reviewed for this entry, alleges that the defendants used Times articles to train their language models without authorisation or payment, and that prompted outputs can closely paraphrase or substantially reproduce specific facts, phrases and structures from published Times reporting. These are the complaint's allegations, not proven findings; a filed pleading states what a plaintiff intends to prove, not what a court has held.

What OpenAI said publicly

OpenAI responded on 8 January 2024 with a public statement saying plainly that it "regard[s] The New York Times' lawsuit to be without merit". The statement argues that training on publicly available material is fair use, that OpenAI offers a documented opt-out mechanism which the Times had already adopted in August 2023, and that any "regurgitation" of training data is "a rare bug" the company is "working to drive to zero". It also asserts, without providing the underlying prompts for independent verification here, that the reproductions the Times identified came from "lengthy excerpts of articles" entered as prompts and that the source articles had already proliferated on multiple third-party websites. That is OpenAI's account of its own product; it is not an independent finding.

Why the dispute matters beyond one publisher

The case turns on two separate, contestable questions the complaint and the reply frame differently: whether training itself is a use requiring permission, and whether specific model outputs reproduce protected expression closely enough to infringe it. Both questions recur across the industry wherever a model is trained on material its owner did not license, which is why data provenance - what a system was trained on, and under what terms - has become a question deployers are increasingly expected to be able to answer, whatever this particular suit's outcome turns out to be.

  • Can the provider of a model in use disclose, even in general terms, the categories of data it was trained on and how consent or licensing was handled?
  • If an output closely resembles a specific published source, does that reflect memorised training data or a coincidental convergence on well-known facts?
  • Does an opt-out mechanism for future training address material already used in past training runs?

Nothing here resolves the underlying legal questions, which remain contested in this and related cases. What the record supports is narrower: a major publisher has alleged unauthorised use in a filed complaint, the defendant has publicly and specifically denied it, and the dispute is proceeding in litigation rather than being settled by either side's public statement.

Sources & reading trail

The New York Times Company v. Microsoft Corporation, 1:23-cv-11195 (Complaint) ↗

The filed complaint, alleging unauthorised training use of Times content and reproduction of Times material in model outputs, and describing the relief sought.

Source published: 27 December 2023 · Retrieved: 16 September 2026

The New York Times Company v. Microsoft Corporation, 1:23-cv-11195 ↗

Confirms the filing date, case number, presiding judge, and copyright-infringement cause of action, and lists Microsoft and multiple OpenAI entities as defendants.

Source published: Not established · Retrieved: 16 September 2026

OpenAI and journalism ↗

OpenAI's public response calling the lawsuit without merit, describing its fair-use position, its opt-out mechanism, and its account of the alleged regurgitation.

Source published: 8 January 2024 · Retrieved: 16 September 2026

Papers and official documents establish the record; the reading and the questions are Model Field Guide editorial analysis. This retrospective draft does not imply the site published on the event date.