RETROSPECTIVE RECORD · PREPARED 16 SEPTEMBER 2026The record · 100 retrospective records ↗

The record / Models

Models / From the record · 4 March 2024 event · prepared 16 September 2026

Claude 3 shipped three sizes with one published card

Anthropic's Claude 3 model card ties benchmark scores to their sampling conditions, including where Opus falls short of expert accuracy.

Visual for this record: Claude 3 shipped three sizes with one published card
Visual published by cdn.sanity.io, shown for identification of the record. Credit: cdn.sanity.io · source page ↗ Rights: owner-review-pending.

Three sizes, one published card, one launch date

On 4 March 2024 Anthropic announced the Claude 3 family: Haiku, Sonnet and Opus, three sizes sharing a 200,000-token context window at launch, with the announcement stating that a version supporting more than 1 million tokens was available to select customers. Listed API pricing at launch was $0.25/$1.25 per million input/output tokens for Haiku, $3/$15 for Sonnet, and $15/$75 for Opus — a historical price record from the announcement, not a current rate. The three-tier structure let a buyer trade capability against cost within one family rather than switching vendors for a cheaper option.

What the model card discloses about the numbers

The accompanying model card, The Claude 3 Model Family: Opus, Sonnet, Haiku, sets out benchmark results with the evaluation conditions attached rather than bare scores. On MMLU (5-shot), Opus scored 86.8%, Sonnet 79.0%, and Haiku 75.2%, alongside a comparison figure of 86.4% the card cites for GPT-4. On GPQA Diamond, a graduate-level science benchmark the card describes as difficult even for non-expert PhDs given thirty minutes and internet access, Opus reached 50.4% zero-shot with chain-of-thought reasoning, rising to 59.5% when the answer was taken by majority vote across 32 samples. The card also states that GPQA showed high variance under chain-of-thought sampling, which is why it reports means across ten evaluation rollouts rather than a single run — a methodological detail a bare leaderboard score would hide.

What the card does and does not claim

The card is explicit that domain-expert humans, given thirty minutes and internet access, score 60–80% on GPQA Diamond questions; Opus's reported 50.4% zero-shot sits below that range, and the card does not claim otherwise. On safety, Anthropic classifies all three Claude 3 models as ASL-2 under its Responsible Scaling Policy, after autonomous-replication and CBRN-related evaluations found no result crossing its pre-specified ASL-3 warning thresholds — but the card adds its own caveat, that 'evaluations are a hard scientific problem, and our methodology is still being improved.' That sentence is worth reading before treating an ASL-2 classification as a permanent clearance rather than a dated assessment against a specific, evolving test suite.

  • Which benchmark in the card most resembles the task actually being evaluated, and at what sampling setting?
  • Does the cheapest tier in the family clear the accuracy bar the task requires, or is the largest model load-bearing?
  • Has the cited safety evaluation methodology been superseded by a later, stricter version of the same framework?

A tiered model family turns a single build-versus-buy decision into three, each with its own cost and accuracy trade-off recorded in the same document. Reading the model card's stated conditions, not just the top-line score, is what makes that comparison meaningful rather than promotional.

Sources & reading trail

Introducing the next generation of Claude ↗

Announces the three-model family, 200K context window, launch pricing per tier, and a >99% needle-in-a-haystack claim for Opus.

Source published: 4 March 2024 · Retrieved: 16 September 2026

The Claude 3 Model Family: Opus, Sonnet, Haiku ↗

Discloses benchmark scores with sampling conditions (MMLU, GPQA Diamond), the ASL-2 safety classification, and the stated limits of the evaluation methodology.

Source published: Not established · Retrieved: 16 September 2026

Papers and official documents establish the record; the reading and the questions are Model Field Guide editorial analysis. This retrospective draft does not imply the site published on the event date.