RETROSPECTIVE RECORD · PREPARED 16 SEPTEMBER 2026The record · 100 retrospective records ↗

The record / Evaluation

Evaluation / From the record · 16 November 2022 event · prepared 16 September 2026

HELM measured many models the same way at once

The 2022 HELM paper standardised scenarios and metrics across 30 models to expose how unevenly prior work had compared them.

Visual for this record: HELM measured many models the same way at once
Visual published by crfm.stanford.edu, shown for identification of the record. Credit: crfm.stanford.edu · source page ↗ Rights: owner-review-pending.

A benchmark of benchmarks, in public

The Holistic Evaluation of Language Models paper, posted to arXiv on 16 November 2022 by a team at Stanford's Center for Research on Foundation Models, did not propose a new task. It proposed a shared protocol: evaluate a fixed list of models on a fixed list of scenarios, under the same few-shot prompting conditions, and publish every raw prompt and completion alongside the scores. The accompanying announcement, published the next day, frames the problem directly: language models were being evaluated on incompatible subsets of tasks, so a comparison between two papers' reported numbers was often not a real comparison at all.

What HELM standardised

The initial release ran 30 models from 12 providers across 42 scenarios, measured against seven metric categories: accuracy, calibration, robustness, fairness, bias, toxicity and efficiency. Sixteen of those scenarios were designated core and evaluated with close to complete metric coverage. The paper's own figure for the state of the field before this exercise is stark: on average, models had previously been evaluated on only 17.9 percent of the scenarios HELM treats as core, and some widely cited models shared no scenario in common at all. Raising that figure to 96.0 percent coverage is the paper's central methodological claim, not a capability claim about any one model.

What the exercise exposed, and what it does not settle

Running the same scenarios across many models surfaced coverage gaps that narrower leaderboards had hidden: dialect variation in question answering, calibration under uncertainty and toxicity were each scenarios that most prior published evaluations had simply skipped. That is a genuine contribution distinct from any single score. It is also, by design, a snapshot: HELM evaluates the models available when a given run is executed, using the prompting conventions current at that time, so a table from 2022 says little about a later model unless the scenario set and evaluation code have both been rerun. The project describes itself as a living benchmark for this reason, and the current state of its taxonomy, as retrieved on 16 September 2026 at the HELM project site, reflects ongoing revisions rather than the fixed 2022 configuration the paper reports.

Questions to carry into your own evaluation

  • Which scenarios and metrics from HELM's taxonomy does a vendor's reported result actually cover, and which categories, such as fairness or calibration, does it omit?
  • Was the comparison run under one shared protocol, or are the two numbers being compared drawn from different papers' own prompting choices?
  • Has the scenario or model list been refreshed since the run being cited, and does the cited table still match the live leaderboard?

HELM's contribution was procedural rather than a new capability test: it made the absence of a shared comparison visible and gave the field a taxonomy for naming what any single benchmark leaves out. Reading it that way is more useful than treating any one HELM column as a verdict.

Sources & reading trail

Holistic Evaluation of Language Models ↗

Establishes HELM's scope: 30 models, 42 scenarios, seven metrics, and the 17.9% to 96.0% core-scenario coverage figure.

Source published: 16 November 2022 · Retrieved: 16 September 2026

Holistic Evaluation of Language Models (HELM) ↗

Contemporaneous announcement stating HELM's three-pillar methodology and the coverage-gap figure in the authors' own words.

Source published: 17 November 2022 · Retrieved: 16 September 2026

HELM (Holistic Evaluation of Language Models) project site ↗

Living project page confirming HELM's continued operation as an ongoing benchmark platform, as retrieved 16 September 2026.

Source published: Not established · Retrieved: 16 September 2026

Papers and official documents establish the record; the reading and the questions are Model Field Guide editorial analysis. This retrospective draft does not imply the site published on the event date.