RETROSPECTIVE RECORD · PREPARED 16 SEPTEMBER 2026The record · 100 retrospective records ↗

The record / Evaluation

Evaluation / From the record · 20 November 2023 event · prepared 16 September 2026

GPQA questions resist a quick web search

The 2023 GPQA paper found PhD-level experts scored far higher than skilled non-experts with unrestricted internet access.

Visual for this record: GPQA questions resist a quick web search
Visual published by lh5.googleusercontent.com, shown for identification of the record. Credit: lh5.googleusercontent.com · source page ↗ Rights: owner-review-pending.

448 questions written by domain experts

The GPQA paper, submitted to arXiv on 20 November 2023, describes a set of 448 multiple-choice questions in biology, physics and chemistry, written and checked by contributors with a PhD or equivalent training in the relevant field. The project repository distributes the dataset, protected against casual scraping, together with baseline evaluation code for prompting strategies including zero-shot, chain-of-thought and retrieval-augmented approaches. The stated goal is a benchmark hard enough, and specific enough, that looking a question up rather than reasoning through it does not reliably work.

What the paper measured: expert, non-expert and model accuracy

The paper reports three comparison points under its own test conditions. Domain experts with PhD-level or PhD-track credentials in the question's own field reached 65 percent accuracy, rising to 74 percent once questions the experts themselves flagged as containing errors were excluded. Highly skilled non-expert validators, given unrestricted web access and spending on average more than thirty minutes per question, reached only 34 percent, the paper's evidence for calling the questions Google-proof: general research skill and search access do not reliably substitute for domain expertise on these items. GPT-4, evaluated under the paper's conditions, reached 39 percent, ahead of the non-expert baseline but well short of the expert one, and notably below the older, web-searchable format that the MMLU paper established for multiple-choice evaluation.

What Diamond means and what the numbers do not establish

The authors also released a smaller, higher-confidence subset commonly called GPQA Diamond, drawn from the questions where expert review gave the most confidence in the correct answer; this piece does not state its exact size, since that detail was not confirmed in the documents opened for it. What the main-set numbers do establish is a difficulty gap between expert and non-expert humans, and a 2023-era model score below both. They do not establish that the gap persists at every later model generation, since GPQA is a fixed question set that has been public since 2023, which creates the same contamination risk documented for other public benchmarks: a later high score could reflect exposure to the questions rather than improved scientific reasoning.

Questions to carry into your own evaluation

  • Is a cited score from the full main set or the Diamond subset, and does the source specify which?
  • How does the model's score compare not just to chance or to other models, but to the paper's own expert and non-expert human baselines?
  • Given how long GPQA has been public, could the model being evaluated have seen these exact questions during training?

GPQA's contribution was pairing a hard science benchmark with a measured human baseline at two skill levels, which lets a reader ask not just whether the model scored higher, but higher than which humans, doing what.

Sources & reading trail

GPQA: A Graduate-Level Google-Proof Q&A Benchmark ↗

Establishes the 448-question dataset and reports the 65%/74% expert, 34% non-expert and 39% GPT-4 accuracy figures.

Source published: 20 November 2023 · Retrieved: 16 September 2026

idavidrein/gpqa ↗

Documents dataset access, protection against scraping, and baseline evaluation code for multiple prompting strategies.

Source published: Not established · Retrieved: 16 September 2026

Measuring Massive Multitask Language Understanding ↗

Provides the contrasting, web-searchable multiple-choice design that GPQA's Google-proof construction was built to respond to.

Source published: 7 September 2020 · Retrieved: 16 September 2026

Papers and official documents establish the record; the reading and the questions are Model Field Guide editorial analysis. This retrospective draft does not imply the site published on the event date.