
A three-size family and a leading exam score
Google's announcement, dated 6 December 2023, introduces Gemini 1.0 in three sizes, Ultra, Pro and Nano, calling it 'our most capable and general model yet'. The headline figure is a '90.0%' score on the MMLU exam benchmark, described as making Gemini Ultra 'the first model to outperform human experts' on that test. The technical report, submitted 19 December 2023, gives the precise figure as '90.04%' and states the benchmark's own human-expert baseline is '89.8%', with the prior state of the art at '86.4%'.
How the reported score was produced
The report is specific about the decoding method behind the 90.04% figure: Gemini Ultra reaches that score 'when used in combination with a chain-of-thought prompting approach that accounts for model uncertainty', generating 'a chain of thought with k samples, for example 8 or 32' and taking the consensus answer if agreement clears a preset threshold, otherwise falling back to a single greedy sample. This is a named, non-default prompting strategy, not the score of a single pass over each question, and the report says a plain chain-of-thought or plain greedy-sampling score is reported separately in its appendix. A benchmark figure produced by sampling and voting across dozens of generations is a different measurement than one produced by a single response, and the two are not interchangeable when comparing to another model's reported number.
What Google's own video describes about itself
Google published a demonstration video the same day, titled 'The capabilities of multimodal AI', showing Gemini responding to spoken and visual prompts. The video's own YouTube description states: 'For the purposes of this demo, latency has been reduced and Gemini outputs have been shortened for brevity.' That is Google's own disclosure, not an outside allegation, and it means the demonstration does not represent unedited, real-time interaction with the model at the timing shown. It does not by itself say anything about the accuracy of the responses shown, only that the pace and length of the exchange were altered for the video.
Questions to carry into your own evaluation
- Is a cited Gemini benchmark score the one produced by simple prompting, or by a sampling-and-consensus method like the one behind the 90.04% MMLU figure?
- Does a product demonstration video disclose whether timing or output length was edited, and does that change what the video can be used to claim?
- When two models are compared on the same benchmark, were both scored under the same decoding strategy and number of samples?
Gemini 1.0's launch combined a genuine benchmark result, achieved under a specific and disclosed decoding method, with a demonstration video that Google itself says was edited for pacing. Neither fact invalidates the other, but treating either the exam score or the video as a plain, unqualified capability demonstration overstates what the source documents themselves say.
Sources & reading trail
States the precise 90.04% MMLU score, the 89.8% human-expert baseline, and the chain-of-thought-with-consensus decoding method used to reach it.
Source published: 19 December 2023 · Retrieved: 16 September 2026
The launch announcement's three model sizes and headline 90.0% MMLU claim.
Source published: 6 December 2023 · Retrieved: 16 September 2026
Google's own video description stating latency was reduced and outputs shortened for the demonstration.
Source published: 6 December 2023 · Retrieved: 16 September 2026
Papers and official documents establish the record; the reading and the questions are Model Field Guide editorial analysis. This retrospective draft does not imply the site published on the event date.