
Four named sizes, no stated parameter counts
Google's announcement, dated 10 May 2023, introduces PaLM 2 as having 'improved multilingual, reasoning and coding capabilities' and names four size tiers, Gecko, Otter, Bison and Unicorn, without quantifying any of them in parameters. The technical report, submitted a week later, states it discusses 'pre-trained models (of various sizes)' and warns readers directly: it explains that 'user-facing products' built on PaLM 2 'typically include additional pre- and post-processing steps' and 'the underlying models may evolve over time', so 'one should not expect the performance of user-facing products to exactly match the results reported in this report'. That sentence is itself a disclosed limit on how the report's own numbers should be used.
What the scaling-law section actually measured
The report's clearest technical contribution is a scaling-law study, run independently of the production PaLM 2 models. Google trained a separate set of models 'from 400M to 15B' parameters across four compute budgets to find the optimal ratio between model size and training tokens, and reports arriving at a conclusion 'strikingly similar' to the Chinchilla study: model size and training data should grow in roughly equal proportion. The report is explicit that these scaling-law models are a separate experiment: it states 'these models were used only for the scaling law study, and do not reflect the model sizes and FLOPs used in PaLM 2 models'. In other words, the one place the paper discloses concrete parameter counts is not the model Google shipped.
What the multilingual and reasoning claims rest on
The report's abstract claims 'significantly improved quality on downstream tasks across different model sizes' and 'large improvements over PaLM on BIG-Bench and other reasoning tasks', trained on data spanning what Google's blog post describes as 'over 100 languages'. These are results on the report's own selected benchmark suite, comparing PaLM 2 to its named predecessor PaLM; the withheld parameter counts mean a reader cannot express the improvement per parameter or attribute it to scale versus training-data changes. The report is a model of a pattern that recurred across several 2023 releases: substantial disclosure of methodology and benchmark results, paired with non-disclosure of the basic size figures that would let those results be placed in context.
Questions to carry into your own evaluation
- Is a cited PaLM 2 result drawn from the technical report itself, or from a downstream product the report warns may not match it?
- Does the report's own scaling-law study apply to the production model, or only to the separate smaller models trained to study the scaling relationship?
- When a size tier like 'Unicorn' or 'Bison' is named without parameters, what is actually being compared against a competitor's model?
PaLM 2's report advanced a genuine scaling-law finding while withholding the figures needed to apply that finding to the model it was ostensibly about. Reading the disclosed methodology as if it described the shipped model conflates two different sets of experiments that the report itself keeps separate.
Sources & reading trail
The report's scaling-law experiment, its statement that the scaling-law models do not reflect PaLM 2's actual sizes, and its caveat about user-facing products diverging from reported results.
Source published: 17 May 2023 · Retrieved: 16 September 2026
The announcement's four named size tiers, multilingual training claim, and absence of stated parameter counts.
Source published: 10 May 2023 · Retrieved: 16 September 2026
Papers and official documents establish the record; the reading and the questions are Model Field Guide editorial analysis. This retrospective draft does not imply the site published on the event date.