
A launch-day arena score attached to a name, not the shipped model
Meta's 5 April 2025 announcement, The Llama 4 herd: The beginning of a new era of natively multimodal AI innovation, introduced three models — Scout, Maverick, and a still-training Behemoth held back from release — and stated that Llama 4 Maverick offers 'a best-in-class performance to cost ratio with an experimental chat version scoring ELO of 1417 on LMArena.' The sentence names two things in the same breath: an arena leaderboard score, and the fact that the version tested to earn it was labelled 'experimental,' distinct from whatever weights a developer could actually download from Meta that day.
What the leaderboard operator's own record shows
LMArena's public model list, maintained on GitHub, names the tested variant explicitly as 'llama-4-maverick-03-26-experimental' and records it among models deprecated from the Text Arena battle mode and leaderboard as of 27 June 2025. It sits on that list alongside similarly tagged experimental or preview checkpoints from other vendors — an early Grok 3 build, several dated 'chatgpt-4o-latest' snapshots, and experimental Gemini builds — which is evidence that submitting a distinctly labelled, non-final chat tune for arena testing was a live practice across the field at the time, not an isolated incident invented for Llama 4. The record does not, on its own, state why any particular entry was removed, only that each aged out of the live leaderboard by the stated date.
Why the naming detail changes what the score means
An Elo-style arena score is a measure of one specific checkpoint's performance in head-to-head human comparisons, not a property of a model family name. If the checkpoint submitted for testing differs from the checkpoint later shipped in an API or a download, the published score describes a model a user cannot obtain, and a separate, possibly different score would be needed to describe the one they can. Meta's own announcement discloses the 'experimental' label; it does not state how the released Maverick's arena performance compares with the tested variant's, which is the specific figure a buyer relying on the headline number would actually need.
- Is the arena score being cited attached to the exact model identifier available for download or API access today?
- Has the tested variant since been deprecated or replaced on the leaderboard, and if so, was a score for the shipped model ever separately reported?
- Does the leaderboard operator's own documentation distinguish vendor-submitted experimental tunes from generally available releases at the time the score was recorded?
A leaderboard score is only as informative as the match between the tested artefact and the one a reader can use. Where a vendor names the tested version as experimental in its own announcement, that is the cue to ask for the shipped model's number before repeating the headline figure.
Sources & reading trail
States the 1417 LMArena Elo score and describes the tested chat version as experimental.
Source published: 5 April 2025 · Retrieved: 16 September 2026
Names 'llama-4-maverick-03-26-experimental' as the tested identifier and records its removal from the live leaderboard as of 27 June 2025.
Source published: Not established · Retrieved: 16 September 2026
Papers and official documents establish the record; the reading and the questions are Model Field Guide editorial analysis. This retrospective draft does not imply the site published on the event date.