Disagreement

Evaluators are not interchangeable. We publish one licensed quality axis and we do not blend a second.

Condition: standard-delivery rows with a measured CI half-width. LMArena data is CC BY 4.0.

Why we refuse to blend evaluators

We never blend evaluators into one quality number. Different labs measure different things. A composite would invent a precision neither source claims, and it would hide the models they would rank differently. The disagreement is the signal. A blend erases it.

A second lab we do not sublicence

A second evaluation lab exists. We do not sublicence those figures, so they do not appear here — not as scores, not as ranks, not as a correlation. Internal use is not a public surface.

Uncertainty we can publish

On the LMArena board, adjacent ranks that sit inside published error bands are ties. The share in the answer is recomputed from this catalogue: standard-delivery rows with a measured CI half-width, CC BY 4.0. Details live on Significance.

Unrated is not zero

196 of 333 standard, publishable models have no LMArena score. Unrated is not zero. Absence is not a favourable verdict. Those models stay unrated until a licensed measurement exists.

See also