Disagreement
Evaluators are not interchangeable. We publish one licensed quality axis and we do not blend a second.
Condition: standard-delivery rows with a measured CI half-width. LMArena data is CC BY 4.0.
Why we refuse to blend evaluators
We never blend evaluators into one quality number. Different labs measure different things. A composite would invent a precision neither source claims, and it would hide the models they would rank differently. The disagreement is the signal. A blend erases it.
A second lab we do not sublicence
A second evaluation lab exists. We do not sublicence those figures, so they do not appear here — not as scores, not as ranks, not as a correlation. Internal use is not a public surface.
Uncertainty we can publish
On the LMArena board, adjacent ranks that sit inside published error bands are ties. The share in the answer is recomputed from this catalogue: standard-delivery rows with a measured CI half-width, CC BY 4.0. Details live on Significance.
Unrated is not zero
196 of 333 standard, publishable models have no LMArena score. Unrated is not zero. Absence is not a favourable verdict. Those models stay unrated until a licensed measurement exists.