---
title: "Disagreement — why we refuse to blend evaluators — Undominated"
canonical: https://undominated.ai/disagreement/
---

# Disagreement — why we refuse to blend evaluators — Undominated

# Disagreement

Evaluators are not interchangeable. We publish one licensed quality axis and we do not blend a second.

As of 2026-09-19, 93% of adjacent LMArena pairs on this catalogue — 127 of 136 — fall inside the models' own published error bands. That is the uncertainty we can publish. We do not blend a second lab's figures into this rank, and we do not invent a crisp first place.

Condition: standard-delivery rows with a measured CI half-width. LMArena data is CC BY 4.0.

## Why we refuse to blend evaluators

We never blend evaluators into one quality number. Different labs measure different things. A composite would invent a precision neither source claims, and it would hide the models they would rank differently. The disagreement is the signal. A blend erases it.

## A second lab we do not sublicence

A second evaluation lab exists. We do not sublicence those figures, so they do not appear here — not as scores, not as ranks, not as a correlation. Internal use is not a public surface.

## Uncertainty we can publish

On the LMArena board, adjacent ranks that sit inside published error bands are ties. The share in the answer is recomputed from this catalogue: standard-delivery rows with a measured CI half-width, CC BY 4.0. Details live on Significance.

## Unrated is not zero

196 of 333 standard, publishable models have no LMArena score. Unrated is not zero. Absence is not a favourable verdict. Those models stay unrated until a licensed measurement exists.

## See also

 - Significance bands
- Independence
- Methodology
- Corrections
