---
title: "Ranges · Undominated.ai"
canonical: https://undominated.ai/hi/ranges/
description: "The leaderboard gives you a position. The benchmark also publishes an interval around it. At rank 40, that interval covers 24 different models — and they cost between $0.077 and $50 per million tokens."
---

# Ranges · Undominated.ai

> The leaderboard gives you a position. The benchmark also publishes an interval around it. At rank 40, that interval covers 24 different models — and they cost between $0.077 and $50 per million tokens.

# Ranges

The leaderboard gives you a position. The benchmark also publishes an interval around it. At rank 40, that interval covers **24** different models — and they cost between **$0.077** and **$50** per million tokens.

A board that publishes a confidence interval for every model also publishes **402** distinct integer positions for 402 models, with **no ties**.

Across the whole board, **399** of 401 neighbouring intervals overlap. That is what the published uncertainty looks like when you draw it instead of sorting by it.

 Each bar is one model. Its length is the span of rank positions the published interval does not rule out; the faint diagonal behind it is the single position the board displays. Colour is price, not quality.

## Stand in a position

Pick a rank. These are the models whose published interval covers it.

 Rank 40

**24** models are not ruled out of rank 40. The cheapest bills **$0.077** per million output tokens and the dearest **$50** — a spread of **647×** at one position.

| Model | Not ruled out of | $/M out |
| --- | --- | --- |
| Qwen3.6 Max Preview | 23–47 | $6.16 |
| GLM 5 | 23–43 | $1.92 |
| GPT-5.6 Terra | 23–44 | $12 |
| Kimi K2.5 | 26–41 | $2.25 |
| DeepSeek V4 Pro 0423 | 23–48 | $0.845 |
| GPT-6 Astra | 21–53 | $50 |
| Claude Sonnet 5 | 28–48 | $10 |
| Gemma 4 31B | 26–52 | $0.34 |
| Hy3 | 27–52 | $0.33 |
| GLM 4.6 | 29–48 | $1.75 |
| Inkling | 29–51 | $4.05 |
| Qwen3.8 27B | 28–53 | $2.55 |
| Claude Sonnet 4.5 | 32–50 | $15 |
| Qwen3.5 397B A17B | 31–51 | $3.5 |
| Qwen3.6 Plus | 32–53 | $1.95 |
| GLM 5V Turbo | 29–54 | $4 |
| GLM 4.7 | 30–54 | $1.75 |
| Gemini 3.5 Flash Lite | 32–54 | $2.5 |
| Gemma 4 26B A4B | 30–58 | $0.3 |
| MiniMax M3 | 33–54 | $1.2 |
| DeepSeek V4 Flash 0423 | 36–56 | $0.077 |
| GLM 4.5 | 37–61 | $2.2 |
| GPT-5.6 Luna | 39–61 | $1.2 |
| Grok 4.6 | 36–61 | $6 |

The widest position on this board is rank 56: 22 models, spanning 971× in price.

These bands are wider than the truth, on purpose. They are built from each model’s own published interval and nothing else, which ignores that the scores are jointly estimated from shared battles — a covariance the published intervals cannot recover. So a band shows what is *not ruled out*, never what is identified, and two models sharing a position are not thereby tied or equivalent.

This is not a claim that the ranking is wrong. A rank column is a deterministic display of a sort over point estimates, and the benchmark never claimed statistical separation. It is also the only board in this market that publishes an interval at all, which is the only reason this page can exist.

196 models in the catalogue carry no entry on this board. They are absent from the figure and named as absent — never drawn at rank zero, and never counted as scoring badly. Absence of a measurement is not a measurement.

## Try it yourself

The blind rank test shows two models with their names hidden and asks you to pick the higher-scoring one — then tells you whether the published intervals can tell them apart at all.

[Blind rank test](/blind-rank/)

Scores and intervals are LMArena Elo, used under CC BY 4.0 from the official dataset. Prices are list output price per million tokens, as published by each provider.
