Methodology

Everything here is reproducible from the same public sources you can check yourself. Where the data is thin or the sources disagree, this page says so — a trust page that reports only good news is marketing.

LMArena runs a public leaderboard. This site is not that leaderboard. We cite LMArena Elo under CC BY 4.0 beside a price. Unrated stays unrated. Artificial Analysis figures are not published here.

60% of the catalogue has no independent quality score. 145 of 364 models are rated. The rest are listed, priced and clearly marked unrated — never ranked.
72/145 rendered positions that are genuinely distinct ranks. 94% of adjacent pairs are inside the benchmark's margin of error.
67/364 models whose headline price is not the whole price — context tiers, separately-billed reasoning, or time-of-day rates.
7/447 entries with variable prices withheld
4/447 entries with suspect latency withheld; may overlap the price group

1 · What "best" means

Quality leads and price breaks ties. That ordering is the whole product: a cheapest-first list is trivially easy to publish and nearly useless, because the cheapest model in this market is 12,500× cheaper than the dearest and generally cannot do the job.

Capability comes from independent evaluators, never from us and never from the vendor. The general scores are LMArena Elo from human preference, on its general board and its document board, used under CC BY 4.0 from the official dataset. The per-task scores under a task lens are Design Arena’s, credited below.

Why there is no single quality score

Capability comes from independent evaluators, never from us and never from the vendor. The general scores are LMArena Elo from human preference, on its general board and its document board, used under CC BY 4.0 from the official dataset. The per-task scores under a task lens are Design Arena’s, credited below.

2 · Ranks you can rely on

The rank column is a significance rank, not a row number. A model’s rank is one more than the number of models that beat it by more than both margins of error combined. Models with the same count share a rank. Two models within error of each other can still hold different ranks, because each is compared with the whole board.

SHA-256 of the rank-relevant projection (slugs, scores, workload prices, frontier flag): 06cd12db54e2968c3de49b45a3a2a6881f374f6e58a462437784450842efb989. Dates and build stamps are excluded. A change here is a change in the ranking, not in the chrome.

Corrections

Most of a leaderboard's ordering is noise

Two arena scores count as separated only when they differ by more than the sum of the two models’ 95% half-widths. That sum averages 10.47 points on this board, and the test is at least as strict as a 95% interval on the difference. 136 of 144 adjacent pairs (94%) fall inside their own sum. Counting each model’s rank as one more than the number of models that clearly beat it, the 145 rendered positions hold 72 distinct ranks, and 4 models are joint first.

"The #1 model" is therefore not a claim this data supports. Ties are shown as ties. Rank 1 is rank 1 among the models in the price feed this site reads: where LMArena ranks a model above all of them that the feed does not carry, the board names that model and does not rank it.

This method is not ours. LMArena has ranked this way for years — their Rank (UB) gives each model one more than the number of models statistically better than it, and they have open-sourced the implementation. What is unusual is applying it to a price comparison, where the incentive is to publish a crisp ordering.

Applied to the general arena board only. Other boards carry their own error, which we do not have, so they are not given a borrowed threshold.

3 · "Better and cheaper" has to mean it

Pareto dominance on price and capability is a two-dimensional projection of a much wider purchase decision, and the projection lies more often than not. A model can score higher and cost less while accepting no images, capping output at half the tokens, or holding half a million fewer tokens of context.

So "both better and cheaper" is reserved for a replacement whose capability envelope is a superset of the original. Of 133 dominance verdicts, 113 are strict and 20 (15%) name what you would give up instead. Without that gate, 61% of raw two-axis verdicts recommended a model that could not do the incumbent's job.

4 · The frontier, not a value score

A model is on the value frontier when no other model scores at least as high, costs no more, and wins on one of the two. 12 of the rated, priced, standard-delivery models qualify. Each of the other 92% has such a model above it, though not always a like-for-like replacement: the section above counts the ones that give something up.

Why not rank by quality per dollar?

Because the ratio rewards cheapness far more steeply than quality. Applied to this catalogue it crowns models nobody should ship — the winner and the runner-up both sit far below the leaders on capability and win purely on being cheap. Any site publishing a "value" column computed that way is giving bad advice. We use the frontier plus an explicit quality floor instead.

5 · Price is a property of a model and a workload

There is no such thing as "the" price of a model. Summarisation is input-heavy, code generation is output-heavy, and an agent loop replays its context every turn; the same catalogue orders differently under each. Rather than pick one blend and hide it, the input:output ratio and cache-hit rate are controls you can see and change.

When a cache-read rate is unlisted, blended estimates use the selected tier’s full input rate.

WorkloadOutput shareCache hitsShape
Balanced 25% 0%3 tokens in per 1 out
Summarise 5% 30%Huge input, tiny output
Chat 40% 50%Long cached system prompt
Code gen 60% 20%Output dominates
Agentic 15% 70%Context replayed each turn

Three things routinely make the advertised rate wrong, and all are surfaced on the row:

  • Context tiers. 51 models change rate past a prompt-length threshold, some of them more than once — Qwen3.7 Flash goes 3.33× past 32K and 6.67× past 256K. Price a long-context workload off the headline and you are out by a multiple.
  • Reasoning tokens billed apart from output. When reasoning tokens are charged separately, cost depends on their billed volume and rate. Include that charge in your estimate.
  • Time-of-day pricing. Models with published price windows: 2. Check the window, timezone and rates; do not assume a fixed discount.

6 · Which price we publish

A score is measured on the model’s own deployment, so the price beside it is the reference price: the cheapest offer, on the balanced blend, that passes every test below. That offer’s input, output and cache-read rates are published together, never mixed with another offer’s.

  1. Serving. An offer OpenRouter lists as not serving is excluded. An offer that reads as degraded at the moment of the read still counts when it served at least 95% of the last day. An offer that served less than 80% of the last day does not count, whatever it reads now.
  2. Standard delivery. Flex, priority, fast and similar tiers are a different product, and are excluded.
  3. Standard rate. An offer on promotion counts at its standard rate: each rate divided by one minus the discount. The promoted price is disclosed, never ranked.
  4. Standard window. A time-of-day price counts at its standard rate; the discount hours are disclosed.
  5. The maker’s precision or better. A 4-bit offer (fp4, nvfp4, mxfp4, int4) counts only when the maker serves 4-bit itself or, where the maker declares no precision, its published weights are 4-bit. An offer that declares no precision counts when it is the maker’s own, or a reseller’s for a model with closed weights: it serves the maker’s deployment. For open weights, or weights nobody has recorded, a reseller’s offer that declares none is excluded: it is never read as full precision.

When no offer passes, the model keeps a fallback price, labelled deal-only with its reason, because an unpriced model would read as absent: among the serving offers of standard delivery, each at its standard rate, the cheapest from a reseller that declares no precision and, failing that, the cheapest 4-bit one. Another delivery tier is used only when nothing else serves. A model with no offer to take a price from keeps OpenRouter’s model-level price.

The cheapest offer the rule excludes, when it is cheaper, is shown on the model page as “Cheaper right now”, with the reason: a promotion and its percentage, 4-bit, a delivery tier, or precision not disclosed. It never sets a rank, the frontier, a verdict or a headline figure. Where the offer that sets the price is itself on promotion, its own promoted price is shown.

On this data 333 models carry a reference price, 15 a deal-only price and 99 the model-level price, batch and free listings among them. 86 show a cheaper offer right now, 11 of them from a reseller that declares no precision.

Before this rule every price here was OpenRouter’s model-level price, which any seller can set, promotions and 4-bit copies included. The change is logged as §96 of the correction log, with its before-and-after figures.

7 · Task-conditional ranking

A single global ordering is intellectually dishonest, because different models win different work. We carry per-task Elo across 25 categories — front-end, data visualisation, game development, SVG, Android, slides and more. Selecting a task re-ranks the entire board on that task's Elo rather than on a general index.

Task scores are Design Arena’s per-task Elo, read from OpenRouter’s model feed: designarena.ai.

Coverage varies by category, and a model absent from a category's leaderboard is shown as unmeasured for that task, not as bad at it.

8 · When the same model is sold twice

Some models are not new models. The seller says so: 11 entries in this catalogue are described by their own vendor as another catalogue model served differently — a faster lane, a different reasoning mode, more compute per query. 10 of the 11 have no entry on any LMArena board. 11 of their 11 base models do.

When the same model is sold twice
Sold asPrice of the tierBase modelWhat the seller says
o1-pro no entry$600 ×10o1 1,365.8 more compute per query
o3 Pro no entry$80 ×10o3 1,409.9 more compute per query
GPT-6 Astra Pro no entry$50 ×1GPT-6 Astra 1,442.3 at max effortsame model, different serving mode
GPT-5.6 Sol Pro no entry$20 ×1GPT-5.6 Sol 1,456.3 at xhigh effortsame model, different serving mode
GPT-5.6 Terra Pro no entry$12 ×1GPT-5.6 Terra 1,446.3 at xhigh effortsame model, different serving mode
GPT-6 Sol Pro no entry$10 ×1GPT-6 Sol 1,395.4 at max effortsame model, different serving mode
GPT-6.1 Sol Pro no entry$10 ×1GPT-6.1 Sol 1,445.6 at max effortsame model, different serving mode
o3 Mini High 1,336.6 $4.4 ×1o3 Mini 1,319.4 same model, different serving mode
o4 Mini High no entry$4.4 ×1o4 Mini 1,353.2 same model, different serving mode
GPT-5.6 Luna Pro no entry$1.2 ×1GPT-5.6 Luna 1,431 at xhigh effortsame model, different serving mode
GPT-6 Luna Pro no entry$0.5 ×1GPT-6 Luna 1,391.5 at max effortsame model, different serving mode

The distinction matters for what you can conclude. Where the seller says identical capabilities — 0 of the 11 — the base model's score is the seller's own claim about the tier too, and the only thing you are buying is speed. Where the seller says more compute or a different mode, no published measurement tells you what the extra spend buys, with the single exception noted above.

This counts only what the seller states. A model whose name ends in "Pro" but which its vendor calls a flagship model in its own right is not counted here, because that would be our claim rather than theirs — the same mistake as treating a vendor's own service tiers as competing sellers, which once turned a 2× price spread into a published 7.2×.

No entry on an LMArena board is not a verdict. It is not evidence of poor quality, and it does not mean nobody has evaluated these — vendors publish their own results. It means the axis this site ranks on has no reading for them, so we do not rank them.

9 · Where the data comes from

Prices are normalised from a machine-readable cross-provider catalogue and cross-checked against vendor pricing pages. Every model page links its source and fetch date.

  • Conflicts are recorded, not averaged. When two sources disagree on a price, that disagreement is the finding. A model listed at $4/$20 by its vendor and sold at $2/$10 through an aggregator has two true prices, and which one applies depends on where you buy.
  • Implausible values are quarantined. An upstream sentinel of −1 for "priced at request time" becomes −$750,000 per million if passed through naively, and wins every cheapest-first sort. Anything outside a plausible range is withheld with a reason attached.
  • Confidence is per field. Not one badge per row. Claiming a record is verified when only its price was checked is the kind of overstatement that ends trust the first time somebody notices.
  • Unrated is not zero. 219 models have no independent score. They are listed by price and excluded from quality ordering, never sorted to the bottom as though they had been measured and failed.

10 · Licensing, and what that constrains

Two of the benchmarks this site relies on have very different redistribution terms, and it changes what can be published.

  • LMArena publishes its leaderboard as a CC-BY-4.0 dataset, which permits redistribution with attribution — and carries the confidence bounds the significance ranks above depend on. That is the channel to use; scraping the site itself is prohibited by their terms.
  • Artificial Analysis's free tier is "internal use only with attribution"; public redistribution requires a commercial licence, and their site terms separately limit use to personal and non-commercial purposes. So we do not publish their figures at all. Every general capability score on this site is LMArena Elo, used under CC BY 4.0 from the official dataset; the per-task scores are Design Arena’s, below. Artificial Analysis remains an internal cross-check, which is what its free tier permits; it reaches no page, export or API here.
  • Design Arena’s API terms make its leaderboard data free to use, commercially included, on one condition: a public display credits Design Arena and links to designarena.ai. Its per-task Elo reaches this site through OpenRouter’s model feed; the task lens, each model page’s rankings by task and this page credit it.
  • Pricing itself is a fact a vendor publishes, which is a different thing from a compiled index score. Prices are cross-checked against vendor pages and every row links its source.

Questions this page answers

Is this the LMArena leaderboard?

No. LMArena runs a public leaderboard. This site is not that leaderboard. We cite LMArena Elo under CC BY 4.0 beside a price. Unrated stays unrated.

Does this site publish Artificial Analysis scores?

No. Artificial Analysis figures are not published here. Language-model preference scores come from Arena. The embeddings instrument uses its separately named MTEB task collection; the benchmark scopes are never combined.

Cite this claim

Plain text and BibTeX
Plain
60% of 364 catalogue models have no independent quality score on the LMArena lens this site ranks on. Unrated is not scored zero. Rank resolution: 72 distinct ranks across 145 rated positions. As of 2026-10-06. Lens: LMArena Elo (CC BY 4.0). Ranking source digest: 06cd12db54e2968c. https://undominated.ai/methodology/
Publication: sha256:ecef83f767cd4d2c21766f6a5b4dffa36c06351f23657542d473bc1a030c0932; calculation SHA-256: 5ae4d0bee9b70298dd437b05df87c32acb407040b50bf88240677f295f25955d; permitted inputs: https://undominated.ai/data/publications/manifests/ecef83f767cd4d2c21766f6a5b4dffa36c06351f23657542d473bc1a030c0932.json
BibTeX
@misc{undominated-methodology-2026-10-06,
  author = {{Undominated.ai}},
  title  = {Methodology integrity — Undominated.ai},
  year   = {2026},
  url    = {https://undominated.ai/methodology/},
  note   = {as of 2026-10-06; LMArena Elo (CC BY 4.0); ranking source 06cd12db54e2968c},
  annote = {Publication: sha256:ecef83f767cd4d2c21766f6a5b4dffa36c06351f23657542d473bc1a030c0932; calculation SHA-256: 5ae4d0bee9b70298dd437b05df87c32acb407040b50bf88240677f295f25955d; permitted inputs: https://undominated.ai/data/publications/manifests/ecef83f767cd4d2c21766f6a5b4dffa36c06351f23657542d473bc1a030c0932.json}
}

LMArena scores used under CC BY 4.0 · numbers recomputed from the catalogue on this page. Input identity & calculation SHA.

11 · Corrections

Prices change without notice and pages get reformatted; some of what is here will be wrong at any given moment. If you find an error, the fastest fix is to say which model, what the figure should be, and where it is published — corrections with a primary source get applied directly.

Report a wrong price

Catalogue last fetched 2026-10-06. Ranking logic is versioned with the site, so a change in how "best" is computed is a visible diff rather than a silent retune.

Evidence & Ask