Glossary

The terms the leaderboard and the methodology lean on, each defined once. Where the methodology or the leaderboard already words a term, the definition keeps those words. No definition holds a figure: each links to the page that computes it.

Blended price

One price per million tokens made from a model’s separate rates, weighted by the workload’s output share and cache-hit rate. When a cache-read rate is unlisted, blended estimates use the selected tier’s full input rate.

Methodology: Price is a property of a model and a workload

Deal

A cheaper offer that the reference rule excludes, shown on the model page as “Cheaper right now” with its reason. It never sets a rank, the frontier, a verdict or a headline figure.

Methodology: reference-price

Deal-only

The label on a model no offer passes the reference rule for: the model keeps a fallback price, labelled deal-only with its reason, because an unpriced model would read as absent.

Methodology: reference-price

Document axis

The second quality lens: LMArena’s document board, scored apart from the general one. The arena boards are separate scales and are never blended.

Methodology: Licensing, and what that constrains

Dominated

Another model scores at least as high, costs no more, and wins on one of the two, though not always a like-for-like replacement. It may accept no images, or hold less context, where the model it beats does.

Check a model

Effective price

A model's cost per million tokens for one workload, at the context tier the prompt reaches, with separately billed reasoning charged where it is dearer than output. There is no such thing as "the" price of a model.

Methodology: Price is a property of a model and a workload

Interval

A score's published confidence bounds, shown as the score plus or minus its half-width, at the confidence level the linked page states. LMArena's dataset carries the confidence bounds, and the leaderboard uses them as published.

Score significance

LMArena

The source of every general score here: LMArena Elo from human preference, on its general board and its document board, taken from LMArena's official dataset under its open licence.

Methodology: Licensing, and what that constrains

Open weights

The catalogue’s hand-researched flag that a model’s weights are available. Licence and commercial-use terms still vary, so it does not mean open source, and a model with no recorded status is unknown, never closed.

Open vs closed

Overpay

What a dominated model costs as a multiple of the cheapest such alternative: its price divided by the price of the cheapest model that dominates it, at the balanced workload. The home page gives the median.

Leaderboard

Reference price

The price this site publishes for a model: the cheapest offer, on the balanced blend, that passes every test below. That offer’s input, output and cache-read rates are published together, never mixed with another offer’s.

Methodology: reference-price

Retirement

A provider’s announcement that a model will stop being served. A date shows when service is expected to end, and a retired model leaves the leaderboard.

Retirement tracking

Service tier

Catalogue entries described by their own vendor as another catalogue model served differently: a faster lane, a different reasoning mode, more compute per query. They are counted only on the seller’s own words, never on a naming convention.

Methodology: When the same model is sold twice

Significance rank

The rank the general leaderboard prints, not a row number: one more than the number of models that beat this one by more than the measurement can resolve. Models share a rank when the same number of models clearly beat each of them.

Methodology: Ranks you can rely on

Standard rate

An offer on promotion counts at its standard rate: each rate divided by one minus the discount. The promoted price is disclosed, never ranked.

Methodology: reference-price

Standard window

The hours in which a time-windowed model bills its ordinary rate. Headline is the standard rate, and a discount window is shown beside it, never as the headline.

Methodology: Price is a property of a model and a workload

Task board

A leaderboard for one kind of work, such as websites or game development, from Design Arena's per-task Elo. Selecting a task re-ranks the entire board on that task's Elo rather than on a general index.

Rankings by task

Tier ladder

The rates a model bills as the prompt grows: some models change rate past a prompt-length threshold, some of them more than once. Every rung is kept, and a price takes the highest rung the prompt actually reaches.

Context price cliffs

Time window

Hours, stated in UTC, in which a seller bills a model at a different rate. Check the window, timezone and rates; do not assume a fixed discount.

Methodology: Price is a property of a model and a workload

Undominated

No model on this board is at least as good for less, or better for the same price. The undominated models together are the value frontier.

Frontier

Unrated

No independent score on the board in view. Unrated models are listed by price and excluded from quality ordering, never sorted to the bottom as though they had been measured and failed.

Methodology: Where the data comes from

Value frontier

A model is on the value frontier when no other model scores at least as high, costs no more, and wins on one of the two. It is drawn among rated, priced, standard-delivery models, at the workload in view.

Frontier

Variant

A delivery or pricing variant of a model already listed, such as its batch or free tier. Its conditions can differ from the standard option.

Batch pricing

Weak Pareto dominance

The dominance rule this site uses: an alternative beats a model when it is higher-scoring or cheaper, with neither dimension worse. It need not be both, so matching the score for less money is enough.

Frontier

Within error

Two scores are within error when they differ by no more than their two published half-widths together: the benchmark can't separate them. The order between them is not a finding.

Score significance

Workload

The shape of the work being priced: how much of it is output, and how much of the input is read from cache. Rather than pick one blend and hide it, the input:output ratio and cache-hit rate are controls you can see and change.

Methodology: Price is a property of a model and a workload

Evidence & Ask