---
title: "Phương pháp — cách quyết định cái “tốt nhất”, và chỗ dữ liệu yếu · Undominated.ai"
canonical: https://undominated.ai/vi/methodology/
description: "Cách trang này xếp hạng 410 mô hình AI: biên giá trị, Elo theo tác vụ, giá theo tải công việc, và tường trình thẳng thắn về những gì dữ liệu không phủ."
---

# Phương pháp — cách quyết định cái “tốt nhất”, và chỗ dữ liệu yếu · Undominated.ai

> Cách trang này xếp hạng 410 mô hình AI: biên giá trị, Elo theo tác vụ, giá theo tải công việc, và tường trình thẳng thắn về những gì dữ liệu không phủ.

# Methodology

Everything here is reproducible from the same public sources you can check yourself. Where the data is thin or the sources disagree, this page says so — a trust page that reports only good news is marketing.

 63% of the catalogue has **no independent quality score**. 150 of 410 models are rated. The rest are listed, priced and clearly marked unrated — never ranked.
 r = 0.806 correlation between our two independent quality sources, over the 62 models both have scored. They largely agree — but not about every model.
 52/108 rendered positions that are **genuinely distinct ranks**. 93% of adjacent pairs are inside the benchmark's margin of error.
 79 models whose headline price is **not the whole price** — context tiers, separately-billed reasoning, or time-of-day rates.
 9 values withheld as implausible rather than displayed: 5 runtime-priced router products, 4 impossible latency figures.

## 1 · What "best" means

Quality leads and price breaks ties. That ordering is the whole product: a cheapest-first list is trivially easy to publish and nearly useless, because the cheapest model in this market is 15.000× cheaper than the dearest and generally cannot do the job.

Capability comes from independent evaluators, never from us and never from the vendor. Two sources are carried: **Artificial Analysis** indices for general, coding and agentic capability, and **LMArena** Elo from human preference.

### Why there is no single quality score

Across the 62 models both sources have measured, they correlate at **r = 0.806** — close agreement. That is not the argument for blending them; it is the argument for looking at where they part company. **4 models** rank very differently between the two, and a composite score would erase exactly those, which are the ones a buyer most needs flagged. So both axes are published side by side and never averaged.

An earlier version of this page reported r = 0.409 and argued the benchmarks broadly disagreed. That figure came from 15 models scraped from a leaderboard page, all of them frontier-class; restricting the range that narrowly flattens a correlation. Re-sourced from Arena's licensed dataset the sample is 62 models spanning the full range, and the correlation is 0.806. The earlier claim was wrong and the argument built on it was too.

| Model | Artificial Analysis | LMArena |
| --- | --- | --- |
| Gemini 2.5 Pro | 18th pct | 71th pct |
| GPT-5.6 Luna | 73th pct | 35th pct |
| Solar Pro 4 | 53th pct | 16th pct |
| Grok 4.6 | 94th pct | 58th pct |
| GPT-5.4 Nano | 45th pct | 11th pct |

Both percentiles are computed over the models scored by *both* sources. Ranking each model within its own source's population would compare a small, entirely frontier-class arena pool against the full-market index pool, and manufacture disagreements that are not real. That was our first implementation, and it was wrong.

## 2 · Ranks you can rely on

The rank column is a **significance rank**, not a row number. Two models the measurement cannot separate share a rank.

### Most of a leaderboard's ordering is noise

The 95% confidence interval on a difference between two Artificial Analysis scores is about ±1.41 index points. On this catalogue, **100 of 107 adjacent pairs (93%)** fall inside it. So 108 rendered positions carry only **52 genuinely distinct ranks**, and 2 models are joint first.

"The #1 model" is therefore not a claim this data supports. Ties are shown as ties.

This method is not ours. LMArena has ranked this way for years — their Rank (UB) assigns the same rank to models that are not statistically separable, and they have open-sourced the implementation. What is unusual is applying it to a *price* comparison, where the incentive is to publish a crisp ordering.

Applied to the Artificial Analysis axes only. Task Elo carries its own error, which we do not have, so it is not given a borrowed threshold.

## 3 · "Better and cheaper" has to mean it

Pareto dominance on price and capability is a two-dimensional projection of a much wider purchase decision, and the projection lies more often than not. A model can score higher and cost less while accepting no images, capping output at half the tokens, or holding half a million fewer tokens of context.

So **"both better and cheaper" is reserved for a replacement whose capability envelope is a superset** of the original. Of 133 dominance verdicts, 111 are strict and **22 (17%)** name what you would give up instead. Without that gate, 61% of raw two-axis verdicts recommended a model that could not do the incumbent's job.

## 4 · The frontier, not a value score

A model is on the **value frontier** when nothing else is simultaneously better and cheaper. 11 of the rated, priced, standard-delivery models qualify; the other 90% are strictly worse deals.

### Why not rank by quality per dollar?

Because the ratio rewards cheapness far more steeply than quality. Applied to this catalogue it crowns a model scoring 37.8 out of a ~63 ceiling, and puts one scoring 14.2 in second place — models nobody should ship. Any site publishing a "value" column computed that way is giving bad advice. We use the frontier plus an explicit quality floor instead.

## 5 · Price is a property of a model and a workload

There is no such thing as "the" price of a model. Summarisation is input-heavy, code generation is output-heavy, and an agent loop replays its context every turn; the same catalogue orders differently under each. Rather than pick one blend and hide it, the input:output ratio and cache-hit rate are controls you can see and change.

| Workload | Output share | Cache hits | Shape |
| --- | --- | --- | --- |
| Cân bằng | 25% | 0% | 3 tokens vào trên 1 ra |
| Tóm tắt | 5% | 30% | Đầu vào rất lớn, đầu ra rất nhỏ |
| Trò chuyện | 40% | 50% | System prompt dài, đã cache |
| Sinh mã | 60% | 20% | Đầu ra chiếm phần lớn |
| Tác tử | 15% | 70% | Ngữ cảnh phát lại mỗi lượt |

Three things routinely make the advertised rate wrong, and all are surfaced on the row:

 - **Context tiers.** 55 models change rate past a prompt-length threshold, some of them more than once — Qwen3.7 Flash goes 3.33× past 32K and 6.67× past 256K. Price a long-context workload off the headline and you are out by a multiple.
 - **Reasoning tokens billed apart from output.** On the affected models this can be 38–42% of a real bill and it does not appear in the advertised price at all.
 - **Time-of-day pricing.** 3 models discount ~50% inside published off-peak windows.

## 6 · Task-conditional ranking

A single global ordering is intellectually dishonest, because different models win different work. We carry per-task Elo across **25 categories** — front-end, data visualisation, game development, SVG, Android, slides and more. Selecting a task re-ranks the entire board on that task's Elo rather than on a general index.

Coverage varies by category, and a model absent from a category's leaderboard is shown as unmeasured for that task, not as bad at it.

## 7 · Where the data comes from

Prices are normalised from a machine-readable cross-provider catalogue and cross-checked against vendor pricing pages. Every model page links its source and fetch date.

 - **Conflicts are recorded, not averaged.** When two sources disagree on a price, that disagreement is the finding. A model listed at $4/$20 by its vendor and sold at $2/$10 through an aggregator has two true prices, and which one applies depends on where you buy.
 - **Implausible values are quarantined.** An upstream sentinel of −1 for "priced at request time" becomes −$750,000 per million if passed through naively, and wins every cheapest-first sort. Anything outside a plausible range is withheld with a reason attached.
 - **Confidence is per field.** Not one badge per row. Claiming a record is verified when only its price was checked is the kind of overstatement that ends trust the first time somebody notices.
 - **Unrated is not zero.** 260 models have no independent score. They are listed by price and excluded from quality ordering, never sorted to the bottom as though they had been measured and failed.

## 9 · Licensing, and what that constrains

Two of the benchmarks this site relies on have very different redistribution terms, and it changes what can be published.

 - **LMArena** publishes its leaderboard as a CC-BY-4.0 dataset, which permits redistribution with attribution — and carries the confidence bounds the significance ranks above depend on. That is the channel to use; scraping the site itself is prohibited by their terms.
 - **Artificial Analysis**'s free tier is *"internal use only with attribution"*; public redistribution requires a commercial licence. Their site terms separately limit use to personal and non-commercial. **This is unresolved.** Until it is, treat any Artificial Analysis figure here as shown for evaluation rather than published under a redistribution right.
 - **Pricing itself** is a fact a vendor publishes, which is a different thing from a compiled index score. Prices are cross-checked against vendor pages and every row links its source.

## 10 · Corrections

Prices change without notice and pages get reformatted; some of what is here will be wrong at any given moment. If you find an error, the fastest fix is to say which model, what the figure should be, and where it is published — corrections with a primary source get applied directly.

Catalogue last fetched 2026-08-24. Ranking logic is versioned with the site, so a change in how "best" is computed is a visible diff rather than a silent retune.
