Corrections

A vendor price change is not a correction. A figure we published that did not match the source is. This list is generated from dated snapshots. It launches thin because the archive is young, and that is the honest state.

Correction log

Every dated entry in the correction log, newest first, reproduced as logged, in English. Entries: 31.

Entries in the log with no date stamp, which cannot be placed on this timeline and are not listed: 16.

Task, review and R-numbers in these entries are internal identifiers of our release process, and commit hashes refer to a private repository.

99. Correction — the homepage did not explain which models its default view leaves out

Logged

The homepage hero counted every distinct model in the catalogue, while the line above the board counted models with a standard row on sale. Its search placeholder repeated the larger census, and the batch/free toggle claimed those tiers were always delivery modes of a model already listed. Models sold only as a batch or free listing make that statement false.

On the preserved release catalogue, measured by src/lib/board/directory.test.mjs, the census is 364 models: 145 ranked, 213 unrated, six sold only as batch/free listings and zero fully retired. Showing batch/free listings yields 145 ranked and 219 unrated. The new line identifies the six separately; a retired count appears only when nonzero. The ranked count also counts distinct models, matching the unrated count.

These extra counts appear only without a narrowing filter, where the line can reconcile to the census. Showing batch/free listings moves those models into the board's own counts. The search placeholder now counts the models its current view can find. The default view, price calculation and ranking rules are unchanged.

The disclosure and conditional toggle hint are translated in all 19 catalogues. The Arabic wording uses its existing Arabic term for the batch tier so the number reads before the description in the right-to-left line. This integrates source commits 50942244959bcd772a0bb416677ec1cbe0ba2faa and 0f67ba33035b338270bec9d183d2eea4dd79e79d without their unrelated pipeline changes.

Focused verification passed all eight directory tests, including fixture partitions, filter/toggle behavior, the current catalogue through the board's own query on general, document and task views, and the homepage's use of these counts. Full integrated build and browser verification are separate release gates; earlier A9 measurements are not claimed for this release.

98. Correction — analytics treated resource filters as additional page views

Logged

The claim that resource filtering caused no network requests was too broad. Filtering the resident catalogue did not fetch data, but Cloudflare's automatic SPA measurement observed the URL updates made by history.replaceState and sent analytics requests. A browser check immediately after networkidle could miss this: the app loads the analytics script two seconds after the document's load event.

After analytics had initialized, a publisher change, reset, search and reset produced eight /cdn-cgi/rum requests on each of the Skills, Agents and MCP hubs: one Ping and one XHR per URL-state change. Cold-load controls identified the separate initial document measurement. The first failing browser run recorded one request without its URL; its identity is unknown and is not inferred from these later measurements.

The manual beacon configuration in src/app.html now uses Cloudflare's documented spa: false. Document-load and Core Web Vitals measurement remain configured. The deliberate tradeoff is that client-only route transitions are no longer counted as additional page views. URL filter state remains intact; no history monkeypatch or custom beacon is used.

The regression in scripts/pagespeed-regression.test.mjs executes the actual inline loader, including its delayed initialization and duplicate protection. It failed before the fix and the full 20-test file passed afterward. A browser candidate with normal caching retained one initial document measurement per hub and made zero requests during settled publisher, reset and character-by-character search interactions on all three hubs. The final deployed release is checked separately against its sealed artifact.

Configuration reference: https://developers.cloudflare.com/web-analytics/get-started/web-analytics-spa/.

97. Correction — models a task board had scored were counted as having no score on it

Logged

Live (release A8, publication e44e3cfe…) prints the homepage task-board select's "{count} models have no score on this board", and the payload behind it lists only the 50 highest-scoring models of each task board, while coverage is the full count of models Design Arena scored. The page counts a model as unscored when it is not in that list. On eight of the 13 boards that offer a select, 13 to 66 models that Design Arena had scored were therefore called unscored: the reverse of "unrated is never 0", measured models called unmeasured.

  • What live prints. Under every one of the eight boards the homepage said 308 models have no score (358 distinct standard models, less the 50 listed). The true counts are in the table.
  • What is true. Every model Design Arena scored on a board is scored on it, and the board's frontier is drawn over all of them. Counts over the 358 distinct models with a standard row on sale, computed with the finder's own query (finderAnswer over runQuery), before and after the payload lists every scored id:

    | Board | Scored | Unscored, said | Unscored, true | Frontier, said | Frontier, true | |---|---|---|---|---|---| | website | 116 | 308 | 242 | 5 | 7 | | codecategories | 109 | 308 | 249 | 7 | 8 | | gamedev | 109 | 308 | 249 | 8 | 9 | | dataviz | 108 | 308 | 250 | 5 | 8 | | uicomponent | 107 | 308 | 251 | 7 | 8 | | 3d | 101 | 308 | 257 | 7 | 8 | | svg | 79 | 308 | 279 | 5 | 5 | | asciiart | 63 | 308 | 295 | 5 | 7 | | mobileapps | 42 | 316 | 316 | 9 | 9 | | fullstack | 41 | 317 | 317 | 5 | 5 | | webapps | 40 | 318 | 318 | 4 | 4 | | androidnative | 36 | 322 | 322 | 6 | 6 | | godotgamedev | 34 | 324 | 324 | 7 | 7 |

    The five boards of 34 to 42 scored models were whole. The models a capped board listed as unscored and undominated were gpt-oss-120b (six boards), gpt-oss-20b (two), DeepSeek V4 Flash 0731 (two) and DeepSeek V4 Flash 0423 (one).

  • How it was found. The final review of the grouped-menu branch compared /finder/'s count under a task with the board's own coverage. The cap is scripts/build-payload.mjs (.slice(0, 50)), written when coverage was added and never revisited; it is on master and live. The finder (this branch) and the homepage select both read the same list.
  • Fix. The payload lists every scored id per task (1,118 entries, from 726). The catalogue's brotli size goes from 28.0 KB to 30.6 KB, inside the 40 KB budget, so it stays in the catalogue. src/lib/run/finder.test.mjs fails if a board lists fewer ids than its coverage, or if a scored model is counted unscored.
  • Not changed. /best/ already says each board ranks the models Design Arena scored; with the full list that is now true.
  • What else moves. The same cap hid scores from the model pages, so 61 model pages gain a "Best rankings by task" block and 34 more gain a category in it (the review's measure on the sealed build), because every scored model is now published. The downloadable catalogue gains 392 Design Arena entries (1,118 against 726) and, with them, the credit it lacked (docs/LICENSING-AND-CLAIMS.md §6).

95. Correction — a manifest the site serves was removed from the repository as "never live"

Logged

Commit 66b628a, which recorded release A7, said: "the pre-merge candidate 1c467dbd was never live and its untracked manifest is removed". Half of that is true. The publication 1c467dbd… was never the site's publication: /data/publication.json on live names a35c5e38…. The other half removed a file the site serves.

  • What was true. A7 was built in a tree that still held that manifest, and build-publication.mjs copies every manifest in data/publications/manifests/ into the release. Release 20260930-164541-40bf8277f317 shipped it. On 2026-10-06 at 20:17Z https://undominated.ai/data/publications/manifests/1c467dbd283a….json, fetched with a cache-busting query, answered 200 with 14,108 bytes, last modified 2026-09-30 16:49:01 GMT, sha256 cc9ce3dd…. The live release's build seal (5ed233ac…) lists the same path, size and hash.
  • What that broke. The archive under /data/publications/ is append-only: a served path stays. From 66b628a on, every build from the repository would have dropped one published file. No release was made from such a build; live is still the A7 release.
  • Measured against the live seal. It serves 66 files under data/publications/: 39 inputs and 27 manifests. 65 were tracked, each byte-identical to the seal's hash. One was absent, that manifest. No other served file was missing or different, and the repository tracked no manifest or input that live does not serve. The 10 inputs the manifest names were all tracked already, which is why no gate noticed: verify-never-live.mjs knew which manifests must be absent and that inputs and manifests must match, and nothing said which files must be present.
  • How it was found. A dry run of the refresh guard, which compares a build with the live seal. The A8 preparation found the same path the same day by the same comparison, and first recommended leaving the file out, on the commit's reasoning. That recommendation was wrong for the same reason: "never live" is a statement about a URL, and this URL answers 200.
  • Now.
    • The manifest is restored, byte-identical, from the live site.
    • data/published-archive.json lists the 66 files with the sha256 of their bytes, taken from the live seal by scripts/published-archive.mjs record. A record only adds.
    • npm test requires every listed file in the tree, byte for byte, and refuses a served manifest on the never-live list (scripts/published-archive.test.mjs). With the manifest moved away the test fails and names it.
    • verify-never-live.mjs requires every listed file in the repository's build.
  • What moves. No page, price or figure. One file returns to the tree.

96. Correction — A model's price was whichever row set OpenRouter's model-level price, promotions and 4-bit copies included

Logged

A methodology change, approved for release A8 (docs/superpowers/specs/2026-10-04-reference-price-design.md). Every price this site published was OpenRouter's model-level price (§47, §92), which any seller can set. A7 accepted a seller's move to it "when traced exactly" to the endpoint that set it. Traced is necessary and not sufficient. A model's score was measured on its maker's deployment, and on A7's committed data the price beside it was, among others:

  • DeepSeek V4 Pro: StreamLake's fp8 row at a 55% promotion, $0.783 / $1.566;
  • DeepSeek V4.1 Flash: OpenInference's fp4 row, $0.0198 / $0.396, where DeepSeek's own endpoint declares no precision;
  • Llama 3.3 70B and Llama 3.1 70B: DeepInfra's turbo tier, $0.10 / $0.32 and $0.40 / $0.40;
  • Ling 3.0 Flash and Ling 3.0 Flash VL: Novita's rows at 65% and 72% promotions;
  • GPT-5.6 Sol and Sol Pro: OpenAI's own row at a 50% promotion, $2 / $10, while Azure lists $4 / $20.

That pairs a quality the buyer may not get with a price for a different product, and the frontier and the headline figures follow promotions.

Revised three times on 2026-10-06 and three times on 2026-10-07, before any release: after the review of this change, for the controller's ruling R121, which changed the rule itself, for R122, which retired the one price hold, for R124, which ranks a promoted row at its standard rate, for R128, which judges "serving" over the day, and for R131, which lets a bad day disqualify a row, with the re-review's two copy findings. What this entry said before each, and why it no longer does, is listed under "Revised" at its end; the figures below are re-measured on the same data.

The rule. The price is the reference price: the cheapest endpoint row on the balanced blend (3 input to 1 output, no cache) that is serving (status 0 or above, or status −2 at the read with a last-day uptime of at least 95%, R128; and a last-day uptime of at least 80% wherever the feed gives one, R131; a row at −5 or with no status is not serving), at standard delivery (no flex, batch, priority, fast, turbo, highspeed or ultrafast in the endpoint tag, as a segment after the seller or as the end of the seller's own segment, sambanova-turbo), at its standard time window (R106), and at the maker's precision or better. A row on promotion (pricing.discount above 0) is ranked and published at its standard rate, each rate ÷ (1 − discount), and its promoted price is its deal (R124). A 4-bit-class row (fp4, nvfp4, mxfp4, int4) passes only when the reference precision is 4-bit-class: the maker's own endpoint's declared precision, else the published weights' precision where the repo records it (data/selfhost-profiles.json weightsPrecision, one record: gpt-oss-120b, mxfp4), else unknown. A row that declares no precision passes when it is the maker's own, or a reseller's for a model whose weights are recorded closed (R124); for open or unrecorded weights a reseller's undeclared row never sets a reference price (R121). The row's whole rate card is published (rates, ladder, cache and reasoning rates, windows), never rates mixed from two rows. The deal is the price row's own promoted price where it has one, else the cheapest serving row with a reason that is cheaper today on the blend, with its reasons; it never sets a rank. src/lib/pricing/reference.mjs is the rule, scripts/apply-reference-price.mjs applies it where endpoint rows enter (scripts/fetch-endpoints.mjs), and data/models.json keeps OpenRouter's model-level record as modelLevelPricing.

Readings the spec left open. The review ratified D1 and D3, and the spec states each of these since its amendment of 2026-10-06, so nothing here is a reading any more:

  • D1, deal-only. When no row passes, the spec kept "the cheapest serving row". It is the cheapest serving row of standard delivery when one exists. Read the old way, Gemini 3.7 Flash and Gemini 3.8 Flash would be priced at Google's flex tier, $0.375 / $1.875, half the standard rate, because every Google row of both carries the same 50% promotion. Since R124 the fallback is a ladder (a reseller's undeclared row, then a 4-bit one, another tier last), and neither Gemini is on it: the standard-delivery row passes at its standard rate, $1.50 / $7.50, a reference price, and its promoted $0.75 / $3.75 is the deal.
  • D2, a negative discount (StreamLake on MiMo V2.5 and V2.5 Pro, −0.2) is a surcharge, not a promotion. It changes no price on this data. The review did not rule on it; the controller ratified it on 2026-10-06, and the spec's rule 3 reads "above 0", which is what the code does.
  • D3, R121 and R124, a row that declares no precision. The maker's own row is the maker's precision by definition and passes (D3). A reseller's does not pass (R121) unless the model's weights are recorded closed (R124): unknown is never full precision, so the row is not known to be like-for-like. It is the deal, with the reason "precision not disclosed", when it is cheaper. 211 published prices are taken from an undeclared row: 192 reference prices, every one the maker's own row (OpenAI, Anthropic, Google and the like), 6 of them at the standard rate of the maker's promoted row, and 19 deal-only prices, every one a reseller's. Until R121 the reseller's row passed and its price carried a label; see "What R121 changed".
  • No serving row. Three models (Solar Pro 4, Mistral Small 3.1 24B, L3.3 Euryale 70B) have endpoint rows and none serves. They keep the model-level price, priceBasis: "model-level", as do the 100 published models with no endpoint row at all (batch and free listings, the two per-unit models, models no endpoint lists). MiMo V2.5 was the fourth until R128: Xiaomi's own fp8 row reads −2 and answered 96.3% of the last day.
  • A held price is not repriced. A keep-live hold on pricing (data/upstream-holds.json) holds the published price: the model skips the rule and keeps the live record's price, basis, row and deal (applyHeldPrice), and heldViolations fails the build if the published pricing is not the live one. The first draft of this change repriced a held model and pinned the hold to modelLevelPricing, which no page prints: the hold held nothing. No price is held on this data any more: the one such hold, GPT-5.6 Sol Pro's, is retired (R122, "The Sol Pro hold" below).
  • Ties on the blend go to the cheaper cache read, then a row that read status 0 or above over one that serves by its day (ruled with R128), then the maker's own row, then a declared precision, then seller and tag. Seven models tie on rows whose rates differ (MiMo-V2.6-Flash, DeepSeek V4 Flash Vision Exp, Kimi K2.5, GLM-4.6V, MiniMax M2, GPT-5 Mini, GPT-5 Nano); on Kimi K2.5 the tie-break takes Novita's standard $0.60 / $3.00 with a $0.10 cache read over Amazon Bedrock's $0.60 / $3.00 with none. It was four before R124 and R128, Qwen3.8 27B among them, and twelve before R121: eight were a tie with a reseller's undeclared row, OpenAI's own against Azure's for five of them. Rows with the same three rates are one price, not a tie: GPT-5.6 Sol's is OpenAI's standard rate and Azure's row alike, and the maker's own row is named.

Before and after, on the same data. A7's committed data/models.json and data/offers.json (endpoints fetched 2026-09-30, regenerated from data-e's raw endpoint cache so each row carries its status: the only change to offers.json is 1,303 status lines). "Before" prices every model at modelLevelPricing; both sides' figures come from scripts/build-payload.mjs itself, run by scripts/measure-reference-price.mjs. Before reproduces the A7 catalogue's stats exactly. GPT-5.6 Sol Pro is not among the 66 rate sets below any more. Its model-level card, $4 / $20 with $8 / $30 past 272K, which A7 never published (it was held at $2 / $10), is field for field what OpenAI's row gives at its standard rate (R122 and R124, below).

beforeafter
prices that change—66 rate sets, 64 balanced prices: 32 up, 32 down; 41 of them rated
price basismodel-level, allreference 320 · deal-only 21 · model-level 103 · 85 deals
frontier1011
dominated133 of 143132 of 143
median overpay7.5× (n 133)6.5× (n 132)
balanced spread12,069×12,500×
  • The frontier. Out: Qwen3 30B A3B Instruct 2507 (1384, $0.0844 → $0.1425 at SiliconFlow fp8). In: Llama 3.1 8B Instruct (1186.6, $0.0575 → $0.025 at DeepInfra fp8) and Gemma 4 26B A4B (1434.1, $0.1425 → $0.1275 at DekaLLM bf16). Staying, repriced: gpt-oss-120b $0.0703 → $0.065, DeepSeek V4 Flash $0.1022 → $0.1125 (DeepInfra fp8), DeepSeek V4.1 Flash $0.1139 → $0.17 (Io Net fp8), GLM-5.3 Flash $0.2375 → $0.1829 (AtlasCloud fp8), MiMo-V2.6-Pro $0.5438 → $0.54. Gemini 3.8 Flash $1.50 → $3.00, the standard rate of Google's row, every Google row of it being at 50% off. GPT-oss-20b, Gemma 3 4B and Claude Opus 5.5 do not move. All eleven are reference prices. Until R124 Gemini 3.8 Flash was on the frontier at the promoted $1.50, a deal-only price.
  • The spread is o1-pro's $262.50 over Mistral Nemo's $0.02175 before and $0.021 after (DekaLLM fp8, $0.018 / $0.03, against the model-level $0.019 / $0.03).
  • The largest moves, balanced: Mercury 2.5 +400% (Inception's own and only row, at 80% off: $0.04 / $0.15 today, $0.20 / $0.75 standard), Qwen3 Coder +311% (DeepInfra's turbo fp4 row at $0.30 / $1.00 to Alibaba's own $0.975 / $4.875, since Novita's fp8 row answered 26.0% of the day, R131), Ling 3.0 Flash VL +189% and Ling 3.0 Flash +186% (Novita's promotions to DeepInfra's $0.06 / $0.18), LongCat 2.0 +150% (AtlasCloud fp8 at 60% off, its standard $0.75 / $3.00), DeepSeek V4 Flash Vision Exp +104% (DeepInfra's 51% promotion to the same row's standard rate), Solar Mini 4, Gemini 3.8 Flash, Gemini 3.7 Flash and GPT-5.6 Sol +100% each (the maker's own row at 50% off), DeepSeek V4 Flash 0731 −88% (Relace's fp4 row at $0.01 / $1.28 to OpenInference fp8 at $0.0099 / $0.1307, a row that reads −2 and answered 96.9% of the day), Llama 3.1 70B +80% (DeepInfra's turbo tier to Amazon Bedrock as a deal-only price), Qwen3 30B A3B Instruct 2507 +69%, Laguna XS 2.1 +67%, DeepSeek V4 Pro +66%, GLM-5.3 −64%, Kimi K3 −63%.
  • The spec's "7.7× to 6.4×" does not reproduce here: before is 7.5× on the committed A7 data.

What R121 changed, measured alone. The same data priced by the rule as it stood before the ruling (data/models.json at 8657661), against the rule now (node scripts/measure-reference-price.mjs --against <that file>):

  • 30 published prices were taken from a reseller's undeclared row: 28 reference prices, 13 of them rated, and 2 that were already deal-only.
  • 12 prices change, 8 of them rated.
    • 10 move up to the next row that passes: DeepSeek V4 Pro 0813 +61% (Wafer to NextBit fp8), Gemma 4 26B A4B +47% (Darkbloom to DekaLLM bf16), GLM-5.3 +39% (Reka to Morph fp8), GLM-5.2 +38% (DigitalOcean to Alibaba fp8), DeepSeek V4 Pro +25% (DigitalOcean to DeepInfra fp8), GLM-5.3 Flash +21% (Relace to AtlasCloud fp8), Llama 4 Maverick +15%, Qwen3 Coder +9%, DeepSeek V4.1 Flash +7% (Wafer to Io Net fp8), Muse Glimmer 30B +5%.
    • 2 move down. Nothing passes for them any more, and the fallback takes the cheapest serving row of standard delivery, whatever it failed for. GPT-5.6 Sol goes from Azure's $4 / $20 to OpenAI's own row at a 50% promotion, $2 / $10 (−50%). Kimi K2.5 goes from Amazon Bedrock's $0.60 / $3 to SiliconFlow's int4 row, $0.45 / $2.25 (−25%). Both are rated, both are back at their model-level price, and each now rests on a row the rule excludes for a known reason, where before it rested on one whose precision was not known.
  • 16 keep their row and their rates and become deal-only, with the reason published: the Azure rows of five GPT-5.x Codex and Chat models and GPT-3.5 Turbo 0613, Amazon Bedrock for Claude Opus 4.1, Claude Sonnet 4, Palmyra X5 and Llama 3.1 70B, Cloudflare for three, and Alibaba, Groq and Phala for one each.
  • Basis: reference 328 → 310, deal-only 11 → 29, model-level 105 → 105. Deals 75 → 76: three new ones are the undeclared row that used to be the price (Gemma 4 26B A4B, Llama 4 Maverick, Muse Glimmer 30B), two are gone because the deal became the price (GPT-5.6 Sol, Kimi K2.5), and five gain a second reason. 8 deals carry "precision not disclosed", 5 of them with a promotion.
  • Headline figures, before the ruling → after it: frontier 9 → 11 (DeepSeek V4 Flash and DeepSeek V4.1 Flash return, because Gemma 4 26B A4B and GLM-5.3 Flash no longer undercut them from an undeclared row), dominated 134 → 132 of 143, median overpay 7.3× → 6.5×, spread unchanged.
  • What it does not change. A deal-only price still enters the ranking. No reference price rests on a reseller's undeclared row any more, but 11 rated models are deal-only (4 before the ruling): Gemini 3.8 Flash, Gemini 3.7 Flash and GPT-5.6 Sol on a promotion, Kimi K2.5 and Gemma 2 27B on a 4-bit copy, DeepSeek Chat on a promoted row of an undeclared reseller, and Claude Opus 4.1, Claude Sonnet 4, Llama 3.1 70B, Qwen 2.5 Coder 32B and Llama 3.2 1B on an undeclared reseller's row. Whether a deal-only price should rank is not decided here. R124, the next day, takes the first six of these off a fallback; see below.

What R124 changed, measured alone (2026-10-07). R121 left that hole: a promoted rate ranked, and sat on the frontier. R124 changes the rule in three places. The measurements were taken before anything was built.

  • Is a promotion uniform across a row's rates? node scripts/measure-promotion-fit.mjs, on the raw endpoint cache. 117 rows on 48 of the 344 models carry a discount, from 0.001 to 0.9, none at 1 or above. Rate ÷ (1 − discount) is exact at four decimals per million on:
    • input 117 of 117, output 117 of 117, cache read 110 of 110 (the rows that state one);
    • all 24 rung rates of the context ladders on promoted rows;
    • cache write 10 of 18. The other 8 are Gemini rows whose feed value is a repeating decimal (0.0000000416666666666667 per token, a twenty-fourth of a dollar per million), so no four-decimal figure is exact. The quotient is published as the pipeline rounds a cache write, at six decimals: $0.083333 for Gemini 3.8 Flash, twice the model-level record's $0.041667.
    • web_search is a per-call fee and is not discounted: the same on 6 of 6 promoted rows that have an undiscounted row to compare (OpenAI's $0.01 beside Azure's $0.01).

    So every rate of every promoted row is derived, and no row is excluded for a rate that does not fit.

  • Is the quotient the seller's standard rate? Where it can be checked, yes.
    • OpenAI, GPT-5.6 Sol: $2 / $10 / $0.20 at 0.5 gives $4 / $20 / $0.40, Azure's undiscounted row exactly, and its $8 / $30 rung is OpenRouter's own model-level rung for Sol Pro. OpenAI's fast row is no evidence: it is itself at 50% off (standard $8 / $40).
    • DeepSeek V4 Pro: StreamLake's $0.783 / $1.566 / $0.06525 at 0.55 and GMICloud's at 0.45 both give $1.74 / $3.48 / $0.145, NextBit's undiscounted row exactly. Baidu's $0.38701 / $0.77402 / $0.03206 at 0.771 gives $1.69 / $3.38 / $0.14: Baidu's own standard rate, not the $1.74 list.
    • Gemini 3.7 Flash: the hand-researched note on its model-level record says the promotional rate "rises to $1.50 input / $7.50 output on 2027-01-01". That is the quotient.
    • 66 of the 117 rows have an undiscounted row of the same model at their standard input and output. No seller lists a promoted and an unpromoted row of one model at one tier.
  • Open weights. /open-weights/ reads the catalogue's ow, which is data/models.json openWeights. Of the 444 published models it is true on 34, false on 72 and not recorded on 338; of the 344 with endpoint rows, 25, 40 and 279; of the 143 rated, 11, 33 and 99. Not recorded is treated as open. The clause changes no price on this data: no reference price is a reseller's undeclared row, and all 21 deal-only models have weights not recorded, the Azure rows of five GPT-5.x models and Amazon Bedrock's Claude Opus 4.1 and Claude Sonnet 4 among them. It will do its work when the field is filled; filling it is a data task with a cite per model, not this change.
  • 13 prices change, against data/models.json at 093fdd9, 6 of them rated. 12 move up:
    • to the standard rate of the row they were already on: Mercury 2.5 +400%, LongCat 2.0 +150%, Solar Mini 4, Gemini 3.8 Flash, Gemini 3.7 Flash, GPT-5.6 Sol and GPT-5.6 Sol Pro +100%, Ling 3.0 Flash Fin +79% (still deal-only: Novita declares no precision), Laguna XS 2.1 +67%, Laguna S 2.1 and DeepSeek Chat +11% (the latter still deal-only);
    • Kimi K2.5 +33%, up the ladder from SiliconFlow's int4 row at $0.45 / $2.25 to Novita's undeclared row at its standard $0.60 / $3.00.
    • 1 moves down: Qwen3.8 27B −9.1%, from Chutes fp8 at $0.24 / $2.20 to AkashML fp8, whose 10% promotion counts at its standard $0.225 / $1.98 and is cheaper than Chutes there.
  • Basis: reference 310 → 319, deal-only 30 → 21, model-level 104. Nine deal-only prices become reference prices. Deals 77 → 86.
  • Headline figures: frontier 11 → 11, the same eleven; dominated 132 of 143 both ways; median overpay 6.5× → 6.7×; spread 12,500× both ways.
  • 14 prices are a promoted row's standard rate, 11 reference and 3 deal-only, 6 rated. Each shows that row's promoted price as "Cheaper right now". On 5 of them another serving row is cheaper today than that promoted price and is not shown, because a model has one deal and the price row's own promotion comes first: Gemini 3.8 Flash and Gemini 3.7 Flash (Google's flex rows at $0.375 / $1.875), GPT-5.6 Sol Pro (OpenAI's flex row at $1 / $5), Qwen3.8 27B and Kimi K2.5.
  • On the two of those 14 whose row is the maker's own and whose model-level record has a note and a batch percentage, Gemini 3.7 Flash and GPT-5.6 Sol, neither is carried beside the standard rate. "Batch discount 50%" beside a $1.50 input reads as $0.75, where the same note says Google's batch rate is $0.375; and Sol's note says the prices "use the accepted OpenRouter listing", which is $2 / $10. Both stay on modelLevelPricing.
  • Still deal-only and rated: 8. Kimi K2.5 (1445.4, Novita, standard $0.60 / $3.00), Claude Opus 4.1 (1418.5, Amazon Bedrock), Claude Sonnet 4 (1339.1, Amazon Bedrock), DeepSeek Chat (1332.5, StreamLake, standard $0.286 / $1.143), Llama 3.1 70B (1260.9, Amazon Bedrock), Gemma 2 27B (1231.4, NextBit int4), Qwen 2.5 Coder 32B (1230, Cloudflare), Llama 3.2 1B (1054.6, Cloudflare). Seven rest on a reseller that declares no precision and one on a 4-bit copy. None is on the frontier, and none is named as the model that dominates another. They still enter the ranking: that question is measured here and not decided.

"Serving" over the day (R128). A row is serving when its status is 0 or above, or when it read −2 and its last-day uptime is at least 95%. A row with no uptime figure is judged by its status alone, and a row at −5 is not serving whatever its day figure. The ruling was first worded for any status below 0; it was measured before it was built, and the controller ratified it for −2 only (node scripts/measure-serving-uptime.mjs):

  • The feed states uptime_last_1d on 1,256 of the 1,303 rows. 85 rows read below 0, 61 at −2 and 24 at −5, every one with a figure. 41 of them, on 28 models, are at 95% or more: 35 at −2 and 6 at −5. Only 2 are at 99% or more.
  • The status is a band of the last half hour's uptime (uptime_last_30m): status 0 is 95.3% to 100% on the 924 rows that state it, −2 is 80.1% to 95.0% on 61, −5 is 0% to 79.6% on 23. So a −2 row at 95% over the day is a row that dipped, which is the flicker A8 saw: rows reading −2 on one fetch and 0 on the next (§93 on work/data-f-20261004). The six −5 rows at 95% over the day answered 69.9% to 79.6% of the last half hour.
  • With every status below 0, 8 prices would change, and one −5 row would move the headline. OpenInference's fp8 row for DeepSeek V4 Flash, $0.005 / $0.1605, 79.6% over the half hour and 97.3% over the day, would price that model 61% lower, take Gemma 3 4B and gpt-oss-120b off the frontier (11 → 9), and move the median overpay from 6.7× to 10.3× (dominated 134 of 143). Solar Pro 4 would be priced at Upstage's row, also −5.
  • With status −2 only, the rule, 5 prices change, all down, 3 of them rated: DeepSeek V4 Flash 0731 −55.5% (OpenInference fp8, $0.0099 / $0.1307, 84.1% over the half hour, 96.9% over the day), Qwen3.6 27B −28.4% (Chutes fp8), Llama 3.3 70B −28.1% (Novita bf16), Qwen3 235B A22B 2507 −18.8% (Novita fp8), Grok 4.7 −9.1% (xAI's own xai row at $2 / $6, where its only standard-delivery row that read 0 is zdr/us at $2.20 / $6.60). MiMo V2.5 goes from the model-level price to a reference price at the same $0.14 / $0.28.
  • Basis: reference 319 → 320, model-level 104 → 103, deal-only 21. Deals 86 → 85. Frontier 11, dominated 132 of 143, median overpay 6.7× → 6.5×, spread unchanged.
  • The reason for the line at −2. The rule admits no row that failed more than a fifth of its last half hour: a row at −5 is down now, and a good day does not make it a price a buyer gets.
  • A tie-break came with it. A −2 row that serves by its day can tie a row that read 0. The first build named it by seller and tag: Gemini 3.8 Flash at Google's google-vertex/global (−2, 97.8% over the day) over Google AI Studio's row at the same three rates, and Qwen3.5 122B A10B at Alibaba's own −2 row over SiliconFlow fp8. The controller ruled that a row that read 0 comes first, then the maker's own. Both models are named at the row that read 0 again; no price and no count changes. The step sits after the cache read and applies to every tie from there, because a step that applied only to equal rate sets would not be an order.

A bad day disqualifies a row (R131). R128 worked in one direction: a day's uptime admitted a −2 row and never removed a row at 0. The re-review of 1418ea8 found the reverse case on a published price (N8), and the controller ruled a floor: a last-day uptime below 80%, where the feed gives one, is not serving, whatever the row reads now. Measured before it was built (node scripts/measure-serving-uptime.mjs):

  • Of the 1,218 rows at status 0 or above, 1,171 state a last-day uptime: 1,125 at 95% or more, 24 between 90% and 95%, 12 between 80% and 90%, 7 between 50% and 80%, 3 below 50%. 47 state none and are judged by their status.
  • 10 rows on 10 models read 0 with a day under 80%: GPT-6 Luna at Amazon Bedrock (1.6%), Qwen3 Coder at Novita fp8 (26.0%), DeepSeek V4 Flash at Baidu fp8 (45.2%), Claude Sonnet 4.6 at Azure global (56.3%), Kimi K3 at Decart mxfp4 (60.2%), Claude Opus 4.6 at Azure global (69.0%), Claude Fable 5 at Azure (69.7%), GPT-5.6 Sol at Amazon Bedrock (72.5%), MiniMax M2.7 at Mara (75.2%), Grok 4.20 Multi-Agent at xAI xai (77.7%).
  • 1 price changes. Qwen3 Coder (unrated) was at Novita's fp8 row, $0.38 / $1.55. Its other rows: DeepInfra turbo fp4 (a tier and 4-bit), Google Vertex $0.22 / $1.80 (an undeclared reseller), Venice fp8 (−5), and Alibaba's own, $0.975 / $4.875: 0.73125 + 1.21875 = 1.95, the price now, +190%. Its deal stays DeepInfra's turbo row. Grok 4.20 Multi-Agent keeps $1.25 / $2.50 and is named at xAI's xai/zdr row.
  • The other eight rows set no price. No count and no headline figure moves: reference 320, deal-only 21, model-level 103, 85 deals, frontier 11, dominated 132 of 143, median overpay 6.5×, spread 12,500×.
  • The re-review listed five prices on a row at 0 with under 95% for the day. Two are under the floor and are the two above. Three stay: Qwen3 VL 235B Thinking (87.4%), GLM-5.1 at Chutes (88.4%) and MiMo V2.5 Pro (91.7%). The floor is 80%, not 95%: at 95% the 36 rows between would go too.

The round (R124, R128 and R131), against 093fdd9: 19 prices change, 13 up and 6 down, 9 of them rated. Reference 310 → 320, deal-only 30 → 21, model-level 104 → 103, deals 77 → 85. Frontier 11 → 11, dominated 132 of 143 → 132 of 143, median overpay 6.5× → 6.5× (6.7× after R124 alone), spread 12,500× → 12,500×.

The Sol Pro hold, retired (R122). data/upstream-holds.json kept GPT-5.6 Sol Pro's pricing at the live publication's $2 / $10 from 2026-09-30, because OpenRouter's model-level price for it had moved to $4 / $20 the day before while GPT-5.6 Sol's stayed at $2 / $10, and publishing both would have shown a factor of two between two models OpenAI prices the same. Under this rule the model-level price no longer decides either, so the hold has nothing left to prevent, and the controller retired it. Both are now priced by the rule, from their own endpoint rows:

GPT-5.6 Sol (1455.6)GPT-5.6 Sol Pro (unrated)
OpenAI openai/flex$1 / $5, promotion 50%, not serving (−5)$1 / $5, promotion 50%, serving
OpenAI openai$2 / $10, promotion 50%$2 / $10, promotion 50%
OpenAI openai/fast$4 / $20, promotion 50%, a tier$4 / $20, promotion 50%, a tier
Azure azure$4 / $20, precision not declared$4 / $20, precision not declared
other rowsAzure us and eu, Amazon Bedrock, $4.40 / $22, not declaredAzure eu, $4.40 / $22, not declared
published price, at R122$2 / $10 / $0.20 cached, $4 / $15 past 272Kthe same
basis, row, reason, at R122deal-only, OpenAI openai, promotion 50%the same
deal, at R122none: the flex row is not servingOpenAI openai/flex, $1 / $5: delivery tier and promotion 50%
published price, since R124$4 / $20 / $0.40 cached, $8 / $30 past 272Kthe same
basis and row, since R124reference, OpenAI openai at its standard rate (discount 0.5)the same
deal, since R124OpenAI openai, $2 / $10: promotion 50%the same
OpenRouter's model-level price$2 / $10$4 / $20, $8 / $30 past 272K
  • The ruling expected Sol at Azure's $4 / $20 with the promotion as its deal. That was the rule's answer until R121, the same day. Azure declares no precision, so after R121 its row could not set a reference price, nothing passed for either model, and the fallback took OpenAI's promoted row: the promotion was the price, labelled deal-only with its percentage. The controller agreed that this was the hole and ruled R124. Both are now a reference price of $4 / $20 at OpenAI's own row, counted at its standard rate, with the promoted $2 / $10 as the deal, which is what R122 expected, by a rule and not by a case.
  • At R122 no published price moved. Sol Pro's rate card was $2 / $10 with $4 / $15 past 272K under the hold and was the same card from OpenAI's row. What changed was what is said of it: basis model-level (a held record) → deal-only, the row and its reason named, one deal more (77), and modelLevelPricing recording the $4 / $20 the hold had kept off the page. Counts: deal-only 29 → 30, model-level 105 → 104. Since R124 both prices are $4 / $20, and Sol Pro's card is its model-level card field for field, so it is no §96 restatement. The 2026-09-30 snapshot has both at 4.00 balanced and this data has 8.00: Sol's move is a §96 restatement (model-level $2 / $10); Sol Pro's is not one the restatement list can carry, and the wire reads it as a move. See "Now".
  • How the record was rebuilt. scripts/build-dataset.mjs stamps every record with the day it runs, so it was not re-run. Sol Pro's model-level pricing was rebuilt from A7's raw models feed (data/.cache/openrouter-raw.json, sha256 1d776d2a…9a8b41, read 2026-09-30T12:21:43Z) by the functions the pipeline uses, normalisePricing then cleanPricingMetadata, and scripts/fetch-endpoints.mjs re-applied the rule from the endpoint cache. The same two functions reproduce the committed model-level pricing of 442 of the 444 other records; the two that differ are the per-unit Lyria models, which a later step rewrites. The tracked data/upstream-baseline.json has the same $4 / $20 for it.

Hand checks, each against data/offers.json:

  • GPT-5.6 Sol (1455.6, weights recorded closed), 7 rows. OpenAI openai, 50% off: $2 ÷ 0.5 = $4, $10 ÷ 0.5 = $20, $0.20 ÷ 0.5 = $0.40; blend 0.75 × 4 + 0.25 × 20 = 3 + 5 = 8. Azure azure, $4 / $20 / $0.40, declares no precision and passes because the weights are closed: blend 8. The two are one rate set, and the maker's own row is named. Dearer: Azure us and eu at $4.40 / $22 (8.8). Out: openai/flex (a tier, and −5), openai/fast (a tier), and Amazon Bedrock's $4.40 / $22 row, which read 0 and answered 72.5% of the day (R131). Reference $4 / $20 / $0.40, cache write $2.50 ÷ 0.5 = $5, past 272K $4 / $15 ÷ 0.5 = $8 / $30, web search $0.01 as the feed states it. Deal: OpenAI openai at $2 / $10, promotion 50%.
  • GPT-5.6 Sol Pro (unrated, weights not recorded), 5 rows. The same OpenAI row at the same standard $4 / $20, blend 8. Azure's rows are a reseller's undeclared rows for a model not recorded closed, so they do not pass; the price is the same without them. Deal: OpenAI openai at $2 / $10, promotion 50%; the flex row at $1 / $5, serving here, is cheaper and is not shown.
  • Gemini 3.8 Flash (1494.4, weights not recorded), 6 rows, all Google's own and all at 50% off. Four are a flex or priority tier. google-ai-studio (status 0) and google-vertex/global (−2, 97.8% of the day): $0.75 ÷ 0.5 = $1.50, $3.75 ÷ 0.5 = $7.50, $0.075 ÷ 0.5 = $0.15; blend 0.75 × 1.5 + 0.25 × 7.5 = 1.125 + 1.875 = 3.00. One rate set; the row that read 0 is named, Google AI Studio's. Reference $1.50 / $7.50 / $0.15, reasoning $7.50, cache write $0.083333, web search $0.014. Deal: the same row at $0.75 / $3.75, promotion 50%. It is on the frontier at $3.00; until R124 it was there at $1.50.
  • Kimi K2.5 (1445.4, weights not recorded), 5 rows, no Moonshot row, no weights record: reference precision unknown, so SiliconFlow int4 $0.45 / $2.25 (0.90) and AtlasCloud int4 (0.9925) fail as 4-bit, and Novita, Amazon Bedrock and Venice declare nothing. Nothing passes. Rung (a), a reseller's undeclared row at its standard rate: Novita at 5% off, $0.57 ÷ 0.95 = $0.60, $2.85 ÷ 0.95 = $3.00, $0.095 ÷ 0.95 = $0.10, blend 0.45 + 0.75 = 1.20; Amazon Bedrock $0.60 / $3.00, no cache read, 1.20; Venice at 5% off, standard $0.56 / $3.50, 1.295. Novita and Bedrock tie; a stated cache read beats none. Deal-only $0.60 / $3.00 / $0.10, reason "precision not disclosed". Deal: Novita today, $0.57 / $2.85, promotion 5% and precision not disclosed. The int4 rows are cheaper today (0.90) and are neither the price nor the deal.
  • DeepSeek V4 Pro, 16 endpoint rows, all serving, weights recorded open. No DeepSeek endpoint, no weights-precision record: reference precision unknown. Not passing, cheapest first: Relace fp4 $0.388 / $3.50 (4-bit, 1.166), DigitalOcean $1.044 / $2.088 (precision not declared, 1.305) and Cloudflare $1.15 / $2.55 (not declared, 1.50). The cheapest row that passes is DeepInfra fp8, $1.30 / $2.60: 0.75 × 1.30 + 0.25 × 2.60 = 0.975 + 0.65 = 1.625. Reka ties it at $1.30 / $2.60 and declares nothing, so it is not a second candidate; Alibaba fp8 is next at 1.77. The three promoted fp8 rows pass and are dearer at their standard rates: Baidu $0.38701 ÷ 0.229 = $1.69 and $0.77402 ÷ 0.229 = $3.38 (2.1125); StreamLake $0.783 ÷ 0.45 = $1.74 and $1.566 ÷ 0.45 = $3.48, and GMICloud $0.957 ÷ 0.55 = $1.74 and $1.914 ÷ 0.55 = $3.48 (2.175, NextBit's undiscounted row). Reference $1.30 / $2.60 / $0.10 cached, +66% on the model-level 0.97875. Deal: Baidu today, $0.387 / $0.774, promotion 77%. R124 and R128 do not move it. Before R121 the price was DigitalOcean's undeclared row at 1.305; that ruling alone moved it by +24.5%.
  • Kimi K3. Moonshot AI's own row declares mxfp4, so 4-bit rows pass. Sail Research fp4 $0.3177 / $7.9431: 0.75 × 0.3177 + 0.25 × 7.9431 = 2.22405, below InferenceNet fp4 (2.55), Relace fp4 (2.80), DeepInfra bf16 (5.70) and Moonshot's own $3 / $15 (6.00). Wafer (3.331275) and DigitalOcean (5.15) declare no precision; Makora and Morph read −2 and are dearer either way; Phala's row at a 15% promotion counts at its standard $3 / $15 (6.00). Reference $0.3177 / $7.9431 / $0.40, −63%. No deal. No ruling moves it.
  • GLM-5.3. Z.AI's own row is fp8, so fp4 and nvfp4 rows are excluded. Baidu fp8 $0.294 / $0.924 is at a 79% promotion and counts at its standard $1.40 / $4.40 (2.15), Z.AI's own rate, and Reka's $0.37 / $1.14 (0.5625) declares no precision. The cheapest row that passes is Morph fp8, $0.1428 / $2.6943: 0.1071 + 0.673575 = 0.780675, below AtlasCloud fp8 (0.9245) and Sail Research fp8 (1.00); InferenceNet (0.87) and Wafer (0.968675) declare nothing. Reference $0.1428 / $2.6943 / $0.1857, −64% on $1.40 / $4.40 (2.15). Deal: Baidu, promotion 79%. Before R121 the price was Reka's row, −74%. Morph's row is cheap on input and dear on output. The rule chooses on the balanced blend only: at one input token to one output, AtlasCloud's fp8 row is cheaper (1.247 against 1.41855).
  • gpt-oss-120b, natively mxfp4. No OpenAI endpoint; the weights record says mxfp4, so 4-bit rows pass. CoreWeave fp4 $0.03 / $0.17: 0.0225 + 0.0425 = 0.065, below DekaLLM bf16 $0.03 / $0.18 (0.0675) and DeepInfra bf16 $0.037 / $0.17 (0.07025, the model-level price). Reference CoreWeave, −7.5%. No deal. Without the weights record the fp4 row would be the deal and DekaLLM the price.
  • Llama 3.3 70B, no maker endpoint (Meta sells none). DeepInfra turbo fp8 $0.10 / $0.32 (0.155) is a delivery tier. Novita bf16, $0.135 / $0.40, reads −2 and answered 96.8% of the day, so since R128 it serves: 0.10125 + 0.10 = 0.20125, below AkashML fp8 $0.20 / $0.52 (0.15 + 0.13 = 0.28) and Parasail fp8 (0.29). Reference Novita, +30% on the model-level 0.155; it was AkashML, +81%, while Novita's row counted as down. Deal: DeepInfra turbo, delivery tier.

The oracle. scripts/reference-price-oracle.mjs reads the spec's text, with no reading of its own, and imports nothing from the rule or from the code that applies it. It agrees with the pipeline on 444 of 444 published models (npm test, and scripts/verify-invariants.mjs): reference 320, model-level 103, deal-only 21, none held. On a tie it accepts any tied row, so the pipeline's tie-break is reported, not assumed (7 models). Since R124 it ranks a promoted row at the standard rate offers.json carries, and checks that rate two ways: standard × (1 − discount) is the row's own rate on 117 of 117 promoted rows, and, with the raw endpoint cache present, all 117 are recomputed from the feed's decimals with 0 differing. With the raw endpoint cache present it also checks every row's status against the raw feed: 1,303 of 1,303 agree. Since R121 it also checks the reasons published with a deal and with a deal-only price against its own, so a ranked price taken from a reseller's undeclared row would fail twice: it is not the oracle's row, and a reference price has no reasons to carry.

  • What it cannot catch. It shares two things with the pipeline: the publication policy, and the maker test isFirstPartySeller, which decides whose row is the maker's. A wrong maker mapping would be wrong in both. The reviewer's own oracle used a maker map written by hand and agreed on the reference precision of 344 of 344 models with endpoint rows.
  • Until 2026-10-06 it was not literal. It built D1 and D2 in by default and offered the spec's words as --literal, where it agreed on 442 of 444: the two Gemini Flash models. An oracle that shares the implementer's readings agrees by construction. The spec was amended instead, and the flag is gone.

A context rule, measured and not applied (spec, "Not in scope"). Excluding rows whose context window is below the model's published one would change 31 published prices (30 before R124 and R128, 31 before R121). Among them Qwen3.8 27B (AkashML serves 262,144 of 1,000,000 tokens; Alibaba's own row at $0.425 / $2.55 would be the price), MiniMax M3 (DeepInfra serves 524,288 of 1,048,576), Llama 4 Scout (327,680 of 1,310,720), GLM-5.2, and three Qwen3 models whose cheapest passing row serves 40,960 of 131,072 tokens.

The A7 precedent is superseded. A promotional or 4-bit headline is no longer accepted because it traces exactly; it is disclosed as a deal. Traceability still applies to every row: the reference row names its seller and tag, and the oracle re-reads it.

Now.

  • Every reader of data/models.json pricing follows: the board, ranking, frontier, dominance, verdicts, snapshots and series, badges and hallmarks, exports, the MCP package and the page companion. /batch/'s standard rate is the reference price (en, de); the model page says where each price is from, in a sentence that claims only what its basis supports (a reference price is "at its standard rate", no longer "with no promotion", and "at the maker's precision or better" only where the maker's precision is known or the row is the maker's own: the 81 reference prices on a reseller's declared row measured against nothing say "at a declared precision that is not 4-bit (fp8)" and that the maker's own is not known; a deal-only price is "a fallback", not "the cheapest serving offer" and not "at the maker's precision"), labels a deal-only price with every reason it is not a reference price, "precision not disclosed" among them, and shows the deal as "Cheaper right now" with its reasons; the board says the same in its price card; /methodology/ states the rule in 19 locales.
  • The MCP package and the companion name the seller. get_model returns priceBasis, priceRow and deal beside the price, each with its reasons, priceRow.discount where the price is a promoted row's standard rate, and provenance.held for a held one; the companion's per-model facts carry the same. Before, both gave a reseller's rate beside the maker's name and called it vendor-published.
  • The wire does not publish the 64 balanced moves as vendor moves: data/price-restatements.json restates them under §96 for the pair that crosses 2026-09-30. GPT-5.6 Sol Pro is not one of them: its price is its model-level price again, so its step from the 4.00 the 2026-09-30 snapshot holds to 8.00 will read on the wire as a price rise. It is one, of OpenRouter's model-level price on 2026-09-29, published late because A7 held it. That list is A7's, and the release that ships the rule refreshes every row. The build now refuses a list that is not the data's: scripts/build-wire.mjs and scripts/verify-invariants.mjs require the §96 entries to be exactly the published models whose stored balanced price is not their model-level one, until the first snapshot after 2026-09-30 is recorded as published. node scripts/measure-reference-price.mjs --write-restatements regenerates it.
  • §92's "No price, rank, verdict or count changes" described §92 and stands as that; the product question it left open is this entry.
  • The upstream watch does not see these prices. scripts/watch-upstream.mjs and its AUTO and HUMAN lanes diff OpenRouter's models feed, the model-level price. 66 published rate cards are no longer that price. When a run refreshes, it re-fetches every endpoint row, so a reseller's row or a seller going down moves a published price the watch never judged. When the models feed has not changed, the workflow refreshes nothing, and a published price stays where it was however far its row has moved. The gate on the published price is scripts/published-diff.mjs, in the daily-refresh work (work/followups-b-20261001), not in this change. AGENTS.md says so beside the lane table.
  • Not done here, recorded so that nobody assumes it:
    • Snapshots and the price-at-floor series store prices without their basis. The step between 2026-09-30 and the first day under the rule reads as a market move in /data/series/price-at-floor.json; only the wire and /now/ say it is a restatement.
    • Weights records: gpt-oss-20b has none, so its 4-bit rows are excluded (no price moves on this data), and prism-ml/ternary-bonsai-2-27b, whose only row is int4, is deal-only for "4-bit" although the model is natively ternary. Each needs a cited weights record, and the second a precision class the rule does not have.
    • pricing.notes is kept when the price is the maker's own row and that row is not on promotion (39 of the 58 records with a note; 41 before R124, where this entry said 40) and dropped otherwise, licence statements included. Licence prose belongs outside pricing.
    • openWeights is not recorded for 338 of the 444 published models, so R124's closed-weight clause reaches almost nothing yet. Each record needs a cite.
    • A promoted row is published at a rate nobody can pay today (Mercury 2.5: $0.20 / $0.75 against $0.04 / $0.15, its only row). Accepted for A8. The model page says so in one sentence, with the rates payable today and the percentage ("Cheaper right now: $0.040 in · $0.150 out per million tokens at Inception, a promotion of 80%. The price on this page is Inception's standard rate …"); the board row shows the standard rate alone. A promotion that has held 30 days or more may graduate to the price; that needs a history per row and is deferred (spec, "Not in scope").
    • One deal per model, the price row's own promotion first, so a cheaper flex row is not shown on 5 models. Accepted for A8; a list of deals is deferred to the next release.

Revised 2026-10-06. This entry was written on 2026-10-05 and reviewed before any release, so no page carried the figures it replaces. What changed, and why:

  • "59 rate sets, 57 balanced prices: 23 up … reference 329 · model-level 104 · 76 deals", "31 prices carry the label" and "191 of the 222 undeclared rows" counted GPT-5.6 Sol Pro at Azure's $4 / $20. It is held, so each count is one lower or, for model-level, one higher. "GPT-5.6 Sol and Sol Pro +100%" is Sol only.
  • "a closed model's own vendor" said of the 191 undeclared maker rows was not measured. They are the maker's own rows; whether each model is closed was never counted.
  • "written from the spec and importing none of the rule" was true of the imports and not of the readings; see "The oracle".
  • "The release that ships the rule regenerates it … before its first snapshot" was a step in a sentence. The check found a stale entry the day it was written: Sol Pro, after the hold was fixed.
  • The delivery-tier test read only tag segments after the seller, so sambanova-turbo was not a tier; it was excluded on this data only because it is also a 25% promotion. A serving, undiscounted $0 / $0 row for a paid model would have become its price and rendered it free; the pipeline now stops on one. Neither changes a price here.

Revised 2026-10-06, ruling R121. The rule itself changed, on the controller's ruling, still before any release. What this entry said the same morning, and what it says now:

  • "unknown precision passes" and "D3, 'precision not disclosed' labels a reseller's undeclared row … 30 prices carry the label, 14 of them rated". A reseller's undeclared row no longer passes. No ranked reference price carries a label; the precisionNotDisclosed field is removed from priceRow and deal, the catalogue's pn flag with it, and precision-not-disclosed is a reason beside the other three.
  • "58 rate sets, 56 balanced prices: 22 up, 34 down; 39 rated … reference 328 · deal-only 11 · model-level 105 · 75 deals", "frontier 10 → 9, dominated 134 of 143, median overpay 7.3×", "Twelve models tie", "31 reference prices" under a context rule, "56 balanced moves" on the wire and "58 published rate cards". Each is re-measured above; "What R121 changed" has the difference and its cause.
  • "DeepSeek V4 Pro … Reference $1.044 / $2.088 / $0.2088, precision not disclosed, +33%" and "GLM-5.3 … Reference $0.37 / $1.14 / $0.067, precision not disclosed, −74%": both were a reseller's undeclared row. They are DeepInfra fp8 at $1.30 / $2.60 (+66%) and Morph fp8 at $0.1428 / $2.6943 (−64%).
  • "GPT-5.6 Sol +100% (OpenAI's promotion to Azure's $4 / $20)" and "This leaves the two siblings a factor of two apart again". Azure declares no precision, so Sol is deal-only at OpenAI's promoted $2 / $10 and does not move from its model-level price.
  • The model page's "Not like-for-like" under a deal said more than is known of an undeclared row. It reads "It does not pass the like-for-like test" in 19 locales.

Revised 2026-10-06, ruling R122. GPT-5.6 Sol Pro's hold is retired; "The Sol Pro hold" above has the two models' rows. What this entry said earlier the same day:

  • "A held price is not repriced. data/upstream-holds.json keeps GPT-5.6 Sol Pro's pricing at the live publication's, $2 / $10 … its basis is model-level, the 105th" and "Retiring the hold … is a decision for the release that ships the rule, not made here". The decision is made. The mechanism stays, with no price under it.
  • "reference 310 · deal-only 29 · model-level 105 · 76 deals", "57 rate sets, 55 balanced prices: 22 up, 33 down", "model-level 104, deal-only 29, held 1", "55 balanced moves" and "57 published rate cards": each is one higher or lower by GPT-5.6 Sol Pro. No published price and no headline figure moves.
  • R122's reason, that the hold "now creates the gap it was meant to prevent", was true of the rule between the review and R121 (Sol at Azure's $4 / $20, Sol Pro held at $2 / $10). After R121 the two already agreed at $2 / $10; retiring the hold makes them agree for the same reason.

Revised 2026-10-07, ruling R124. The rule changed again, still before any release. What this entry said the day before, and what it says now:

  • "with no promotion (pricing.discount not above 0)" and "The cheapest excluded serving row that is cheaper on the blend is published as the deal". A promoted row is ranked at its standard rate; the deal is the price row's own promoted price where it has one.
  • "A row that declares no precision passes only when it is the maker's own". A reseller's also passes for a model whose weights are recorded closed. No price turns on it on this data.
  • "D1 … The standard row, $0.75 / $3.75, is their deal-only price; the flex row is their deal", "Gemini 3.8 Flash ($1.50, deal-only on Google's promoted row) … is the one frontier price that rests on a fallback". Both Gemini Flash models are reference prices at $1.50 / $7.50.
  • "58 rate sets, 56 balanced prices: 22 up, 34 down; 37 of them rated … reference 310 · deal-only 30 · model-level 104 · 77 deals", "187 reference prices … and 24 deal-only prices, 6 of them the maker's own promoted row and 18 a reseller's", "Four models tie", "30 published prices" under a context rule, "56 balanced moves" on the wire and "58 published rate cards", here and in AGENTS.md. Each is re-measured above.
  • "published price $2 / $10 / $0.20 cached … deal-only, OpenAI openai, promotion 50%" for GPT-5.6 Sol and Sol Pro, and "it is the 56th §96 restatement" of Sol Pro. Both are $4 / $20, and Sol Pro is no restatement.
  • "The largest moves … DeepSeek V4 Flash Vision Exp +104% (DeepInfra's 51% promotion to GMICloud fp8), Llama 3.3 70B +81%, DeepSeek V4 Flash 0731 −73%, Qwen3 235B A22B 2507 +71%". The first is the same +104% at DeepInfra's own standard rate; the other three moved with R128.
  • The price card and the model page said "the cheapest offer that is serving, at standard delivery, with no promotion and at the maker's precision" and, of a deal-only price, "the cheapest serving offer; none passes the like-for-like test". Neither is true of a standard rate or of a ladder: 19 locales now say "at its standard rate" and "a fallback".

Revised 2026-10-07, ruling R128. "Serving" was "status 0 or above". It is that, or status −2 with at least 95% uptime over the last day; the ruling's first wording, any status below 0, is measured above and was not ratified. The same day this entry said the narrower rule was "not the ruling as given … a deviation awaiting one", and that Gemini 3.8 Flash "is named at Google's google-vertex/global row": the rule is ratified at −2, and a tie now names the row that read 0. "Four models … have endpoint rows and none serves" is three, "Novita bf16 is not serving" in the Llama 3.3 70B hand check is no longer so, and the figures in the table are after both rulings.

Revised 2026-10-07, ruling R131 and the re-review of 1418ea8. Still before any release:

  • "serving (status 0 or above, or status −2 … with a last-day uptime of at least 95%)". A row whose last day was under 80% no longer serves, whatever it reads. Qwen3 Coder was "+42%" on the model-level price, at Novita's fp8 row; it is +311%, at Alibaba's own. "18 prices change, 12 up and 6 down" for the round is 19 and 13.
  • N1. Under a price taken at a promoted row's standard rate the model page said "Cheaper right now: $0.068/M at Inception — promotion 80%. It does not pass the like-for-like test, so it never sets a rank." That offer does pass, and it sets the rank, on 11 reference prices; on 3 deal-only prices it sets the fallback that ranks. Nothing said in words that the price on the page is 5× what Mercury 2.5 costs today. The page now says, for those 14: "Cheaper right now: $0.040 in · $0.150 out per million tokens at Inception, a promotion of 80%. The price on this page is Inception's standard rate: the ranking uses the standard rate, never a promotion."
  • N2. "at the maker's precision" was printed on all 320 reference prices. On 273 the reference precision is unknown; 192 of those are the maker's own row, where it holds by definition, and 81 are a reseller's fp8 (55), bf16 (22) or fp16 (4) row. Rule 5 passes such a row for declaring a precision that is not 4-bit; it does not show that precision to be the maker's. Those 81 now read "at a declared precision that is not 4-bit (fp8) … The maker's own precision is not known", on the model page and in the board's price card (catalogue pq), and scripts/verify-invariants.mjs holds the card to the record. The fix round before had rewritten this sentence and kept the phrase.
  • F1, found by the follow-up check of 0a2b051: two definitions of "serving". The "like-for-like at the best declared precision" note (model page, and the board's label) is computed by the offer summary in src/lib/publication-policy.mjs, which had its own test, "status not below 0". R128 and R131 changed rule 1 and left it behind, so the note and the price on the same page disagreed on 7 models:
    • Qwen3 Coder: the note named Novita fp8 at $0.38, the row R131 had removed (26.0% of the day). It names no row now: no declared row of standard delivery serves.
    • Llama 3.3 70B: CoreWeave fp16 at $0.71 → Novita bf16 at $0.135, the row that sets the price.
    • Qwen3 235B A22B 2507: Parasail fp8 at $0.14 → Novita fp8 at $0.09, the price row.
    • Qwen3 235B A22B Thinking 2507: Venice fp8 at $0.45 → Novita fp8 at $0.30.
    • Qwen3.6 27B: SiliconFlow fp8 → Chutes fp8, both $0.30; Chutes is the price row.
    • MiMo V2.5 and DeepSeek Chat had no note and have one: Xiaomi fp8 at $0.14, the price row, and DeepInfra fp4 at $0.32.

    Four of the seven date from R128 in 1418ea8, where this entry's "Every reader … follows" was already too wide. isServing now lives in src/lib/pricing/row-terms.mjs and both read it. src/lib/publication-policy.test.mjs asks each through its own entry point, over a grid of statuses and day figures and over the 432 committed rows the note can name (25 of them not serving), and fails if they part. No price, count or headline figure changes; the catalogue's obq and obs change for those seven.

  • And where no row qualifies, the page made one up. 86 published models have offers that differ in precision; 84 have a best-precision row and 2 have none, Qwen3 Coder and Llama 3.1 70B. For those the model page fell back to the first listed row and printed the same sentence: "Like-for-like at the best declared precision is $0.220/M from Google" for Qwen3 Coder, a row that declares no precision, and DeepInfra's turbo tier at $0.40 for Llama 3.1 70B. The second was already so on this branch; under the old serving test MiMo V2.5 and DeepSeek Chat had the same fallback. The note, and the FAQ answer built from it, are now printed only where a row qualifies (likeForLikeNote in src/lib/board/cards.mjs). The board's badge never had the fallback: with no row its premium is 1 and it is not shown.
  • Seen and not changed. The note's first sentence, "Headline cheapest is a lower precision", is printed whenever precisions differ. On 25 of the 84 the best-precision row is the first listed row itself, and where the cheapest row declares nothing its precision is not known to be lower. The sentence predates §96 and needs its own wording ruling.

92. Correction — OpenRouter's model-level price is not always its cheapest endpoint

Logged

Live (release A7, publication a35c5e38…) says on /batch/, in all 19 locales (18 print the English, German its translation): "The other 48 have competing sellers, so the standard rate is the cheapest of the market and the ratio measures competition rather than delivery". The standard rate is OpenRouter's model-level price. §47 said the same thing on 2026-09-11: "OpenRouter's model-level price already is its cheapest endpoint".

It is not, and not only because of rows that are not like-for-like.

  • What §47 measured. All 422 priced rows carried the upstream model-level pricing.prompt; 0 differed. That shows the published price is the model-level price, which still holds. It never compared the model-level price with the endpoint rows. "Is its cheapest endpoint" was an inference.
  • What was measured now. scripts/measure-model-level-price.mjs, on A7's committed data/models.json and data/offers.json (endpoints fetched 2026-09-30). Status is not in offers.json. It comes from the raw endpoint cache those offers were built from, and every one of the rows below matched a raw row on tag and rates. Rates are compared at four decimals on both sides, the precision offers.json stores.
    • The population is 344 models: standard delivery, publishable, priced per token, with endpoint rows. They have 1,303 endpoint rows between them. Every one of the 344 has a row at its model-level price.
    • 277 rows are cheaper on input, output or cache read, on 98 models. On input or output it is 92 models. 82 models have a row cheaper on input.
    • Each row is classified by the first reason that applies:

      | reason | rows | |---|---| | cheaper on the cache read only | 30 | | status −2, not serving | 27 | | a delivery tier (flex and the like) | 45 | | a live promotion (pricing.discount above 0) | 44 | | a declared precision no row at the model-level price declares | 27 | | no declared precision to compare (unknown) | 62 | | none of these | 42 |

    • 42 rows on 24 of the 344 models (7%) are like-for-like and cheaper. Each declares the precision of a row at the model-level price, is standard delivery, carries no promotion, and is serving. 40 of them (23 models) are cheaper on the balanced blend. 34 of them (20 models) are dearer on neither rate.
    • Examples. MiniMax M3: DeepInfra fp8 at $0.28 in / $1.10 out against $0.30 / $1.20, which six fp8 sellers charge. DeepSeek V3.2: AtlasCloud fp8 at $0.1344 / $0.2016 against Baidu fp8's $0.28 / $0.42. Mistral Nemo: DekaLLM fp8 at $0.018 in against $0.019.
    • 13 of the 27 precision rows (5 models) are cheaper at a higher declared precision than every row at the model-level price. That is not like-for-like either, but it is in the buyer's favour.
    • A review counted another seller cheaper on 50 of 343 rows on an earlier day's offers.json. That count was not reproduced. On this data the nearest reading is 82 of 344 models with a row cheaper on input.
  • On /batch/. Of the 48 cross-seller pairs, 4 have a like-for-like seller below the standard rate:
    • GLM-5.3, 6 rows;
    • GLM-5.3 Flash, 5;
    • Kimi K3, 3;
    • gpt-oss-120b, 1. DekaLLM bf16 charges $0.03 in against $0.037. It is $0.18 out against $0.17, and cheaper on the blend.
  • Now.
    • /batch/ says the standard rate "is OpenRouter's model-level price, which any of those sellers can set and which is not always the cheapest offer, so the ratio measures competition as well as delivery", in English and German.
    • The same claim is corrected in the comments of src/lib/run/batch.mjs and its test.
    • §47 keeps its words and points here.
  • What moves. One sentence on /batch/ in 19 locales. No price, rank, verdict or count changes: the site prices every model at its model-level price, as before. Whether a cheaper like-for-like seller should be shown next to it is a product question, not this correction.

Pointer, 2026-10-05: §96 answers that question. From release A8 the published price is the reference price, the cheapest like-for-like endpoint row, and the model-level price travels as modelLevelPricing. This section's measurement stands as made on A7.

91. Correction — the wire rounded a price down that the board rounds up, and linked a page that is gone

Logged

Live (release A7, publication a35c5e38…) published, in /data/wire/events.json and /wire.rss, three price summaries for Qwen3.8 27B (qwen/qwen3.8-27b) that print its balanced price as $1.06 per million:

  • price-change-2026-09-03-qwen-qwen3-8-27b: "$0.9562/M on 2026-09-02, $1.06/M on 2026-09-03";
  • price-change-2026-09-15-qwen-qwen3-8-27b: "$1.06/M on 2026-09-11, $0.798/M on 2026-09-15";
  • price-change-2026-09-22-qwen-qwen3-8-27b: "$0.798/M on 2026-09-15, $1.06/M on 2026-09-22".

The snapshots record $1.065. The site rounds a price once, half-up, on the decimal (A4 review L5, R109), so $1.065 is $1.07, which is what the board prints for it. The wire's usd() used toFixed(2), which rounds the binary expansion: 1.065 is 1.06499… as a float. One price, two figures: the defect R109 removed everywhere else.

  • What else was checked. The 158 price and frontier events in the feed print 313 prices and 131 percentages. These three prices are the only figures where toFixed and half-up disagree. No title changes: every percentage in a title already agreed. Titles now go through the shared rounding too (roundHalfUp, src/lib/format/usd.mjs), so a move the float carries as 5.0499…% prints 5.1%, not 5%.
  • The link. price-change-2026-09-02-mistralai-devstral-2512 pointed at /models/mistralai__devstral-2512/, which answers 404: the model has left the catalogue. The days that move is read from, 2026-08-30 and 2026-09-02, are not snapshots the citation fence publishes, so no archived copy is served to point at instead. A later published day that still lists the model (2026-09-15, 09-22) would show its price then, not the move. The event now carries no link and says why ("No link: … is no longer in the catalogue …"). The rule covers every model event, archived context events included; this is the only one today.
  • What moves. Three summaries, $1.06 to $1.07, and one link replaced by a note. No price, rank or verdict changes. scripts/build-wire.test.mjs holds the three binary edges ($1.065, $0.30005, a 5.05% move) and the removed-model event.

90. Correction — the price spread was a ratio of input rates, while every other price on the site is the balanced blend

Logged

Live (release A6, publication 1da61ea4…) printed a price spread of 8,824×.

  • /methodology/, in all 19 locales: "the cheapest model in this market is 8,824× cheaper than the dearest".
  • /llms.txt: "the spread between the cheapest and most expensive model is about 8,824x".
  • The catalogue's stats.priceSpread, which the page companion also reads.

The sentence compares what models cost. The figure did not measure that.

  • What it was. The dearest input rate over the cheapest, across every published row priced above $0 (419 rows, batch delivery included). On live that was OpenAI o1-pro's $150 per million input tokens over IBM Granite 4.0 H Micro's $0.017: 8,824×.
  • What every other price on the site is. The balanced blend: three input tokens to one output token, no cache. The frontier, the dominance verdicts and the overpay multiple take it from effectivePricePerMillion (src/lib/ranking.mjs), and the board computes the same blend (priceOf, src/lib/data/store.svelte.ts). An input rate alone is not what a model costs: it leaves out output and separately billed reasoning.
  • What live should have printed: 12,069×. On the balanced blend, over the 344 priced standard-delivery models, the dearest is o1-pro at $262.50 per million and the cheapest is Mistral Nemo at $0.02175. Live's 8,824× understated that by 27%.
  • Why it moved in the other direction in the next release. In the 2026-09-30 refresh, DeepSeek V4 Flash 0731's input rate fell from $0.018 to $0.01 while its output rate rose from $0.32 to $1.28. Its balanced price rose 250%, from $0.0935 to $0.3275. Its input rate became the cheapest, and the input ratio became $150 / $0.01 = 15,000×: 24% above the balanced 12,069×, which did not move. An input ratio can move against the prices it claims to describe.
  • How it happened. scripts/build-payload.mjs has computed the spread from pricing.input since the first pipeline commit (a3123fa, 2026-08-24). §1 of this log measured it on input and said so; the published sentence never did.
  • Now. stats.priceSpread is the balanced price the frontier uses, from the same function, over priced standard-delivery models. Publication-denied providers are excluded, and so are models billed per unit (§88), whose token rates are null. On the 2026-09-30 data it is 12,069×. scripts/build-payload.test.mjs holds a catalogue where the two disagree: 1,000× on input over the same models, 250× on the blend. It also carries a batch row, a free row, a per-unit row and a denied row, each cheaper than the rest, and none of them enters the spread. The old formula gives 10,000× on it.
  • What moves. Only the spread: 8,824× on live, 12,069× in this release. No other figure uses it.

89. Correction — models within error of each other were said to share a rank

Logged

Five published sentences said or implied that models the measurement cannot separate share a rank. The board's rank card and rank cell said the same thing with a count:

wherewhat it saidsince
/methodology/, ranks, 19 locales"Two models the measurement cannot separate share a rank."2026-08-24 (26fd25e)
/methodology/, the same section, 19 localesLMArena's Rank (UB) "assigns the same rank to models that are not statistically separable"2026-08-24 (26fd25e)
the homepage board's rank line, 19 locales"Ranks are significance ranks: tied models share one."2026-09-24 (673097a)
/llms.txt"Models the benchmark cannot separate share a rank."2026-08-24 (922489f)
/methodology/, the paragraph below, 19 locales"The 95% confidence interval on a difference between two arena scores is about ±10.57 points. On this catalogue, 125 of 134 adjacent pairs (93%) fall inside it. So 135 rendered positions carry only 69 genuinely distinct ranks"2026-08-24 (26fd25e); arena wording since 2026-08-26 (32a3adb)
the board's rank card, on a shared rank"Tied for #N with K others", where K and the names listed were every model within error, at any rank2026-09-24 (854c2f5)
the board's rank cell, its screen-reader label"Rank N, tied with K others", with the same K2026-09-24 (a2bdae1)
  • What is true. A model's rank is one more than the number of models that beat it by more than both margins of error combined (sigRank in src/lib/data/store.svelte.ts). Models with the same count share a rank. The count is per model and does not chain. Two models within error of each other can still hold different ranks, because each is compared with the whole board. LMArena's Rank (UB) is the same rule.
  • Measured with the production store, on the default board of the live A6 catalogue:
    • Which catalogue: catalogue.json with sha256 8251392d…, byte-identical to the one served on undominated.ai.
    • Which board: the general board, 135 rated standard models, each with a published interval.
    • "Within error" means that the two intervals overlap, |Δq| ≤ acw_a + acw_b. 809 pairs are within error. 697 of them hold different ranks, and 112 share one. The sentence was false for 86% of the pairs it describes.
      • Under the stricter reading, each score inside the other's interval (|Δq| ≤ min(acw_a, acw_b)), it is still false: 266 of those 345 pairs hold different ranks.
    • Qwen2.5 72B Instruct (1268.9 ± 4.06) is rank 122, and Mistral Large 2407 (1266.1 ± 3.9) is rank 123. They are 2.8 apart, inside 7.96.
    • Gemini 3.8 Flash (1494.7 ± 8.53) is rank 1. Claude Fable 5 (1492.6 ± 4.76), 2.1 below it, is rank 4.
      • Three models clear Fable 5's interval. Claude Opus 4.6 (1503 ± 3.49), for one, is 10.4 above it, beyond 8.25.
      • No model clears Gemini 3.8 Flash's wider interval.
    • The widest rank gap within error is 20 ranks, on four pairs.
      • One is GPT-6 Astra (1443.7 ± 11.64) at rank 21 and R1 0528 (1427.5 ± 5.67) at rank 41. They are 16.2 apart, inside 17.31.
      • The other three pair DeepSeek V3.1 Terminus (rank 50) with MiniMax M2.7, Step 3.5 Flash and Qwen3 VL 235B A22B Thinking (rank 70).
    • The widest score gap within error at different ranks is 19.7: Solar Pro 4 (1377.3 ± 12.18) at rank 83 and Mercury 2 (1357.6 ± 10.56) at rank 89, inside 22.74.
    • Of the 134 pairs adjacent in score (score ties broken by rank), 80 are within error and hold different ranks.
  • The converse always holds. Two models that share a rank are always within error of each other.
    • Why: if A clearly beats B, every model that clearly beats A also clearly beats B, and so does A. B's count is then larger, and its rank lower.
    • On this data, all 112 same-rank pairs are within error.
    • A test checks it on 300 random boards and on the live one.
    • This is why "tied" is true of the other holders of a rank, and of no one else.
    • src/lib/board/tiers.mjs and the gotcha tier-sentence-needs-all-pairs-tied said that two holders of one rank can be apart. Under runQuery's rule they cannot, so every tier is all-tied.
  • The rank card counted the wrong population.
    • K was tiedWith: every model within error of this one, at any rank. The card listed the same models.
    • Of the 102 models whose card had the title, 90 counted and named a model of another rank. The rank cells' labels were wrong on the same 90 of 102.
    • Claude Fable 5, at #4 with three other holders, read "Tied for #4 with 4 others". The first model it named was Gemini 3.8 Flash, at #1.
    • Gemini 3.8 Flash read "Tied for #1 with 7 others", and four of the seven are at #4.
    • Live, the four rank-4 cells read "Rank 4, tied with" 4, 6, 4 and 8 "others". Four models hold rank 4.
  • The same mistake was corrected once before. §57 (2026-09-23) removed it from the modality introduction, which said every overlapping pair shares a published rank (UB). It survived on the methodology page, on the board and in llms.txt. docs/LICENSING-AND-CLAIMS.md §3.2 glossed Arena's rule as "statistically indistinguishable models tie by construction", the same misreading. That gloss is corrected too.
  • Now.
    • /methodology/ says: "A model's rank is one more than the number of models that beat it by more than both margins of error combined. Models with the same count share a rank. Two models within error of each other can still hold different ranks, because each is compared with the whole board."
    • Its Rank (UB) sentence now says the rule gives each model one more than the number of models statistically better than it.
    • The board's rank line says "A model's rank is 1 + the number of models that clearly beat it."
    • Those sentences are in 19 locales. The 18 translations are model translations, checked by two blind back-translation reads.
    • /llms.txt says what the methodology page says.
    • The paragraph below it drew the same inference with "So". The 69 ranks do not follow from the 125 pairs: 80 to 82 of them hold different ranks, depending on how six exact score ties are ordered. Its ±10.57 was not a 95% confidence interval on a difference. It is the mean of acw_a + acw_b, a test at least as strict. The paragraph now states the test, the two counts and the rule, with no "so":
      • "Two arena scores count as separated only when they differ by more than the sum of the two models' 95% half-widths. That sum averages 10.57 points on this board, and the test is at least as strict as a 95% interval on the difference."
      • "125 of 134 adjacent pairs (93%) fall inside their own sum."
      • "Counting each model's rank as one more than the number of models that clearly beat it, the 135 rendered positions hold 69 distinct ranks, and 4 models are joint first."
      • It is in 19 locales. Since 32a3adb dropped their older translations, 17 of them had shown the English.
    • /press/ asked writers never to say "these models are tied", yet the board says "Tied for #4". It now says what that means: the models share rank 4, and the same number of models clearly beat each of them. /press/ is in English and German.
    • The rank card and the rank cell's label count and name only the other holders of the rank. On this catalogue that took them from 90 of 102 wrong to 0, for the titles, the names and the labels alike.
      • Fable 5 reads "Tied for #4 with 3 others", and names Gemini 3.7 Flash, Claude Opus 4.7 and Muse Spark 1.3.
      • Gemini 3.8 Flash reads "Tied for #1 with 3 others".
      • Where there is one other holder, the label reads the card's singular title, "Tied for #122 with 1 other". The label has no singular form, and "tied with 1 others" would otherwise be on 38 cells; on live it was on 3.
      • A rank held alone still shows the bare rank and names the models within error of it, as before.
    • src/lib/data-query.test.mjs has three tests:
      • A three-model board on which two models within error hold ranks 1 and 2. It fails if any English string, or the llms.txt text, says such models share a rank.
      • The converse, on 300 random boards.
      • Every rank card and label on the live board.
  • Known, and not fixed here. /significance/ lists models in score order under its "Sig. rank" column. Ranks do not chain, so on this catalogue the rank number falls 20 times as you read down the column. That is the order problem of non-chaining-ranks-invert-in-score-order, on a page that does not sort by rank first.

88. Correction — Lyria 3 Pro and Lyria 3 Clip were published as free; they are billed per song and per clip

Logged

Live (release A6, publication 1da61ea4…) published two Google models as free: Lyria 3 Pro Preview (google/lyria-3-pro-preview) and Lyria 3 Clip Preview (google/lyria-3-clip-preview).

  • /free/ listed both among its "19 rows are free in this catalogue".
  • The board printed "free" for each in all 19 locales, and /models/ printed "free / free".
  • Each model page was titled "free route, unrated".
  • The catalogue and the per-model files carried isFree and fr 1. The marketplace fields carried a Google AI Studio offer at $0 per token.

Neither model is free.

  • What the sources say.
    • OpenRouter's feed lists both with token prices "0" and "0".
    • Its own descriptions say "Full-length songs are priced at $0.08 per song." and "30 second duration clips are priced at $0.04 per clip."
    • Google's Gemini API pricing page (https://ai.google.dev/gemini-api/docs/pricing, fetched 2026-09-30 15:05Z) reads: "Free Tier Paid Tier, per request in USD Lyria 3 Clip Preview (30s) Not available $0.04 per song Lyria 3 Pro Preview (Full Song) Not available $0.08 per song". The free tier is "Not available".
  • How it happened. scripts/build-dataset.mjs set isFree whenever both token prices were 0. A per-song price is not in the token fields, so its absence read as $0, and $0 read as free: absence rendered as a favourable verdict. Both rows have carried this since the first pipeline commit (a3123fa, 2026-08-24).
  • Now.
    • data/per-unit-prices.json records each model with the verbatim quotes, both sources and its unit. It carries no rate. The per-song price is never converted to a per-token one, stored in a price field or ranked.
    • The pipeline publishes both as not free, with no token rate (pricingKind: 'per-unit'), and without the $0 offer.
    • The build fails if either model gains a per-token price or leaves the feed, or if its description stops carrying the quote.
    • The board and /models/ print "not priced per token", translated in 18 locales.
    • Each model page says "Priced per song, not per token" ("per clip" for Clip), quoting both sources. Its title and description say the same.
    • Both leave /free/.
    • Neither is rated, so neither was on the frontier or in a dominance verdict. With no token price they cannot enter either at $0.
  • What moves.
    • /free/ goes from 19 rows to 17, and the free-flagged rows from 19 to 17.
    • Priced models in the catalogue (stats.priced) go from 438 to 436.
    • /openrouter-vs-direct/ counts models with offers 346 → 344, and lab-only models 157 → 155. The lab-only share stays at 45%.
    • No figure of the ranking, the frontier or the dominance verdicts changes.
  • The other $0/$0 rows. 17 of the 19 published rows priced $0 in and out remain free.
    • 16 are the :free delivery rows of models that also have a paid row.
    • The 17th is openrouter/free, OpenRouter's Free Models Router. Its description reads "The simplest way to get free inference. openrouter/free is a router that selects free models at random from the models available on OpenRouter."
    • None of the 17 descriptions names a per-unit price.

87. Correction — Kimi K3 was published under the Modified MIT licence, and its licence is its own

Logged

Live (release A6, publication 1da61ea4…) labels Kimi K3 (moonshotai/kimi-k3) and its batch row modified-mit: on /open-weights/, on both model pages, on the compare pages that show its licence, and in the data files. §84 left this label under review. The licence file settles it, and the label was wrong.

  • What the file says. Kimi K3's LICENSE is titled "Kimi K3 License" (https://huggingface.co/moonshotai/Kimi-K3/raw/main/LICENSE, fetched 2026-09-30 14:42Z, sha256 20c797ce…). It is not the Modified MIT License, which Kimi K2.7 Code's repository carries (sha256 9cb73476…).
  • The condition Modified MIT does not have. "If the Licensee or any of its affiliates operates a Model as a Service business, and the aggregate revenue of the Licensee and its affiliates exceeds 20 million US dollars (or the equivalent in other currencies) in total over any consecutive 12 months, the Licensee must enter into a separate agreement with Moonshot AI before using the Software or its derivative works for any commercial purpose."
  • The rest of the file.
    • Its display clause names "Kimi K3", above 100 million monthly active users or US$20 million monthly revenue.
    • Internal use, and use through Moonshot AI's own products or certified inference partners, are exempt from both clauses.
    • The repository's metadata names license_name: kimi-k3. The card's tag is the licence's name, as it was for Qwen3.8-Max (§84).
  • Now. Both rows read kimi-k3, citing the file's URL, fetch time and sha256 (data/models-handresearched.json). Open weights stays true, and open-weight models stay at 25.
  • What moves.
    • /open-weights/ prints kimi-k3 in Kimi K3's licence column, and Kimi K2.7 Code keeps modified-mit.
    • The two Kimi K3 model pages and their data files print kimi-k3. So do the two compare pages that show its recorded licence, against Kimi K2.6 and DeepSeek V4.1 Flash.
    • The payload's licence facet gains kimi-k3 and keeps modified-mit.
    • The price note "Licence not independently verified" is replaced by the terms above.
  • The other labels. Checked the same way the same day:
    • 22 standard models carry an open licence label, and 20 of them name a Hugging Face repository in the feed.
    • Of those 20, 19 carry metadata that agrees with the label. Llama 4's two say llama4, Meta's name for the same licence; they and Command A are gated.
    • Command A Plus and Mistral Large 3 name no repository.
  • How it was found. The finding came from building a licences page (menu Task 5, fetched 13:24Z, the same sha256). It was re-fetched for this release.

86. Correction — task scores were credited to LMArena, and ranked out of a count they exceed

Logged

Three published statements named the wrong source for the per-task scores:

wherewhat it saidsince
the board's quality card under a task lens"Source: LMArena, CC BY 4.0."2026-09-24 (673097a)
/methodology/, licensing"Every capability score on this site is LMArena Elo, used under CC BY 4.0 from the official dataset."2026-08-26 (32a3adb)
/methodology/, how quality is measured"Published scores are LMArena Elo from human preference, on its general board and its document board."2026-08-26
  • What is true. The task lens and each model page's "Best rankings by task" show Design Arena's per-task Elo, read from OpenRouter's models feed (benchmarks.design_arena), and have since the first prototype (b153b75, 2026-08-24). The general and document scores are LMArena's. No page credited Design Arena or linked to it.
  • Its licence. Design Arena's API documentation: "Data from the Design Arena API is free to use for personal and commercial projects", and a public display must "Credit Design Arena as the source" and "Provide a visible link to designarena.ai". We read the figures through OpenRouter, not that API; docs/LICENSING-AND-CLAIMS.md §6 records the gap.
  • Now. The task lens's rank line links "Task scores: Design Arena, via OpenRouter" to designarena.ai and its card names Design Arena. The 108 model pages with task rankings credit it under the list, with a link. /methodology/ credits and links it in its task section and in its licensing list, in 19 locales, and its two sentences say general scores.
  • A rank out of the wrong count. Each model page printed a task rank as "#{rank} of {count}". The rank is the feed's, and it counts models the feed does not list: game development lists 109 models with ranks up to #127 and 18 numbers missing. The count was the models the feed does list. 16 of the 653 cells on the build of 687e548, on 8 model pages, read as impossible: GPT-5.3 Codex "#44 of 42" in mobile apps, Gemini 3 Pro Image Preview "#3 of 1" in image editing, three image models "#8" to "#11" "of 5" in graphic design, image and logo. The count is gone, and the credit line says whose rank it is.
  • Caught before it shipped. A first draft of the credit said the task ranks "count the models OpenRouter lists". That was the same mistake in prose, and it did not ship.
  • Not yet carried. Two downloads are the pipeline's and change with the next data release. catalogue.json holds the task Elo in tasks, and its attribution names only Arena. The 444 per-model files under /data/models/ carry taskStandings, each with the feed's rank and the listed count as of, and no Design Arena credit.

85. Correction — workload prices were rounded twice before they were displayed

Logged

Pages built on the server priced a workload with effectivePricePerMillion (src/lib/ranking.mjs) and eff (scripts/build-payload.mjs). Both returned the blend already rounded to four decimals, and the page then rounded that to its display precision. MiMo-V2.5-Pro's code-generation price is $0.661488/M. It was stored as 0.6615 and printed $0.662. Live (release 20260930-033430-834a81b1a6f0) also formatted with Number#toFixed, which rounds the binary value, not the decimal.

Each figure below names what it counts. "Displayed" means a figure printed on a page. "Computed" means every value the code works out, whether or not a page shows it.

  • What readers saw on /alternatives/ (126 pages, each figure checked against exact decimal arithmetic on the published rates):

    | displayed | live, on its own data | this build | |---|---|---| | the model's own price per workload | 40 of 625 wrong | 0 of 630 | | a candidate's price | 12 of 417 wrong | 0 of 416 | | savings chips ("Balanced −57%") | 22 of 5,417 wrong; 65 are exact halves | 0 of 5,400; 64 are exact halves |

    /check/ changed four displayed savings on three pages. Each is an exact half, now rounded up.

  • Computed prices. These are the 1,735 workload prices at prompt 0: 347 published standard models, five workloads each.
    • Live's method gets 90 of them wrong on this data.
    • The A4 fix round's L5 (e425d59) rounded half-up on the decimal, but still from the four-decimal value. That fixed 86 of the 90, broke 12 that toFixed had got right by accident, and left 16 wrong (A4 re-review, R109).
    • Now 0 are wrong, and no context-tier rung is either. The engine returns the exact decimal (twelve significant digits keep every digit a six-decimal rate and two-decimal weights make). The payload uses the engine rather than a copy of it.
    • src/lib/format/workload-prices.test.mjs checks all 1,735.
  • Computed savings. A saving is (price − alternative) / price, rounded. The /alternatives/ route computes 15,082 savings on this data, one per candidate and workload, most never shown as a chip.
    • 183 are exact halves, such as 57.5%.
    • From four-decimal prices and Math.round, 55 came out wrong.
    • Exact prices alone would still leave 15 wrong. (2 − 0.85) / 2 × 100 is 57.49999999999999 in floats, which Math.round makes 57.
    • Every saving computed from prices now rounds half-up on the decimal (roundHalfUp, the L5 rule), and 0 of the 15,082 are wrong. So are 0 of the 1,099 verdict savings in the published model records (/data/models/).
    • Against the payload's four-decimal rates rather than the published ones, 181 are exact halves, the re-review's count. The 8 rates with five or six decimals make the difference.
  • The four-decimal form stays only where it is stored: data/snapshots/, whose published days are hashed. storedPricePerMillion writes it, and today's snapshot is byte-identical under the new code.
    • Pages that show a past day (a stamp, the floor series, the wire) round that stored value once.
    • 31 computed prices would display differently from a four-decimal value (every workload, delivery mode and context rung). None of them is a balanced price, the only workload those pages show.
  • Row order. With exact prices, Qwen3.6 Plus and Inkling improve on GPT-5 and GPT-5.1 by exactly the same margin. A float difference no longer decides which is listed first; the name does, as the comparator says.

84. Correction — three licences contradicted the models' own public repositories

Logged

Live (release 20260930-033430-834a81b1a6f0) published three researched licences that the models' Hugging Face cards contradict. The two batch or free rows that copy them since Task 22 would have repeated the error. Each card was fetched on 2026-09-30:

modellivethe model card
GLM 5.3 (z-ai/glm-5.3, and its batch row)proprietary, not open weightsopen weights at zai-org/GLM-5.3, under Z.ai's own licence ("license: other", "license_name: glm-5.3")
Gemma 4 31B (google/gemma-4-31b-it, and its free row)gemma"license: apache-2.0" at google/gemma-4-31B-it
DeepSeek V4 Flash Vision (Exp) (deepseek/deepseek-v4-flash-vision-exp)proprietary, not open weights"license: mit" at deepseek-ai/DeepSeek-V4-Flash-Vision-Exp, "This repository is licensed under the MIT License"
  • How it happened. The researched records said "proprietary" (GLM 5.3's note: "No Hugging Face repo listed at fetch time, so open-weight status is unconfirmed") and "gemma", and nothing checked them against the feed. The feed names each model's repository (hugging_face_id), and the pipeline discarded that field. The A4 review found GLM 5.3 and Gemma 4 31B. DeepSeek V4 Flash Vision (Exp) turned up in the check that followed, when every researched licence with a repository in the feed was compared with its card.
  • Now. The three records carry the card's licence, each with the card's URL, fetch time and sha256. The pipeline keeps huggingFaceId on every record. A researched "proprietary" or "not open weights" that the feed contradicts by naming a public repository is no longer published: the licence stays unknown, and the enrichment run reports it (scripts/enrich-dataset.mjs). Unknown is not a verdict; a wrong "proprietary" is.
  • What moved. Open-weight models: live 19. This release gains Task 22's four Mistral models and these two, for 25. /open-weights/ lists GLM 5.3 and DeepSeek V4 Flash Vision (Exp). /self-host/ profiles both: GLM 5.3 at 753B and V4 Flash Vision (Exp) at 304.6B, both Hugging Face listings, with no active count because neither card states one. GLM 5.3 becomes the self-host ceiling at a 24 GiB card, about 455 GiB at Q4, in place of Kimi K3.
  • Not corrected here, under review. The same comparison found three more licence labels that differ from their cards, on models already open weights:
    • Qwen3.8 2.4T-A95B: apache-2.0 against "qwen3.8-max";
    • Nemotron 3 Ultra: nvidia-open-model against "openmdw-1.1";
    • Kimi K3: modified-mit against "kimi-k3".

    (Llama 4's llama-4-community against "llama4" is the same licence under two names.) The open-weights flag of all three is right. Which label is right needs the licence texts, not the card's tag.

  • Amended 2026-09-30 (A4 re-review, R109). The texts settle two of the three, and both were wrong on live. Both are fetched 2026-09-30 07:20Z and cited with their sha256 in data/models-handresearched.json.
    • Qwen3.8 2.4T-A95B is not Apache 2.0. The repository's LICENSE (sha256 a136bc41…) is the "Qwen3.8-Max License". Its grant is permissive, but it adds two conditions:
      • a product or service with more than 100,000,000 monthly active users, or more than US$20,000,000 monthly revenue, must display the model name prominently;
      • a model-as-a-service or AI work assistant business whose revenue exceeds US$50,000,000 over any twelve months needs a separate licence from Qwen before any commercial use.

      The row now reads qwen3.8-max. The card's tag, which §84 read as a label question, is that licence's name. Open weights stays true.

    • Nemotron 3 Ultra is OpenMDW-1.1. The card (sha256 8ac2780e…) says "Use of this model is governed by the OpenMDW License Agreement, version 1.1". The text it links is cited alongside (sha256 2ab44b68…). The row, and its free row, now read openmdw-1.1 in place of nvidia-open-model. No other row carried that label.
    • Kimi K3 (modified-mit against "kimi-k3") is still under review.
    • Counts. Open-weight models stay at 25. On /open-weights/ the licence column changes on those two rows, and apache-2.0 covers 9 standard open-weight models, not 10. The payload's licence facet loses nvidia-open-model and gains openmdw-1.1 and qwen3.8-max. Against live, it has also lost gemma, since Gemma 4 31B's row is apache-2.0 above.
  • Amended 2026-09-30 (A7 review, R112). The text settles the third, and live has it wrong: Kimi K3 is under the "Kimi K3 License", not Modified MIT. It is logged as its own dated correction, §87, because live published the wrong label.
    • The menu branch's /licences/ (Task 5) found it first (fetched 13:24Z, sha256 20c797ce…) and, while the catalogue still said modified-mit, listed Kimi K3 apart under a mismatches entry of data/licences.json. A7 relabelled the rows kimi-k3, so on the merge with master (2026-10-07) that entry became a kimi-k3 licence of its own, and the mismatch list is empty. licenceGroups fails the build on a catalogue label with no reviewed entry, and on a mismatch whose label has moved; this is the case its test names.
    • Every other licensed standard model's Hugging Face metadata agrees with its label (21 models, checked the same day). 19 were found through the feed's huggingFaceId. Command A+ and Mistral Large 3 2512 have none; they were checked at the repositories their records' licence sources name (CohereLabs/command-a-plus-05-2026-fp8, mistralai/Mistral-Large-3-675B-Instruct-2512). Llama 4's two repositories and Command A's are gated; their public metadata names llama4 and cc-by-nc-4.0.

83. Correction — Hy3 and DeepSeek V4 Pro 0813 were priced at their discount windows

Logged

The live site (release 20260930-033430-834a81b1a6f0) priced two time-windowed models at their cheaper window. It gave no rule for choosing a window, and none was disclosed:

modelwhat the site showedthe standard window (R106)windows
Tencent Hy3 (tencent/hy3)$0.0825 in / $0.33 out, cached $0.0206; $0.1444/M on the balanced blend; on the value frontier$0.132 / $0.528, cached $0.033; $0.231/M00:00–16:00 UTC at the standard rate; 16:00–24:00 UTC at $0.0825 / $0.33
DeepSeek V4 Pro 0813 (deepseek/deepseek-v4-pro-0813)$0.66 / $1.98, cached $0.022; $0.99/M$1.32 / $3.96, cached $0.044; $1.98/Mpeak 01:00–04:00 and 06:00–10:00 UTC Monday to Friday; off-peak all other hours at half
  • How it happened. OpenRouter's top-level price for a windowed model is the window in force when the feed is read (§82). Hy3's price was held at the live record of 2026-09-23, read in the evening window. DeepSeek V4 Pro 0813 was accepted at the 00:15Z refresh, in DeepSeek's off-peak hours. Neither page said that the price shown was a discount window.
  • The rule now (R106). The headline is the vendor's standard (non-discount) window, chosen from the window set and never from the clock (scripts/price-windows.mjs). It is the set's dearest window; every cheaper window is a discount off it.
    • DeepSeek's pricing page says "Off-peak rates are half of the peak rates. Peak hours are 01:00 - 04:00 and 06:00 - 10:00 UTC, Monday through Friday, excluding Chinese public holidays. All other hours are off-peak" (fetched 2026-09-30).
    • Hy3 charges its dearer rate for sixteen hours and the other for eight.
    • A window set in which no window is at least as dear as every other fails the build rather than being guessed.
    • The discount is disclosed as data. The board flags it ("time-of-day rates"), and the price card, the model page and each endpoint row state its hours and rate. The upstream watch compares window sets, so a read in the other window is no change.
  • What moved on the board (measured on the build of this change):
    • Hy3 left the value frontier: 12 → 11 undominated models. Its verdict is now a trade-off: Gemma 4 31B scores 1441.7 against 1440.6 at $0.153/M against $0.231/M, with 16K max output against 128K.
    • Models with a better-and-cheaper alternative: 123 → 124 of 135 rated and priced (91% → 92%). Trade-off verdicts: 18 → 19. The strictly dominated count stays at 105.
    • The median overpay drops from 6.7× to 6.5×.
    • GLM 4.7 Flash and Step 3.5 Flash had Hy3 as their better-and-cheaper alternative. It is now Gemma 4 26B A4B (1434.5 at $0.121/M).
    • Hy3's cheapest seller becomes DeepInfra at $0.13 (fp4) instead of Tencent's discount window. Its seller spread drops from 2.42× to 1.54×.
    • DeepSeek V4 Pro 0813 is unrated, so it has no rank and no verdict. Its prices double on every surface.
  • Not a claim change: DeepSeek V4.1 Flash. Its headline was already DeepSeek's peak rate, $0.30 / $1.20. The models feed carries no windows for it. DeepSeek's own endpoint row in the endpoints feed does, and so does DeepSeek's page: off-peak $0.15 / $0.60. The discount is now disclosed from that row. The row itself had read $0.15 / $0.60 at 00:15Z and is now priced at its standard window, like every windowed endpoint row.
  • Dated stamps are not rewritten. /now/2026-09-30/ and /data/snapshots/2026-09-30.json were published by the release above with Hy3 on the frontier. They stay as published (data/published-snapshots.json). The corrected board is the current one, and the next dated stamp carries it. The move between those two stamps is this correction, so the wire does not publish it as a vendor price change (data/price-restatements.json).
  • Known gap. DeepSeek's page makes Chinese public holidays off-peak all day. The feed's windows do not encode holidays, so no page states it.
  • Amended 2026-09-30 (A4 review, R108). This entry left out the following:
    • The wire. Live's wire published "DeepSeek V4 Pro 0813 is 15.1% cheaper on the balanced blend" on 2026-09-30 ($1.17/M on 2026-09-29, $0.99/M on 2026-09-30). The $0.99 was DeepSeek's off-peak window. At its standard window the model is $1.98/M, so the same pair is a 70% rise ($1.1661 → $1.98). The event stays, as §82's do: it is published history (R103), and this entry is its correction.
    • Other figures that change on the corrected board.
      • /frontier/: "64% of frontier … (5 move)" becomes "69% … (4 move)". Hy3 was one of the five frontier models that move off the frontier under some workload.
      • /methodology/: the models whose score sits within the benchmark's margin of error of a better-and-cheaper alternative go from 67 to 68 of 358. Models with published price windows go from 2 to 3 (DeepSeek V4.1 Flash).
    • Pages.
      • Three pages are removed. /alternatives/google__gemma-4-26b-a4b-it/ existed only because Hy3, at its discount window, undercut Gemma 4 26B A4B under the agentic workload. Two frontier-pair compare pages, /compare/google__gemma-4-26b-a4b-it--vs--tencent__hy3/ and /compare/google__gemma-4-31b-it--vs--tencent__hy3/, go because Hy3 left the frontier.
      • One is added, /alternatives/tencent__hy3/, now that Hy3 has a better-and-cheaper alternative.
    • Sellers. DeepSeek V4.1 Flash's DeepSeek seller row read $0.15 in / $0.60 out at the 00:15Z fetch, DeepSeek's off-peak window. At its standard window it is $0.30 / $1.20.
    • Tencent's own price. Tencent Cloud's price page (https://www.tencentcloud.com/document/product/1300/78937, fetched 2026-09-30 06:05Z) lists Hy3 flat: $0.132 in, $0.528 out, $0.033 cached, with no peak or off-peak condition. It gives time-of-day rates only for its DeepSeek rows. Hy3's $0.0825 window exists only on OpenRouter's Tencent endpoint (tencent/fp8), so it is that endpoint's off-peak label, not Tencent's price. The headline R106 now publishes, $0.132 / $0.528, is the price Tencent lists.
  • Amended again 2026-09-30 (A4 re-review, R109). Two statements in the amendment above were false. Each explained a figure it had not traced.
    • The frontier movers. It said Hy3 was one of the five frontier models that move off the frontier under some workload. It was not. On live, priced at its discount window, Hy3 was on the frontier under all five workloads: one of the 9 always on, out of 14 ever on. At its standard window it is on under none, so the frontier's union drops to 13. The mover that went away is Gemma 4 26B A4B. It moved under live's prices, and at Hy3's standard price it is on under every workload. The always-on count stays 9, now of 13: 64% becomes 69%, and 5 movers become 4 (integrity.frontierStability, live catalogue against this build).
    • The methodology figure. It said 67 → 68 of 358 counts models whose score sits within the benchmark's margin of error of a better-and-cheaper alternative. It does not. It counts models whose headline price is not the whole price: context tiers, separately billed reasoning, or time-of-day rates (caveatedModels, src/lib/public-facts.mjs). It moved because DeepSeek V4.1 Flash gained time-of-day rates. That is the only model added to the set, and none left it.
  • Amended 2026-09-30 (Task 23 merge, R110). "How it happened" cited a key of the project's private agent notes beside §82. /corrections/ publishes this entry, and the key names nothing a reader can open, so it is removed; §82 carries the finding.

82. Correction — Hy3's price "changes" on the wire were the time of day we read the feed

Logged

The wire (/wire/, /wire.rss) published five price changes and four frontier moves for Tencent Hy3 (tencent/hy3). None was a market move. Hy3 has one rate card with two UTC windows (§24): $0.132 in / $0.528 out per million tokens from 00:00 to 16:00 UTC, and $0.0825 / $0.33 from 16:00 to 00:00. The daily snapshot records OpenRouter's model-level price, and that is the window in force when the feed is read. A snapshot taken before 16:00 UTC recorded the day rate, $0.231/M on the balanced blend. A later one recorded the evening rate, $0.1444/M. The wire read each switch as news.

datewhat the wire publishedwhat it was
2026-08-26"Hy3 is 37.5% cheaper on the balanced blend" ($0.231 → $0.1444/M)read in the evening window, after 08-25 was read in the day window
2026-09-02"Hy3 is 60% dearer" ($0.1444 → $0.231/M) and "Hy3 left the frontier"read in the day window
2026-09-03"Hy3 is 37.5% cheaper" and "Hy3 joined the frontier"read in the evening window
2026-09-15"Hy3 is 60% dearer" and "Hy3 left the frontier"read in the day window
2026-09-23"Hy3 is 37.5% cheaper" and "Hy3 joined the frontier"read in the evening window
  • Measured. Each step is exactly $0.231 ↔ $0.1444/M, the balanced blends of the two windows. The tracked data/models.json records the same two windows, rate for rate, on all 12 fetch dates from 2026-08-28 (the first run that read UTC windows) to 2026-09-30. On each date the top-level price is one of the two windows. The records before 2026-08-28 carry only the top-level price, but §24, written before that run, quotes the same two rates. Hy3's LMArena score was 1441.2 in every snapshot to 2026-09-08, 1439.9 on 2026-09-11 and 1440.6 from 2026-09-15. It moved at one flip, 09-11 → 09-15, by +0.7 with the board's rescoring, which is a rise and cannot move a model off the frontier; the recorded price did.
  • The frontier moves followed from the recorded price. At $0.231/M Hy3 is dominated; at $0.1444/M it is not, so the frontier changed each time the recorded window did. The frontier flip on 2026-08-26 was not published, because that pair was over the wire's board-turnover ceiling.
  • The events stay. They are published history, and they are not deleted (R103). This entry is published on the wire beside them.
  • Now. In the next data release, Hy3's price and its Tencent endpoint row are held at the live record of 2026-09-23, the evening window ($0.0825 / $0.33), in data/upstream-holds.json. scripts/build-wire.mjs (commit a810b16) publishes no price movement for a held price after the live record's date, so no new Hy3 event can come from a window flip while the hold stands. Two things are not fixed by that:
    • The hold covers Hy3 only. DeepSeek V4 Pro 0813, the catalogue's other time-windowed model, is not held. Its accepted headline is DeepSeek's off-peak rate, $0.66 / $1.98, and a refresh during DeepSeek's peak hours (01:00–04:00 and 06:00–10:00 UTC, weekdays) reads $1.32 / $3.96.
    • The headline of any time-windowed model is still chosen by the refresh clock. A deterministic, disclosed choice of window is the follow-up before the next refresh (Task 22).

    The trap is recorded in the project's private agent notes.

  • Amended 2026-09-30 (Task 23 review, R110). The line above named the key of that record, which /corrections/ publishes. The key is replaced by a description; nothing else changed.
  • Amended 2026-09-30 (A3 review, R104). The Measured paragraph first said Hy3's score "did not change across the flips (1441.2 until 2026-09-15, then 1440.6)". The snapshots say 1439.9 on 2026-09-11, and the live wire's 2026-09-15 event says the same ("against 1439.9 at $0.1444/M on 2026-09-11"), so the score did change at that flip. The paragraph is corrected above; the conclusion does not change.

80. Correction — /effort/ counted catalogue rows as models

Logged

Live on aeff187 (release 20260929-214451), /effort/ said "99 of 437 models cannot disable thinking" in its meta description, Open Graph and JSON-LD, and "99 of 437 publishable models (23%)" in its lede and scorecard. 437 is the catalogue's rows: a model's batch and free delivery variants are rows of their own. It is the R58 error §78 fixed on the methodology page, left on a page that review did not reach.

  • Recomputed from build/data/catalogue.json (437 rows; a model is fk, else the slug, as src/lib/public-facts.mjs counts the hero's 352) joined to data/models.json on the slug: every row's fk equals its record's familyKey, and no family mixes mandatory and optional rows.

    | | rows (was published as models) | models | |---|---|---| | publishable | 437 | 352 | | cannot disable thinking | 99 (23%) | 77 (22%) | | reasoning supported | 307 | 232 | | supported and optional | 208 | 155 |

  • Now. src/lib/effort/census.mjs counts family keys; the page and src/lib/seo/identity.ts say "77 of 352 models". The table still lists all 99 entries, and the lede and the evidence claim say so ("the 77 are 99 of the catalogue's 437 entries"). src/lib/knee-effort.test.mjs pins the family count on a fixture with delivery variants.
  • Amended 2026-09-30 (A2 review, R102): this entry misquoted the live page. Its first paragraph said the page read "99 of 437 publishable models (23%)" in its lede and scorecard. Read again in the aeff187 build (its homepage is byte-identical to the live one) and on the live /effort/: only the lede carried "(23%)". The scorecard read "99/437 cannot disable thinking", with 307 "reasoning supported" and 208 "supported and optional" beside it, and no percentage. The evidence claim under the table already said entries: "99 of 437 publishable catalogue entries (23%)". The paragraph above stays as first written, so the misquote is on the record.

79. Correction — "38–42% of a real bill" for reasoning tokens had no source

Logged

README.md's table of what breaks the headline price said that separately-billed reasoning tokens are "38–42% of a real Gemini Flash bill, invisible in the advertised rate". No receipt supports the figure. §78 is left for the other half of the same fix round.

  • Where it came from. It entered with the data-pipeline commit a3123fa (2026-08-24), at the same time in four places: README, this log's price-reliability summary (the "reasoning-tokens 28 38-42% of a real bill" line, near the top, which stays as the record of what was claimed), a comment in src/lib/cost.test.mjs, and a comment in src/lib/signals.mjs, whose price-reliability signal also said "this can be 40% of the bill". Neither signals.mjs text is in the tree now. None computes the figure. No script, data file or artifact in the repo has an invoice, a measured reasoning-token volume, or any other denominator for a share of a bill.
  • The only related computation is synthetic. The test fixture bills 1,000 of 3,000, which is 1/3, and asserts only that the share is over 0.3. Its comment said "on real Gemini Flash models, 38-42% of the actual bill".
  • It was found once already. On 2026-09-24 the §71 review traced the same claim, took it out of the methodology page, and retained the trace: research/artifacts/2026-09-23-recovery/methodology-time-window-proposal/reasoning-claim-trace.json (main checkout, untracked; sha256 830f03a6…). It found "no observed real-bill denominator". README and the generated catalogues kept the figure. The final claims review of 2026-09-29 (S3) found it again in the catalogues. The other half of this fix round handles those.
  • Searched again on 2026-09-29. docs, data, scripts, src, research and the git history (git log -S) for "38–42" and "38-42". The figure appears only as an assertion. The only code that reads reasoning tokens (src/lib/audit/imports.mjs) parses a visitor's own usage export and publishes nothing.
  • Now. README says what is traceable: reasoning tokens are billed per token on top of the visible output, the advertised price does not say how many a request produces, and no fixed share of a bill is claimed. The test comment gives the fixture's own 1/3 and points here. The count beside it (28 models) is the August catalogue's, as README's section now says.

78. Correction — counts, a denominator and a rounding the final claims review caught (2026-09-29)

Logged

The final whole-branch claims review of the leaderboard recomputed every figure a visitor can read (final-review-claims.md). Five of the wrong ones are live on undominated.ai today; the branch did not introduce them:

  • "All 8 of their base models do." The methodology's service-tier paragraph. The 10 tiers have 10 base models, and 8 of them have an LMArena entry; GPT-6 Sol and GPT-6 Luna have none on any board. It now says "8 of their 10", both counted from the payload, in 19 locales.
  • "437 models". The hero, the search box, the SEO descriptions, the header and footer census and the press figures counted catalogue rows. A batch or free service tier is a row of its own: 437 rows are 352 models, and the "69 with tiered pricing" are 49 models. With variants on, the status line's "301 unrated" was 216 models. Every "N models" now counts models (R58).
  • "90% of the market". The homepage's search description in 17 of the 18 generated catalogues, and nothing hid it in 16 of them. The figure is 123 of the 136 rated, priced models, which is what the English says.
  • "1.1M". A 1,050,000-token window printed as 1.1M on 18 models, 4.8% more than it holds, and 32,768 as 33K. A capacity now rounds down. The data pipeline's loss text ("1.1M → 1M context", 13 model and alternatives pages) still rounds up until npm run data regenerates it.
  • "The other 90% are strictly worse deals." Methodology section 4, against section 3's 14 of 123 verdicts that give something up. Section 4 now defines the frontier exactly and says each of the others has a model that scores at least as high for no more, not always a like-for-like one.

Two more were new in the branch and never live: "Newest on the frontier" under a task lens named a join to the general board's frontier, and the document board's cards stated 95% intervals for a board that publishes none. Both were fixed before release, in commits e80b123 and bdd37d0.

77. Correction — nine published performance figures flattered the site

Logged

The final whole-branch review measured every figure in the AGENTS.md budget table and its notes, README.md's "Measured, not claimed" and "Why it is fast", and DESIGN.md §10.11, against full builds of 21f9e1e, 5934e20 and a1b84a8. Nine figures were wrong in the flattering direction: they showed the site smaller or faster than measured.

WhereSaidMeasured
AGENTS, client JS (initial graph)37,033 B37,195 B brotli, 162 B more. Since 5934e20 (37,039 B) the router manifest for /badges/ and /reasoning-price/ added 156 B
AGENTS, catalogue payload22.4 KB26.6 KB brotli: the client chunk that carries it, 26,560 B (catalogue.json 26,441 B). Already wrong on live, whose data these are
AGENTS, 2026-08-27 notethe initial graph "is 34.1 KB"37,195 B; the note now speaks of 2026-08-27 in the past tense
AGENTS, 2026-08-27 notethe homepage's inline script "is 3.2 KB"4,127 B
AGENTS long tasks and its 2026-09-28 note, README, DESIGN §10.11header-click sorts with a long task, "1 of 42" at load ≤ 2.11–4 of 42 over three low-load runs: the Task 15 re-review measured 4 of 42 at load 1.6–2.4 (a1b84a8 0 of 36). The best run had been quoted as the figure
DESIGN §10.11homepage HTML "4.7 KB gzip smaller"3.9 KB (3,909 B en; 3,909–5,715 B over 19 locales)
DESIGN §10.11JS before hydration "+15.6 KB brotli"+17.4 KB (165,322 against 147,905 B)
DESIGN §10.11first-paint CSS "+6.4 KB"; initial graph "37.1 KB"+6.5 KiB (6,657 B); 37.2 KB
README"the entire 410-model catalogue is 19.1 KB brotli"; catalogue.json "19.5 KB brotli"26.4 KB brotli, 264.6 KB raw, 437 rows

Stale the other way, or without a source:

  • FCP, "217 ms home · 68 ms methodology" (AGENTS). No receipt, and older than a1b84a8. At load 4–10, FCP from requestStart is 315–431 ms at home and 239–307 ms on /methodology/, on every build including a1b84a8; that page's first style-and-layout task alone takes 182–251 ms before it paints. The review could not measure at a load comparable to the figure's, so it is not proven wrong, only unsupported. The row now says "not re-measured; no receipt" and gives no new figure.
  • Client JS, "52 KB brotli" (README). No source. The initial graph is 37,195 B.
  • Homepage HTML, "42.9 KB" (AGENTS). That was the 2026-08-27 page. It is 40,964 B (en), and 40,132–42,953 B over 19 locales.
  • Tap targets, "0 of 183". The claim holds, the count was stale: 0 of 1,414 targets at 360 px.
  • README, "Every figure below was measured from the live catalogue". The figures under it date from the August catalogue (410 models); no line of that section had changed since 2026-08-26. It now says so, and the catalogue counts elsewhere in README.md and REPO-MAP.md give 2026-09-29's: 437 rows of 352 distinct models, 441 model pages.

Not re-verified: the low-load figures (the select's input task 33 ms, header INP 80 ms, keystrokes 7.1 ms, hydration 91 ms). The review's machine never fell under a 1-min load of 4. Their relation to a1b84a8 reproduces at load 5–11 (INP +24 ms, the input task ×2.5, hydration +20–25 ms), so they stay, with their load.

How it was found: by measuring each published figure against a build, not by reading it. The byte figures reproduce with node scripts/measure-home-budgets.mjs after npm run build. The review is a working document of the release process, kept outside the repository.

Amended 2026-09-30 (Task 23 review, R110). The paragraph above gave the review's path in a private working tree, which /corrections/ publishes. The path is replaced by a description; nothing else in this entry changed.

76. Correction — the "expiring" frontier hallmark had not expired since 2026-08-30

Logged

A hallmark (/hallmark/<key>.svg) is meant to state this week's verdict and strike its claim when a model leaves the frontier. Production served the stamps computed on 2026-08-30, a month after the verdicts had moved on.

  • How. Two things write /hallmark/<key>.svg.
    • The route src/routes/hallmark/[slug].svg prerenders one stamp per stamped verdict from the current data.
    • scripts/build-hallmark.mjs writes the same files into static/hallmark/, which is gitignored. Only npm run data:hallmark ran it, and nothing ran that after 2026-08-30. It is not in npm run data or npm run build.

    Where both existed, the static file won: every built stamp that had a static twin was byte-identical to it.

  • Measured on the leaderboard build at 5934e20 (hallmark-shadow.mjs, which renders every stamp with the route's own hallmarkFor and hallmarkSvg and compares):
    • 136 stamped verdicts. 130 built stamps differed from what the route renders, and all 130 were the 2026-08-30 files.
    • 7 stated the opposite frontier claim. "On the undominated frontier" for Gemini 3.7 Flash, Solar Pro 4 and MiMo-V2.5-Pro, which are off it in the 2026-09-23 data. "Not on the undominated frontier" for gpt-oss-20b, Gemma 3 12B, Gemma 3 4B and GLM 5.3, which are on it.
    • 4 more stamps belonged to keys with no verdict left: allenai__olmo-3-32b-think, anthropic__claude-opus-4, ibm-granite__granite-4.1-8b and qwen__qwen3.8-max.
  • Production is the same. A cache-busted GET of undominated.ai on 2026-09-28 returned the 2026-08-30 stamps for gpt-oss-20b ("Not on"), Gemini 3.7 Flash and Solar Pro 4 ("On"), and 200 for the orphan anthropic__claude-opus-4.svg.
    • Each stamp carries its "as of 2026-08-30" date, so no stamp misstates its own date.
    • But a hallmark is sold as expiring, and these did not expire.
    • The new /badges/ page would have shown gpt-oss-20b's "Not on the frontier" stamp as its live example, beside a badge saying "undominated".
  • Why badges were fine. build-badges.mjs has always run in npm run build, and it clears static/badge/ first. Badges: 346 built, 346 current.
  • Fix. npm run build runs build-hallmark.mjs right after build-badges.mjs. build-hallmark.mjs clears the folder first, so orphans go too. After the fix, the build has 136 stamps, all equal to the route's rendering and all "as of 2026-09-23": 13 live, 123 struck, 0 orphans. A test in src/lib/hallmark/contract.test.mjs fails if the build stops running the script before vite build.
  • The hallmark README was wrong too. Its README-badge section said badges use a nested provider/model path, "not __". The flat __ key has been canonical since db19a53 (2026-08-30), as the root README shows; build-badges.mjs writes the nested path only as an alias for early embeds. The project's private agent notes had copied the README's claim. Both are corrected.
  • Amended 2026-09-30 (Task 23 review, R110). The bullet above named the key of a record in the project's private agent notes, which /corrections/ publishes. The key is replaced by a description; nothing else in this entry changed.

75. Correction — "contrast failures: 0 of 1502" had no source, and the audit behind it had stopped running

Logged

AGENTS.md, README.md and docs/TOOLING.md said the site had 0 contrast failures, "0 of 1502". It was never true of anything this repo can re-measure.

  • Where it came from. The row entered with the repo foundation commit (1418eda, 2026-08-24 17:19). The token contrast audit, scripts/audit-contrast.mjs, did not exist until 26fd25e, five and a half hours later. No script, receipt or trace produces 1502. The one other mention, research/2026-08-24-undominated-ui-ux-transformation-review.md:94, cites AGENTS.md itself.
  • What the audit has said. Run at each revision on its own tokens (git archive <rev> scripts/audit-contrast.mjs src/styles): 38 failing pairs at 26fd25e, 32 on master (a99ca4a). It has never reported 0.
  • It then stopped running, unnoticed. a1b84a8 moved the frontier row's re-tuned faint ink from a literal light-dark() in app.css into a token, --frontier-faint in tokens.css. The script matched only the literal and threw ("frontier-row --fg-faint override not found in app.css") at every revision from a1b84a8 to 06b97da. Nothing executed it: it is not in package.json, the justfile or CI, and the one test that named it read its source text.
  • Repaired and measured (Task 14, d907d76). The script now follows var() to the token. On master's CSS its output is byte-identical to master's own script, apart from the pairs added since. On the branch it lists 29 failing pairs, the same set at a1b84a8, 06b97da and the Task 14 tip. That commit's body and the Task 14 report called them "all outside the board rows". That was wrong: the 2 px --signal frontier rule (light 2.28:1) is drawn on the board's frontier rows. It is not their only cue; they also carry the wash and their verdict text.
  • One of the 29 was live text. The top-3 task rank on 23 model pages was 12 px --signal on the white panel, 2.37:1 by axe and 2.38:1 by the audit. It is now --signal-text (9.66:1 light, 12.10:1 dark), in Task 14 fix round 1.

The state now: 28 of 150 gated token pairs fail, and none is live text.

  • 17 are live borders, strokes and markers: chip and input borders, the bright signal on paper, the frontier chart, the board's frontier rule.
  • 9 are in two components nothing imports (RowBadges, QualityBar).
  • 2 are not usages (a claim under test, and a declared token with no text use).

On the leaderboard, axe color-contrast finds 0 with the score lens off and on (audit-board check 8). AGENTS.md, README.md and TOOLING.md now say this instead of "0 of 1502". scripts/audit-contrast.test.mjs runs the script in npm test, so a crash fails the suite. It checks the score lens's pairs and the task rank, not the count. The script still exits 0 with failures; the design-system plan gates the count with its own ratchet. Receipt: research/artifacts/2026-09-24-leaderboard/contrast-audit-history.json (every revision's run, and the 28 pairs by class).

74. Correction — "zero requests during interaction" missed the browser's own refetches

Logged

AGENTS.md lists network requests during interaction as a budget of 0, currently 0. The leaderboard audit (Task 12) counted requests at the static server instead of through Playwright's request events, and the figure is not 0 on either the live application (a1b84a8) or the leaderboard branch.

  • Every URL change re-requests /favicon.ico, in both builds. The page writes its view to the URL with history.replaceState on each sort and each keystroke in the board's search, and every write that changes the URL brings one GET /favicon.ico. Over 20 sequences per build: 120 of 120 sorts and 480 of 480 keystrokes, at a 400 ms and at an 80 ms pace alike. An operation that leaves the URL as it is fetches nothing: in the interaction A/B's base quality↔price phase, 42 sorts of which 28 changed the URL brought 28 requests, and 4 keystrokes into the site header's search, which never writes the URL, brought none (Task 12 re-review). Chromium makes this fetch itself, so Playwright's page.on('request') never reported it. Resource timing does record it (initiator other).
  • On the branch, widening a search often re-requests the provider logos. A search typed and then deleted removes the unscored directory's maker groups and rebuilds them, logo <img> elements included. The rebuilt logos were requested again at these rates:
    • The first type-and-delete in the page, 26–28 logos each time, including the 73 KB aion-labs.svg. Here it came after six sorts, at a 400 ms pause between keystrokes: 19 of 20 sequences. In the Task 12 re-review it came on a freshly loaded page: 5 of 5 at 400 ms, 5 of 5 at 250 ms, and 3 of 5 at 80 ms.
    • A second type-and-delete straight after, in the same page: 1 of 20 here at 80 ms, and 0 of 5 in the re-review.
    • Searches typed after other use of the page: the review's earlier harnesses saw it in 2 of 12 sequences.

    The live application keeps its unrated rows in a pool. It made no such request in 50 sequences: 40 here and 10 in the re-review. What decides whether the logos are requested again is not established.

Amended 2026-09-26, before merge, after the Task 12 review. The first version of this entry said four keystrokes added no favicon request, and that a type-and-delete sequence re-requested 28 logos. Both were wrong. The keystroke figure came from one short probe that typed into the site header's search box, which is the first input[type=search] on the page and never writes the URL. It was not reconciled with the Task 12 interaction receipt, which already counted 84 favicon requests for 84 board keystrokes in each build. The 28 was one sequence's count, not a rate. The figures above are from a re-measurement with 20 interleaved sequences per build, counted at the server (research/artifacts/2026-09-24-leaderboard/request-rates.json).

The first amendment then put the logo refetch down to typing pace ("19 of 20 at 400 ms, 1 of 20 at 80 ms"). That was itself confounded by order: the 80 ms phase always ran second, in the same page, straight after the 400 ms one. The re-review's control, a fresh page at 80 ms, refetched in 3 of 5 sequences. The rates above now say which condition each one was measured under, and make no claim about the cause. Its "none without a URL change" rested on a count in which every operation changed the URL. The two observations cited in the first bullet now support it.

All of these measurements were local, on a server that sends no cache headers, where each request fetched the whole file. In production both paths fall under nginx's location /, which sends max-age=0, must-revalidate. So each should cost a conditional request (a 304) rather than a full download. That follows from the configuration; it was not measured on the live site.

No application code changed. Two ways to meet the budget: give /providers/ and the favicons a cache lifetime, as /logo.png already has (max-age=86400); or keep the directory's maker groups in the DOM through a search. Until one of them ships, the budget line should read "0 application fetches". Receipts: research/artifacts/2026-09-24-leaderboard/request-rates.json and research/artifacts/2026-09-24-leaderboard/perf-ab-a1b84a8.json (requests).

Amended 2026-09-28 (Task 15 fix round 1).

  • The 304s were measured, then removed. The Task 15 reviewer served the build with the origin's real headers for /providers/*.svg (max-age=0, s-maxage=300, must-revalidate, an ETag and a Last-Modified). Every search then sent 16 conditional requests, each answered 304, in 5 of 6 rounds. The logos behind them are the makers folded under hidden="until-found".
  • Through Cloudflare today the count is 0. Its 4-hour browser TTL serves the logos from the disk cache.
  • Task 15's report was wrong. It said a Last-Modified validator, "which nginx sends", made it 0. That holds only for a validator without max-age=0. The origin sends max-age=0.
  • The fix is in the page, not the cache headers. The directory keeps each maker's logo node for the page's life and moves it into the group a search or filter re-creates (keepMark, src/lib/ui/board/UnscoredDirectory.svelte).
  • Measured against the origin's headers, 6 interleaved rounds:
    • requests during a search: 0 (none reached the browser's network stack), against 16 × 304 before;
    • logo <img> nodes that survive a search: 48 of 48, against 1 of 48 before;
    • the directory's pixels, folded, after a search, after a filter, and unfolded at two depths: identical.
  • Hiding groups instead of re-creating them was rejected. A hidden="until-found" group keeps its box, so it would hold a grid cell, and a filter would reorder the makers.
  • Still open: the favicon. It is fetched again on each URL change (first bullet above).

Receipt: research/artifacts/2026-09-24-leaderboard/perf-r82-fix1.json (I3_logos).

53. Correction — the paper mixed specifications, matching states and historical populations

Logged

The publication assessment and independent code/data audits found contradictions in the unpublished manuscript. A threshold-five, sixteen-group description survived beside threshold-eight, ten-group estimates. The active combined McFadden R² was printed as 0.163; the retained fit is 0.126539, displayed as 0.127. Threshold seven also converges, and BIC favors the namespace model among the four displayed candidates. The revised manuscript and statistical register now use one explicit specification, with conventional exploratory tests distinguished from clustered uncertainty.

The displacement helper compared replay scores using intervals borrowed from the catalogue's historical matching state. Its selected-row comparison now uses the repaired source record's score and exact interval endpoints together: the ten-record median half-width is 4.18, and displacement/half-width spans 0.28–5.33. The historical stored population's 4.62 is not relabelled as the repaired population's 4.67. Solver differences remain baseline-relative; they do not establish true matching errors or recovered evaluation coverage.

Standard delivery is an eligibility policy, not complete deduplication. Six of the 78 excluded variants have no retained identifier/family counterpart. The new eligibility ledger records every exclusion. Namespace and display-label counts are separate, and a versioned crosswalk supports a labelled grouping sensitivity rather than asserting that namespaces are independent producers.

Historical validation also used a moving HEAD population for exclusions. With the dated August 28 population, the comparison is 114/118, not 113/116; all four disagreements are printed. The required historical input is now a frozen, hashed projection, so reproduction requires neither Git history nor a live feed. Restricted commercial benchmark fields are removed from the candidate input projection with original and projected hashes retained. Arena's 908 normalized records were retrospectively corroborated against pinned upstream revision 79c11360ced70bee2c635f18a7b9ec17e38a247f. This does not invent an original request receipt or establish redistribution rights for other fields.

The prospective adjudication apparatus missed source-row disagreements and could accept incomplete or conflicting annotations. The amended frame contains 38 disputes and 158 blank cases. Its scorer validates both labels and required source rows. Simple random samples replace zero-probability namespace quotas; exact hypergeometric inversion replaces an invalid interval that could exclude zero at zero observed errors. No completed human annotations or accuracy estimates exist, and software-test fixtures are not counted as such evidence.

The main text is now a conditional operator case study. The historical correction ledger remains separate. Source/PDF receipts and complete-cell table checks address artifact drift; citation generation now actually respects the unpublished flag. The flag remains false, with no deployment in this task. Field-level distribution rights remain separately documented and unresolved where no permission or reviewed basis was established.

Detailed responses: paper/REVIEW-RESPONSE-2026-09-22.md. Inputs, executable counterexamples, model review receipts, revised rendering proof and reproduction logs: research/artifacts/2026-09-22-paper-readiness/.

The private publication-browser rehearsal caught a further source-to-web defect: the renderer handled [@key] but left narrative @key and author-suppressed [-@key] citations unformatted. The web bibliography therefore contained 16 works while the PDF cited 20. All three forms now render, and the regression compares the complete source citation-key set with rendered links and references, rather than checking only the links the renderer happened to create.

52. Correction — catalogue HTML existed without its matching JavaScript release

Logged

During verification of the workflow and interface release, /skills/ and /mcp-servers/ at the origin referenced an older Svelte build. The active release was 20260922-023654, but those catalogue pages retained HTML dated 2026-09-21 22:01:54 UTC. Origin and public HTML hashes matched, so this was not an edge-only stale response. One referenced entry module, /_app/immutable/entry/start.BvWBGga3.js, returned 404 directly from the origin. Its availability in an existing browser or edge cache could conceal the incomplete release.

The previous activation contract proved that catalogue indexes, pages and attributed downloads existed. It did not prove that the pages referenced JavaScript and stylesheets present in that release. Preserving old HTML alone therefore did not preserve its functionality.

Correction. Build the catalogues together with the rest of the application. The activation helper now checks prerendered immutable-asset references before changing the active release, in addition to its existing catalogue checks. New workflow JSON must also resolve to catalogue records and rendered guides; a later build cannot silently omit the published workflow feature. Regression tests exercise valid and missing absolute/relative assets without mutating the simulated live release.

This changes no model score, price or ranking rule. Build, browser, source reconciliation and eventual release evidence are recorded in research/artifacts/2026-09-22-platform-workflows/.

51. Correction — feature coverage implied a tools recommendation it could not establish

Logged

/tools/ used a Skills CLI target list dated 2026-08-25: 77 entries, with 21 authored featured flags. Its six-bit coverage calculation highlighted Cline, Goose, OpenCode and Qwen Code in frontier green. The page qualified this as feature coverage, but the selection and colouring still implied a purchasing recommendation without a comparable quality measurement.

The calculation treated desktop apps as IDEs, and an undocumented capability as absent when comparing tools. Search also changed the comparison population. The OS selector could silently substitute a Linux command for missing Windows instructions. A brand-wide feature union could additionally combine a CLI's licence and model backend with an unrelated desktop surface.

Source correction. The rebuilt catalogue removes feature-dominance verdicts and the arbitrary featured subset. Explicit workflow requirements match current tools using documented interface scopes; missing evidence and known unsupported or unverified combinations cannot qualify. Editorial starting points state the job, rationale, limitation and alternatives; they are not performance winners. All catalogue matches remain visible in alphabetical order, with per-profile source dates, lifecycle notices and comparable fields.

Current research uses the live Skills registry, ACP registry, MCP client documentation and individual official repositories/docs. Generic installation targets and duplicate renamed identities are excluded from product counts. Retired and unresolved records remain discoverable. Exact installation commands are selected only for the chosen OS; locally detected device information never supplies a command or substitutes another OS's instructions.

No benchmark score or exact price is inferred by this correction. A refreshed HTTP response is recorded separately from semantic review and cannot renew a profile's review date by itself. Research, counts, validation and release status: research/artifacts/2026-09-21-tools-rebuild/.

48. Correction — the model page still flattened the ladder that §17 un-flattened

Logged

The client payload has carried every context-tier rung since §17 (tr on the catalogue row, pricing.tiers on the model file). The homepage price cell walks that array. The model page did not.

src/routes/[[locale=locale]]/models/[slug]/+page.svelte printed:

Past {first threshold} tokens the price changes. Input goes to {last rung}/M ({max multiple}×).

For Qwen3.7 Flash that is "Past 32,000 tokens … $0.200/M (6.67×)". The raw OpenRouter overrides, fetched 2026-09-19, are still two rungs: $0.10/M past 32K (3.33×) and $0.20/M past 256K (6.67×). The middle band was invisible on the page that exists to explain the price.

The same pairing lived in src/lib/signals.mjs (tiers.at(-1) against tierThreshold), so the warning string agreed with the lie.

The ranking engine and the payload were right. The sentence was not. Same class of bug as §17, one layer further out. The model page now lists every rung; the warning names every threshold; a test refuses the flattened sentence.

22. Correction — the delayed analytics beacon used the wrong collection origin

Logged

The first production audit after moving Cloudflare Web Analytics out of the critical path scored 96 Best Practices, not 100. The page logged two deterministic errors: the manually loaded beacon posted to https://cloudflareinsights.com/cdn-cgi/rum, and the POST response did not admit the https://undominated.ai origin. The browser reported both the CORS failure and the failed resource. Allowing that host in CSP was therefore necessary but not sufficient; it authorised a request that still failed at Cloudflare's ingestion boundary.

The collection target was tested without changing production by substituting only the beacon configuration in a browser route. With send.to: '/cdn-cgi/rum', the same script and site token posted to https://undominated.ai/cdn-cgi/rum, received 204, and produced zero console errors. The app now carries that explicit same-origin target and the CSP removes the no-longer needed remote connect-src origin. A regression test binds both halves of that contract.

The same Lighthouse run reported 86 Performance, but explicitly warned that the test device's CPU was slower than Lighthouse's calibration; host load average was 13.68. That number is not used as product evidence. The CORS failure was independent of throttling and was reproduced in a separate browser pass.

Lesson: successful script loading does not prove successful telemetry. Observe the beacon POST and the console on the public origin; manual and automatically injected Cloudflare beacons choose different default collection origins.

19. Correction — "let Cloudflare compress it" made every asset larger

Logged

AGENTS.md and docs/DEPLOYMENT.md both recorded that the origin nginx has no brotli module, "so the edge must do it," with Cloudflare Brotli set On. Two things were wrong with that.

First, the edge never got the chance. nginx has gzip_static on, so it served the precompressed .gz and the response arrived at Cloudflare already Content-Encoding: gzip. Cloudflare does not re-encode an already-encoded response, so it passed gzip straight through. Measured across the homepage's 16 critical-path assets: 61,116 bytes served where 53,743 was available — 12% wasted on every cold load, with the .br files the build produces sitting unused because nginx cannot serve them without brotli_static.

Second — and this is the part that matters — the obvious fix made it worse. Disabling gzip_static and gzip for the immutable location so the origin sent plain bytes did cause Cloudflare to compress, with zstd. Measured on the same 31,624-byte stylesheet:

PathBytes
build-time gzip -9 (what we had)7,448
Cloudflare dynamic brotli7,629
Cloudflare dynamic zstd (what browsers then got)8,081
build-time brotli -q 11 (unreachable — no module)6,367

Cloudflare compresses on the fly at speed-optimised quality; the build compresses offline at maximum. A static gzip -9 beats an edge brotli. Reverted.

The genuine win is brotli_static serving the 6,367-byte artefact the build already produces — a further 15% — which needs the nginx brotli module. It is not packaged for Debian 13 and is not worth compiling on a box shared with seven other production sites. Do it on the dedicated VPS.

Lesson: "the CDN will handle it" is an assumption, not a mechanism. Both halves of this were found by measuring bytes on the wire, and the second half only because the first fix was measured after shipping rather than assumed to have worked.

18. Correction — r = 0.409 was a small-sample artefact; it is 0.806

Logged

Section 12 of this document, the README, the methodology page, the model pages and several of my own reports all claimed the two independent quality sources correlate at only r = 0.409, and built the "never blend them" argument on that number.

It was wrong. That figure came from 15 models read off a leaderboard page — a sample that was both unlicensed and, more damagingly, entirely frontier-class. Restricting the range of a variable that severely attenuates its correlation; the models spanned roughly 1476–1508 Elo, a 32-point window inside a 450-point scale.

Re-sourced from Arena's official CC-BY-4.0 dataset, the overlap is 62 models spanning 1055–1504 Elo, and the correlation is r = 0.806. The benchmarks largely agree.

What survives, and what does not

  • Does not survive: "two reputable benchmarks substantially disagree." They do not.
  • Survives, and is now sharper: showing both axes rather than blending. The reason changes from they disagree wholesale to they agree in aggregate and diverge on 4 of 62 — and a composite erases exactly the models worth flagging.

The bug behind the bug

The first re-computation on the licensed data returned r = −0.012, which would have been a far more dramatic finding. It was also false. Arena publishes several boards on different scales: text, search and document are Bradley-Terry Elo (~830–1520), while agent uses ips, which ranges −0.2 to 0.1. The merge fell back to any available board when no text row matched, storing ips values of 0.1 in the Elo field. Near-zero "Elo" paired against strong models drove the correlation to nothing.

Near-zero correlation is the signature of randomly paired records, which is what prompted the check rather than the publication. The merge now accepts only Bradley-Terry rows from the primary board, inside a plausible range, with no cross-scale fallback.

Two lessons, both cheap to state and expensive to learn: a statistic from a convenience sample is worth approximately nothing even when it is arithmetically correct; and a number that is dramatic enough to build a product argument on deserves one more check before it becomes a headline, not after.

17. Correction — the tier ladder is not one rung

Logged

An earlier version of this document, the README and the shipped price cell all claimed "Qwen3.7 Flash is 6.67× dearer past 32K tokens". That was wrong, and it was wrong in the product as well as the prose.

Qwen3.7 Flash has two tiers: $0.10/M from 32K (3.33×) and $0.20/M from 256K (6.67×). The lean client payload flattened the ladder to a single {threshold, rate} pair, taking the first threshold with the last tier's rate. Five models carry more than one tier, and every one of them was overcharged across its entire middle band — Qwen3.7 Flash quoted $0.35/M at a 50K prompt where the truth is $0.175/M, a clean 2×.

Found by an adversarial agent reading scripts/build-payload.mjs:64-66 during the workflow's verification phase, not by a test. The payload now carries the full ladder and priceOf walks it; a regression test pins the three bands.

The lesson worth keeping: the ranking engine had this right all along — effectivePricePerMillion in ranking.mjs iterates tiers correctly. The bug lived only in the projection built for the client, which is exactly the code path no unit test covered because it looked like plumbing. A wrong price is the one bug this product cannot ship, and it reached the browser through the least interesting file in the repo.

Investigated corrections

Recorded changes to a published claim or calculation. A source update and a correction are different events.

2026-09-24 · Equal-price alternatives are no longer labelled cheaper

Before: Some alternatives headlines and column labels promised both a higher score and a lower price even when the comparison included equal-price candidates. Search descriptions also attributed a union of workloads to one workload.

Corrected: The pages distinguish higher measured score or lower cost, with neither worse under at least one listed workload. Matching recorded capability requirements is kept separate from a universal drop-in guarantee.

Visible copy and search metadata. Candidate sets, counts, prices, scores and workload calculations are unchanged. Inspect evidence →

2026-09-24 · Family coverage reconciled with the accepted catalogue

Before: The Claude and Command family pages said Opus 5.5 and Command A+ were absent even though the accepted catalogue included them. Some other accepted family members were also missing from the curated lists.

Corrected: Known family memberships now include the accepted additions and their separately labelled service variants. Being unrated remains distinct from being absent. Family ranges are recomputed over the corrected membership.

Membership and commentary corrections using the existing September 23 catalogue; no new rate, context limit or benchmark score. Inspect evidence →

2026-09-24 · Stale measurements removed from model commentary

Before: Some Our take paragraphs repeated older prices, scores and rankings, or inferred speed from latency figures already flagged as suspect. Muse Spark 1.2 commentary quoted a score despite the accepted record being unrated.

Corrected: Model commentary explains fit and constraints. The current proof and price tables supply measurements, with their sources and conditions. Withheld latency cannot support a speed comparison.

All authored model notes were reviewed against accepted records. This review does not renew vendor-source observation dates or add benchmark results. Inspect evidence →

2026-09-24 · Billing annotations now follow the structured rate evidence

Before: Historical prose created time-of-day flags for three DeepSeek marketplace rows without native window evidence. A duplicate long-context summary for GPT-5.6 Sol disagreed with its actual tier table, and several Grok notes described a strict threshold as inclusive. Methodology copy also claimed a real-bill reasoning share without a traced invoice.

Corrected: Structured API ladders and published windows govern billing annotations. Unsupported time flags and duplicate long-context summaries are removed, threshold wording matches the calculation, and reasoning cost is described as depending on billed volume and rate.

Accepted rates, complete tier ladders, native Tencent windows, benchmark scores and source observation dates are preserved. No fixed time discount or reasoning share is assumed. Inspect evidence →

2026-09-23 · Frontiers now stay within the measured benchmark task

Before: Some image and video recommendations compared a model on one task with a model scored on another task.

Corrected: Dominance is recomputed within each selected benchmark task. Secondary task scores remain visible without entering the primary frontier.

Accepted image and video snapshots; source scores and quoted prices are unchanged. Inspect evidence →

2026-09-23 · MiniMax M3 licence corrected

Before: The catalogue and self-host profile labelled the model MIT.

Corrected: The official model card identifies the MiniMax Community licence. The source URL and its content hash are retained; no broad commercial-use assurance is inferred.

MiniMax M3 licence metadata. Inspect evidence →

2026-09-23 · Gemini 3.5 Pro announcement separated from catalogue availability

Before: A stub headline said no source named the model, based on an older check.

Corrected: Google’s current Gemini page names 3.5 Pro as coming soon. This catalogue still has no accepted priced entry for that exact model.

Announcement language does not establish a price, availability date or independent score. Inspect evidence →

2026-09-23 · Empty-result counts respect the selected delivery lane

Before: The search-empty message attributed all catalogue entries, including already-excluded service variants, to the active filter.

Corrected: The baseline clears only the filters counted in the message and preserves the selected delivery lane and workload.

Filter explanations; rankings and prices are unchanged. Inspect evidence →

Catalogue snapshot: 2026-10-06. Publication ecef83f767cd. The table below compares available daily observations; it is not a complete error history.

ModelFieldWasIsSeen
z-ai/glm-5v-turbolmarena1436.51437.22026-10-06
z-ai/glm-5.3-flashlmarena1471.91469.62026-10-06
z-ai/glm-5.3price.balanced2.1512026-10-06
z-ai/glm-5.3lmarena1475.11471.42026-10-06
z-ai/glm-5.2price.balanced1.10621.122026-10-06
z-ai/glm-5.2lmarena1466.91470.42026-10-06
z-ai/glm-5.1price.balanced2.151.82752026-10-06
z-ai/glm-5.1lmarena1462.41461.52026-10-06
z-ai/glm-5price.balanced0.931.0852026-10-06
z-ai/glm-4.7-flashprice.balanced0.14540.1452026-10-06
z-ai/glm-4.7-flashlmarena1352.21350.92026-10-06
z-ai/glm-4.7price.balanced10.73752026-10-06
z-ai/glm-4.7lmarena1435.91435.52026-10-06
z-ai/glm-4.6vlmarena1376.51376.82026-10-06
z-ai/glm-4.6lmarena1440.51439.92026-10-06
z-ai/glm-4.5vlmarena1332.91332.82026-10-06
z-ai/glm-4.5-airlmarena1383.71383.92026-10-06
z-ai/glm-4.5lmarena1430.21430.32026-10-06
xiaomi/mimo-v2.6-proprice.balanced0.54370.542026-10-06
xiaomi/mimo-v2.6-prolmarena—14912026-10-06
xiaomi/mimo-v2.6-flashlmarena—1456.42026-10-06
xiaomi/mimo-v2.5-prolmarena1464.814652026-10-06
xiaomi/mimo-v2.5lmarena1427.41427.52026-10-06
x-ai/grok-4.7lmarena—1399.92026-10-06
x-ai/grok-4.6lmarena1429.91427.42026-10-06
x-ai/grok-4.5lmarena1450.11448.12026-10-06
x-ai/grok-4.3lmarena1397.71396.72026-10-06
upstage/solar-pro4price.balanced0.15750.5252026-10-06
upstage/solar-pro4lmarena1377.31386.22026-10-06
upstage/solar-mini4price.balanced0.08750.1752026-10-06
thinkingmachines/inkling-smalllmarena1412.31413.72026-10-06
thinkingmachines/inklingprice.balanced1.76251.7252026-10-06
thinkingmachines/inklinglmarena1439.71441.82026-10-06
tencent/hy3price.balanced0.14440.2312026-10-06
tencent/hy3lmarena1440.61440.72026-10-06
stepfun/step-3.5-flashlmarena1403.714032026-10-06
qwen/qwen3.8-27bprice.balanced1.0650.5622026-10-06
qwen/qwen3.8-27blmarena1439.31440.72026-10-06
qwen/qwen3.7-pluslmarena1454.21454.62026-10-06
qwen/qwen3.6-max-previewlmarena1446.31446.82026-10-06
qwen/qwen3.6-35b-a3bprice.balanced0.36250.32026-10-06
qwen/qwen3.6-27bprice.balanced1.041.01252026-10-06
qwen/qwen3.5-flash-02-23lmarena1397.713972026-10-06
qwen/qwen3.5-397b-a17bprice.balanced1.28750.87752026-10-06
qwen/qwen3.5-397b-a17blmarena1438.314382026-10-06
qwen/qwen3.5-35b-a3bprice.balanced0.44690.3552026-10-06
qwen/qwen3.5-35b-a3blmarena1395.51394.42026-10-06
qwen/qwen3.5-27blmarena1408.11408.92026-10-06
qwen/qwen3.5-122b-a10blmarena1417.91417.12026-10-06
qwen/qwen3-vl-30b-a3b-thinkingprice.balanced0.750.46752026-10-06
qwen/qwen3-vl-30b-a3b-instructprice.balanced0.26250.22752026-10-06
qwen/qwen3-vl-235b-a22b-thinkinglmarena1400.81400.72026-10-06
qwen/qwen3-vl-235b-a22b-instructprice.balanced0.63250.372026-10-06
qwen/qwen3-vl-235b-a22b-instructlmarena1420.61420.12026-10-06
qwen/qwen3-next-80b-a3b-thinkinglmarena1367.81368.32026-10-06
qwen/qwen3-next-80b-a3b-instructprice.balanced0.350.26812026-10-06
qwen/qwen3-next-80b-a3b-instructlmarena1417.61417.32026-10-06
qwen/qwen3-maxlmarena1413.11412.42026-10-06
qwen/qwen3-coder-30b-a3b-instructprice.balanced0.12250.122026-10-06
qwen/qwen3-coderprice.balanced0.4750.63752026-10-06
qwen/qwen3-32blmarena13401340.12026-10-06
qwen/qwen3-30b-a3b-instruct-2507price.balanced0.08440.14252026-10-06
qwen/qwen3-30b-a3b-instruct-2507lmarena1383.61384.12026-10-06
qwen/qwen3-30b-a3bprice.balanced0.22750.2152026-10-06
qwen/qwen3-30b-a3blmarena13171316.62026-10-06
qwen/qwen3-235b-a22b-thinking-2507lmarena1415.11415.22026-10-06
qwen/qwen3-235b-a22b-2507price.balanced0.15310.2052026-10-06
qwen/qwen3-235b-a22b-2507lmarena1419.81419.32026-10-06
qwen/qwen3-235b-a22blmarena1366.11366.62026-10-06
qwen/qwen-2.5-coder-32b-instructlmarena1229.912302026-10-06
qwen/qwen-2.5-72b-instructlmarena1268.912692026-10-06
poolside/laguna-xs-2.1price.balanced0.0750.1252026-10-06
poolside/laguna-s-2.1price.balanced0.11250.1252026-10-06
openai/o3-mini-highlmarena1336.51336.62026-10-06
openai/o3-minilmarena1319.11319.42026-10-06
openai/o1lmarena1365.71365.82026-10-06
openai/gpt-oss-20blmarena1287.31287.72026-10-06
openai/gpt-oss-120bprice.balanced0.07030.0652026-10-06
openai/gpt-oss-120blmarena1365.91365.42026-10-06
openai/gpt-6.1-sollmarena—1445.62026-10-06
openai/gpt-6-sollmarena—1395.42026-10-06
openai/gpt-6-lunalmarena—1391.52026-10-06
openai/gpt-6-astralmarena1443.71442.32026-10-06
openai/gpt-5.6-terralmarena1446.21446.32026-10-06
openai/gpt-5.6-sol-proprice.balanced482026-10-06
openai/gpt-5.6-solprice.balanced482026-10-06
openai/gpt-5.6-sollmarena1455.11456.32026-10-06
openai/gpt-5.6-lunalmarena1429.914312026-10-06
openai/gpt-5.5lmarena1465.61466.82026-10-06
openai/gpt-5.4-nanolmarena1372.91372.12026-10-06
openai/gpt-5.4-minilmarena1412.11411.42026-10-06
openai/gpt-5.4lmarena1452.61452.22026-10-06
openai/gpt-5.2lmarena1412.41412.32026-10-06
openai/gpt-5.1lmarena1422.61422.22026-10-06
openai/gpt-5-nanolmarena1319.713202026-10-06
openai/gpt-5-minilmarena13731373.22026-10-06
openai/gpt-4o-mini-2024-07-18lmarena1286.31286.42026-10-06
openai/gpt-4o-2024-05-13lmarena1300.31300.42026-10-06
openai/gpt-4.1-nanolmarena1284.71284.82026-10-06
openai/gpt-4.1-minilmarena1340.21340.42026-10-06
openai/gpt-4.1lmarena1382.813832026-10-06
openai/gpt-4-turbolmarena1271.61271.72026-10-06
nvidia/nemotron-3-ultra-550b-a55bprice.balanced1.051.252026-10-06
nvidia/nemotron-3-super-120b-a12bprice.balanced0.17250.16382026-10-06
moonshotai/kimi-k3price.balanced63.762026-10-06
moonshotai/kimi-k3lmarena1472.31475.92026-10-06
moonshotai/kimi-k2.7-codeprice.balanced1.31711.34092026-10-06
moonshotai/kimi-k2.6price.balanced1.340.96132026-10-06
moonshotai/kimi-k2.6lmarena1454.91455.52026-10-06
moonshotai/kimi-k2.5price.balanced0.91.22026-10-06
moonshotai/kimi-k2.5lmarena1445.61445.22026-10-06
mistralai/mistral-small-3.2-24b-instructprice.balanced0.13280.10622026-10-06
mistralai/mistral-small-24b-instruct-2501lmarena1233.51233.62026-10-06
mistralai/mistral-nemoprice.balanced0.02170.0212026-10-06
mistralai/mistral-medium-3-5lmarena1420.61421.22026-10-06
minimax/minimax-m3price.balanced0.5250.4852026-10-06
minimax/minimax-m3lmarena1433.51432.12026-10-06
minimax/minimax-m2.7price.balanced0.36750.5252026-10-06
minimax/minimax-m2.5price.balanced0.47250.52122026-10-06
minimax/minimax-m2.5lmarena1358.91359.32026-10-06
minimax/minimax-m2lmarena1342.11339.92026-10-06
minimax/minimax-m1lmarena1342.71342.82026-10-06
microsoft/phi-4lmarena1216.61216.72026-10-06
meta/muse-spark-1.3lmarena1489.71489.92026-10-06
meta/muse-spark-1.2lmarena—1483.22026-10-06
meta/muse-spark-1.1lmarena1480.21479.22026-10-06
meta/muse-glimmer-30bprice.balanced0.63750.5252026-10-06
meta-llama/llama-4-maverickprice.balanced0.30370.4152026-10-06
meta-llama/llama-3.3-70b-instructprice.balanced0.1550.20132026-10-06
meta-llama/llama-3.3-70b-instructlmarena12741274.12026-10-06
meta-llama/llama-3.2-3b-instructlmarena1109.41109.52026-10-06
meta-llama/llama-3.2-1b-instructlmarena1054.51054.62026-10-06
meta-llama/llama-3.1-8b-instructprice.balanced0.05750.0252026-10-06
meta-llama/llama-3.1-70b-instructprice.balanced0.40.722026-10-06
meta-llama/llama-3.1-70b-instructlmarena1260.81260.92026-10-06
meituan/longcat-2.0price.balanced0.5251.31252026-10-06
inclusionai/ling-3.0-flash-vlprice.balanced0.03120.092026-10-06
inclusionai/ling-3.0-flash-finprice.balanced0.090.11122026-10-06
inclusionai/ling-3.0-flashprice.balanced0.03150.092026-10-06
inception/mercury-2.5price.balanced0.06750.33752026-10-06
inception/mercury-2lmarena1357.61355.22026-10-06
ibm-granite/granite-4.2-8blmarena1316.81317.82026-10-06
google/lyria-3-pro-previewprice.balanced0—2026-10-06
google/lyria-3-clip-previewprice.balanced0—2026-10-06
google/gemma-4-31b-itprice.balanced0.15250.2052026-10-06
google/gemma-4-31b-itlmarena1441.71443.42026-10-06
google/gemma-4-26b-a4b-itprice.balanced0.12110.12682026-10-06
google/gemma-4-26b-a4b-itlmarena1434.51433.92026-10-06
google/gemma-3-4b-itlmarena1290.71290.82026-10-06
google/gemma-3-27b-itprice.balanced0.17250.12026-10-06
google/gemma-3-27b-itlmarena1357.71357.82026-10-06
google/gemini-3.8-flashprice.balanced1.532026-10-06
google/gemini-3.8-flashlmarena1494.714972026-10-06
google/gemini-3.7-flashprice.balanced1.532026-10-06
google/gemini-3.7-flashlmarena1490.514872026-10-06
google/gemini-3.6-flashlmarena1476.11479.52026-10-06
google/gemini-3.5-flash-litelmarena1435.51434.72026-10-06
google/gemini-3.5-flashlmarena1475.71477.72026-10-06
google/gemini-3.1-pro-previewlmarena1480.11480.22026-10-06
google/gemini-3.1-flash-lite-previewlmarena1415.31415.72026-10-06
google/gemini-2.5-flashlmarena1417.314172026-10-06
deepseek/deepseek-v4.1-flashprice.balanced0.5250.18222026-10-06
deepseek/deepseek-v4.1-flashlmarena—1462.52026-10-06
deepseek/deepseek-v4-pro-0813price.balanced0.991.6252026-10-06
deepseek/deepseek-v4-proprice.balanced1.19411.20752026-10-06
deepseek/deepseek-v4-prolmarena14441445.22026-10-06
deepseek/deepseek-v4-flash-vision-expprice.balanced0.32340.662026-10-06
deepseek/deepseek-v4-flash-0731price.balanced0.09350.092026-10-06
deepseek/deepseek-v4-flashprice.balanced0.1750.11252026-10-06
deepseek/deepseek-v4-flashlmarena1431.81432.12026-10-06
deepseek/deepseek-v3.2-explmarena1422.61422.42026-10-06
deepseek/deepseek-v3.2price.balanced0.3150.292026-10-06
deepseek/deepseek-v3.2lmarena1424.81424.52026-10-06
deepseek/deepseek-v3.1-terminusprice.balanced0.4750.45252026-10-06
deepseek/deepseek-v3.1-terminuslmarena1416.91417.12026-10-06
deepseek/deepseek-r1-0528price.balanced0.91250.922026-10-06
deepseek/deepseek-r1-0528lmarena1427.51427.92026-10-06
deepseek/deepseek-r1lmarena1372.61372.72026-10-06
deepseek/deepseek-chat-v3.1price.balanced0.4250.45252026-10-06
deepseek/deepseek-chat-v3-0324price.balanced0.50250.43752026-10-06
deepseek/deepseek-chatprice.balanced0.45020.50022026-10-06
deepseek/deepseek-chatlmarena1332.41332.52026-10-06
cohere/command-r-plus-08-2024lmarena1228.81228.92026-10-06
arcee-ai/trinity-large-thinkinglmarena1341.813402026-10-06
anthropic/claude-sonnet-5.5lmarena—1466.62026-10-06
anthropic/claude-sonnet-5lmarena1442.21442.92026-10-06
anthropic/claude-sonnet-4.6lmarena1458.31457.82026-10-06
anthropic/claude-sonnet-4.5lmarena1438.31438.62026-10-06
anthropic/claude-sonnet-4lmarena1339.31339.12026-10-06
anthropic/claude-opus-5.5lmarena—1511.72026-10-06
anthropic/claude-opus-5lmarena1504.91502.12026-10-06
anthropic/claude-opus-4.8lmarena1460.814612026-10-06
anthropic/claude-opus-4.7lmarena14901489.92026-10-06
anthropic/claude-opus-4.6lmarena15031503.22026-10-06
anthropic/claude-opus-4.5lmarena1450.514512026-10-06
anthropic/claude-opus-4.1lmarena1418.11418.62026-10-06
anthropic/claude-haiku-4.5lmarena1396.61396.12026-10-06
anthropic/claude-fable-5.1lmarena1507.61510.62026-10-06
anthropic/claude-fable-5lmarena1492.61491.72026-10-06
amazon/nova-2-lite-v1lmarena1362.61361.82026-10-06
z-ai/glm-5.2price.balanced1.14371.10622026-09-30
z-ai/glm-5.1price.balanced1.48132.152026-09-30
qwen/qwen3.8-27bprice.balanced1.10621.0652026-09-30
qwen/qwen3-30b-a3bprice.balanced0.2150.22752026-09-30
openai/gpt-5.6-sol-proprice.balanced842026-09-30
meta/muse-glimmer-30bprice.balanced0.5250.63752026-09-30
deepseek/deepseek-v4-pro-0813price.balanced1.16610.992026-09-30
deepseek/deepseek-v4-proprice.balanced1.14491.19412026-09-30
deepseek/deepseek-v4-flashprice.balanced0.09540.1752026-09-30
z-ai/glm-5.3price.balanced1.292.152026-09-29
z-ai/glm-5.2price.balanced0.99761.14372026-09-29
z-ai/glm-5.1price.balanced1.48351.48132026-09-29
z-ai/glm-4.7price.balanced0.737512026-09-29
x-ai/grok-4.7price.balanced2.432026-09-29
qwen/qwen3.8-27bprice.balanced1.0651.10622026-09-29
qwen/qwen3.6-27bprice.balanced0.9151.042026-09-29
qwen/qwen3.5-35b-a3bprice.balanced0.54690.44692026-09-29
qwen/qwen3-vl-30b-a3b-instructprice.balanced0.22750.26252026-09-29
qwen/qwen3-next-80b-a3b-instructprice.balanced0.34250.352026-09-29
openai/gpt-oss-120bprice.balanced0.26250.07032026-09-29
openai/gpt-5.6-sol-proprice.balanced482026-09-29
nvidia/nemotron-3.5-lightningprice.balanced0.110.0852026-09-29
moonshotai/kimi-k2.7-codeprice.balanced1.35461.31712026-09-29
moonshotai/kimi-k2.6price.balanced1.71251.342026-09-29
minimax/minimax-m2.7price.balanced0.5250.36752026-09-29
minimax/minimax-m2price.balanced0.44630.5252026-09-29
inclusionai/ling-3.0-flash-vlprice.balanced0.090.03122026-09-29
google/gemma-4-26b-a4b-itprice.balanced0.14250.12112026-09-29
deepseek/deepseek-v4.1-flashprice.balanced0.20.5252026-09-29
deepseek/deepseek-v4-pro-0813price.balanced0.6931.16612026-09-29
deepseek/deepseek-v4-proprice.balanced1.1831.14492026-09-29
deepseek/deepseek-v4-flash-vision-expprice.balanced0.330.32342026-09-29
deepseek/deepseek-v4-flash-0731price.balanced0.190.09352026-09-29
deepseek/deepseek-v4-flashprice.balanced0.10310.09542026-09-29
deepseek/deepseek-v3.2price.balanced0.30180.3152026-09-29
deepseek/deepseek-v3.1-terminusprice.balanced0.45250.4752026-09-29
deepseek/deepseek-chat-v3-0324price.balanced0.43750.50252026-09-29
deepseek/deepseek-chatprice.balanced0.46250.45022026-09-29
z-ai/glm-5.3:batchprice.balanced1.0750.83752026-09-23
z-ai/glm-5.3-flash:batchprice.balanced0.11870.0952026-09-23
tencent/hy3price.balanced0.2310.14442026-09-23
qwen/qwen3.6-27bprice.balanced0.7250.9152026-09-23
openai/gpt-oss-20bprice.balanced0.0550.0362026-09-23
nvidia/nemotron-3.5-lightningprice.balanced0.10250.112026-09-23
moonshotai/kimi-k2.7-codeprice.balanced1.33211.35462026-09-23
deepseek/deepseek-v4.1-flashprice.balanced0.26250.22026-09-23
deepseek/deepseek-v4-pro-0813price.balanced0.990.6932026-09-23
deepseek/deepseek-v4-proprice.balanced1.19411.1832026-09-23
deepseek/deepseek-v4-flashprice.balanced0.11080.10312026-09-23
z-ai/glm-5.3price.balanced2.151.292026-09-22
z-ai/glm-5.2price.balanced2.150.99762026-09-22
qwen/qwen3.8-27bprice.balanced0.7981.0652026-09-22
qwen/qwen3.6-35b-a3bprice.balanced0.30.36252026-09-22
qwen/qwen3.5-35b-a3bprice.balanced0.44690.54692026-09-22
qwen/qwen3-vl-30b-a3b-instructprice.balanced0.26250.22752026-09-22
openai/gpt-oss-120bprice.balanced0.07030.26252026-09-22
nvidia/nemotron-3.5-lightningprice.balanced0.110.10252026-09-22
moonshotai/kimi-k3price.balanced5.306862026-09-22
moonshotai/kimi-k2.7-codeprice.balanced1.40751.33212026-09-22
mistralai/mistral-small-3.2-24b-instructprice.balanced0.10620.13282026-09-22
minimax/minimax-m1price.balanced0.96250.852026-09-22
meta/muse-glimmer-30bprice.balanced0.63750.5252026-09-22
gryphe/mythomax-l2-13bprice.balanced0.060.08752026-09-22
deepseek/deepseek-v4-pro-0813price.balanced0.9880.992026-09-22
deepseek/deepseek-v4-proprice.balanced21.19412026-09-22
deepseek/deepseek-v4-flash-0731price.balanced0.0750.192026-09-22
deepseek/deepseek-v4-flashprice.balanced0.10890.11082026-09-22
deepseek/deepseek-chatprice.balanced0.45020.46252026-09-22
z-ai/glm-5v-turbolmarena1436.81436.52026-09-15
z-ai/glm-5.3-flashlmarena1470.81471.92026-09-15
z-ai/glm-5.3lmarena1473.51475.12026-09-15
z-ai/glm-5.2price.balanced0.952.152026-09-15
z-ai/glm-5.2lmarena1466.71466.92026-09-15
z-ai/glm-5.1lmarena1462.71462.42026-09-15
z-ai/glm-5lmarena1446.41446.32026-09-15
z-ai/glm-4.7-flashlmarena1352.41352.22026-09-15
z-ai/glm-4.6vlmarena1376.41376.52026-09-15
z-ai/glm-4.6lmarena1440.41440.52026-09-15
z-ai/glm-4.5vlmarena1332.81332.92026-09-15
z-ai/glm-4.5-airlmarena1383.51383.72026-09-15
z-ai/glm-4.5lmarena1430.11430.22026-09-15
xiaomi/mimo-v2.5-prolmarena1465.31464.82026-09-15
xiaomi/mimo-v2.5lmarena1427.51427.42026-09-15
x-ai/grok-4.6lmarena—1429.92026-09-15
x-ai/grok-4.5lmarena1452.91450.12026-09-15
x-ai/grok-4.3lmarena1398.41397.72026-09-15
upstage/solar-pro4lmarena1377.51377.32026-09-15
thinkingmachines/inkling-smalllmarena1413.91412.32026-09-15
thinkingmachines/inklinglmarena1438.51439.72026-09-15
tencent/hy3price.balanced0.14440.2312026-09-15
tencent/hy3lmarena1439.91440.62026-09-15
stepfun/step-3.5-flashlmarena1403.81403.72026-09-15
qwen/qwen3.8-27bprice.balanced1.0650.7982026-09-15
qwen/qwen3.8-27blmarena1437.91439.32026-09-15
qwen/qwen3.7-pluslmarena1453.81454.22026-09-15
qwen/qwen3.6-pluslmarena1437.21436.72026-09-15
qwen/qwen3.6-max-previewlmarena1446.41446.32026-09-15
qwen/qwen3.5-flash-02-23lmarena1397.81397.72026-09-15
qwen/qwen3.5-397b-a17blmarena1437.61438.32026-09-15
qwen/qwen3.5-35b-a3bprice.balanced0.54690.44692026-09-15
qwen/qwen3.5-35b-a3blmarena1395.61395.52026-09-15
qwen/qwen3.5-27blmarena14081408.12026-09-15
qwen/qwen3.5-122b-a10blmarena1418.11417.92026-09-15
qwen/qwen3-vl-235b-a22b-thinkinglmarena1400.51400.82026-09-15
qwen/qwen3-vl-235b-a22b-instructlmarena1420.41420.62026-09-15
qwen/qwen3-next-80b-a3b-instructlmarena1417.31417.62026-09-15
qwen/qwen3-maxlmarena14131413.12026-09-15
qwen/qwen3-32blmarena1339.913402026-09-15
qwen/qwen3-30b-a3b-instruct-2507price.balanced0.14250.08442026-09-15
qwen/qwen3-30b-a3b-instruct-2507lmarena1383.51383.62026-09-15
qwen/qwen3-30b-a3blmarena1316.913172026-09-15
qwen/qwen3-235b-a22b-thinking-2507lmarena1414.81415.12026-09-15
qwen/qwen3-235b-a22b-2507price.balanced0.3850.15312026-09-15
qwen/qwen3-14bprice.balanced0.39810.152026-09-15
qwen/qwen-2.5-coder-32b-instructlmarena1229.81229.92026-09-15
qwen/qwen-2.5-72b-instructlmarena1268.81268.92026-09-15
openai/o4-minilmarena1353.11353.22026-09-15
openai/o3-mini-highlmarena1336.41336.52026-09-15
openai/o3-minilmarena1318.91319.12026-09-15
openai/o3lmarena1409.81409.92026-09-15
openai/o1lmarena1365.61365.72026-09-15
openai/gpt-oss-20blmarena1287.21287.32026-09-15
openai/gpt-6-astralmarena—1443.72026-09-15
openai/gpt-5.6-terralmarena1446.81446.22026-09-15
openai/gpt-5.6-sollmarena1454.91455.12026-09-15
openai/gpt-5.6-lunalmarena1430.31429.92026-09-15
openai/gpt-5.5lmarena1466.71465.62026-09-15
openai/gpt-5.4-nanolmarena13731372.92026-09-15
openai/gpt-5.4-minilmarena1412.21412.12026-09-15
openai/gpt-5.2lmarena1412.51412.42026-09-15
openai/gpt-5-nanolmarena1319.61319.72026-09-15
openai/gpt-5-minilmarena1372.913732026-09-15
openai/gpt-5lmarena1405.91406.12026-09-15
openai/gpt-4o-mini-2024-07-18lmarena1286.21286.32026-09-15
openai/gpt-4o-2024-08-06lmarena1282.31282.52026-09-15
openai/gpt-4o-2024-05-13lmarena1300.21300.32026-09-15
openai/gpt-4.1-nanolmarena1284.61284.72026-09-15
openai/gpt-4.1-minilmarena1340.11340.22026-09-15
openai/gpt-4.1lmarena1382.71382.82026-09-15
openai/gpt-4-turbolmarena1271.51271.62026-09-15
nvidia/nemotron-3-ultra-550b-a55bprice.balanced1.251.052026-09-15
nvidia/nemotron-3-super-120b-a12bprice.balanced0.16380.17252026-09-15
moonshotai/kimi-k3price.balanced4.685.30682026-09-15
moonshotai/kimi-k3lmarena14761472.32026-09-15
moonshotai/kimi-k2.6lmarena1455.11454.92026-09-15
moonshotai/kimi-k2.5lmarena1445.81445.62026-09-15
mistralai/mistral-small-24b-instruct-2501lmarena1233.41233.52026-09-15
mistralai/mistral-medium-3-5lmarena14211420.62026-09-15
mistralai/mistral-large-2407lmarena1265.91266.12026-09-15
minimax/minimax-m3lmarena1435.21433.52026-09-15
minimax/minimax-m2.5lmarena1359.21358.92026-09-15
minimax/minimax-m2lmarena1341.91342.12026-09-15
minimax/minimax-m1lmarena1342.51342.72026-09-15
microsoft/phi-4lmarena1216.51216.62026-09-15
meta/muse-spark-1.3lmarena—1489.72026-09-15
meta/muse-spark-1.1lmarena1479.41480.22026-09-15
meta/muse-glimmer-30bprice.balanced0.50.63752026-09-15
meta-llama/llama-4-maverickprice.balanced0.3240.30372026-09-15
meta-llama/llama-3.3-70b-instructlmarena1273.912742026-09-15
meta-llama/llama-3.2-3b-instructlmarena1109.31109.42026-09-15
meta-llama/llama-3.2-1b-instructlmarena1054.41054.52026-09-15
meta-llama/llama-3.1-8b-instructlmarena1186.41186.52026-09-15
meta-llama/llama-3.1-70b-instructlmarena1260.71260.82026-09-15
ibm-granite/granite-4.2-8blmarena1318.41316.82026-09-15
google/gemma-4-31b-itlmarena1441.81441.72026-09-15
google/gemma-4-26b-a4b-itprice.balanced0.08650.14252026-09-15
google/gemma-4-26b-a4b-itlmarena1434.91434.52026-09-15
google/gemma-3-4b-itlmarena1290.61290.72026-09-15
google/gemma-3-27b-itlmarena1357.51357.72026-09-15
google/gemma-3-12b-itlmarena13341334.22026-09-15
google/gemma-2-27b-itlmarena1231.21231.32026-09-15
google/gemini-3.8-flashlmarena1495.21494.72026-09-15
google/gemini-3.7-flashlmarena1491.41490.52026-09-15
google/gemini-3.6-flashlmarena1476.31476.12026-09-15
google/gemini-3.5-flash-litelmarena1436.41435.52026-09-15
google/gemini-3.5-flashlmarena1476.81475.72026-09-15
google/gemini-2.5-flashlmarena1417.21417.32026-09-15
deepseek/deepseek-v4-pro-0813price.balanced0.86920.9882026-09-15
deepseek/deepseek-v4-proprice.balanced1.184722026-09-15
deepseek/deepseek-v4-prolmarena1439.814442026-09-15
deepseek/deepseek-v4-flash-0731price.balanced0.09380.0752026-09-15
deepseek/deepseek-v4-flashprice.balanced0.10710.10892026-09-15
deepseek/deepseek-v4-flashlmarena1431.91431.82026-09-15
deepseek/deepseek-v3.1-terminuslmarena1416.71416.92026-09-15
deepseek/deepseek-r1-0528lmarena1427.41427.52026-09-15
deepseek/deepseek-r1lmarena1372.51372.62026-09-15
deepseek/deepseek-chat-v3-0324lmarena1374.913752026-09-15
deepseek/deepseek-chatlmarena1332.31332.42026-09-15
cohere/command-r-plus-08-2024lmarena1228.71228.82026-09-15
cohere/command-r-08-2024lmarena1187.11187.32026-09-15
arcee-ai/trinity-large-thinkinglmarena13421341.82026-09-15
anthropic/claude-sonnet-5lmarena1443.21442.22026-09-15
anthropic/claude-sonnet-4.6lmarena14581458.32026-09-15
anthropic/claude-sonnet-4lmarena1339.11339.32026-09-15
anthropic/claude-opus-5lmarena15051504.92026-09-15
anthropic/claude-opus-4.8lmarena1461.71460.82026-09-15
anthropic/claude-opus-4.7lmarena1490.314902026-09-15
anthropic/claude-opus-4.6lmarena1503.115032026-09-15
anthropic/claude-opus-4.5lmarena1450.41450.52026-09-15
anthropic/claude-opus-4.1lmarena14181418.12026-09-15
anthropic/claude-opus-4lmarena1365.61365.72026-09-15
anthropic/claude-haiku-4.5lmarena1395.11396.62026-09-15
anthropic/claude-fable-5.1lmarena1513.71507.62026-09-15
anthropic/claude-fable-5lmarena1494.11492.62026-09-15
anthropic/claude-3-haikulmarena1194.51194.62026-09-15
amazon/nova-2-lite-v1lmarena1362.51362.62026-09-15
z-ai/glm-5.3-flashprice.balanced0.11280.23752026-09-11
z-ai/glm-5.3-flashlmarena—1470.82026-09-11
z-ai/glm-5.3lmarena1476.91473.52026-09-11
z-ai/glm-5.2price.balanced1.48350.952026-09-11
z-ai/glm-5.2lmarena1465.41466.72026-09-11
z-ai/glm-5.1lmarena1464.11462.72026-09-11
z-ai/glm-5lmarena1445.21446.42026-09-11
z-ai/glm-4.7-flashlmarena1352.91352.42026-09-11
z-ai/glm-4.7lmarena1435.31435.92026-09-11
z-ai/glm-4.6vlmarena1374.71376.42026-09-11
z-ai/glm-4.6price.balanced0.96250.762026-09-11
z-ai/glm-4.6lmarena1439.81440.42026-09-11
z-ai/glm-4.5vlmarena1333.61332.82026-09-11
z-ai/glm-4.5-airlmarena1382.81383.52026-09-11
z-ai/glm-4.5lmarena1429.41430.12026-09-11
xiaomi/mimo-v2.5-prolmarena14651465.32026-09-11
xiaomi/mimo-v2.5lmarena1427.31427.52026-09-11
x-ai/grok-4.6lmarena1443.7—2026-09-11
x-ai/grok-4.5lmarena1452.31452.92026-09-11
x-ai/grok-4.3lmarena1397.51398.42026-09-11
upstage/solar-pro4price.balanced0.05250.15752026-09-11
upstage/solar-pro4lmarena1376.21377.52026-09-11
thinkingmachines/inkling-smalllmarena1411.71413.92026-09-11
thinkingmachines/inklinglmarena1439.21438.52026-09-11
tencent/hy3lmarena1441.21439.92026-09-11
qwen/qwen3.8-27blmarena1440.81437.92026-09-11
qwen/qwen3.7-pluslmarena1456.21453.82026-09-11
qwen/qwen3.6-pluslmarena1436.81437.22026-09-11
qwen/qwen3.6-max-previewlmarena1446.31446.42026-09-11
qwen/qwen3.5-flash-02-23lmarena1397.61397.82026-09-11
qwen/qwen3.5-397b-a17blmarena1438.31437.62026-09-11
qwen/qwen3.5-27blmarena1407.914082026-09-11
qwen/qwen3.5-122b-a10bprice.balanced0.81750.7152026-09-11
qwen/qwen3.5-122b-a10blmarena1417.91418.12026-09-11
qwen/qwen3-vl-235b-a22b-thinkinglmarena1400.61400.52026-09-11
qwen/qwen3-vl-235b-a22b-instructlmarena1420.91420.42026-09-11
qwen/qwen3-next-80b-a3b-thinkinglmarena1367.51367.82026-09-11
qwen/qwen3-next-80b-a3b-instructprice.balanced0.350.34252026-09-11
qwen/qwen3-next-80b-a3b-instructlmarena1418.61417.32026-09-11
qwen/qwen3-maxlmarena1412.714132026-09-11
qwen/qwen3-32blmarena1340.11339.92026-09-11
qwen/qwen3-30b-a3b-instruct-2507price.balanced0.08440.14252026-09-11
qwen/qwen3-30b-a3b-instruct-2507lmarena1384.31383.52026-09-11
qwen/qwen3-235b-a22b-thinking-2507lmarena1413.81414.82026-09-11
qwen/qwen3-235b-a22b-2507price.balanced0.2050.3852026-09-11
qwen/qwen3-235b-a22b-2507lmarena1419.41419.82026-09-11
qwen/qwen3-235b-a22blmarena13661366.12026-09-11
qwen/qwen-2.5-coder-32b-instructlmarena1230.21229.82026-09-11
qwen/qwen-2.5-72b-instructlmarena1269.11268.82026-09-11
openai/o4-minilmarena1352.51353.12026-09-11
openai/o3-mini-highlmarena1336.61336.42026-09-11
openai/o3-minilmarena1319.21318.92026-09-11
openai/o3lmarena1409.21409.82026-09-11
openai/o1lmarena1365.91365.62026-09-11
openai/gpt-oss-20blmarena1287.81287.22026-09-11
openai/gpt-oss-120blmarena1365.61365.92026-09-11
openai/gpt-5.6-terralmarena14461446.82026-09-11
openai/gpt-5.6-sollmarena1454.21454.92026-09-11
openai/gpt-5.6-lunalmarena1428.51430.32026-09-11
openai/gpt-5.5lmarena1465.81466.72026-09-11
openai/gpt-5.4-nanolmarena1372.813732026-09-11
openai/gpt-5.4-minilmarena1412.11412.22026-09-11
openai/gpt-5.2lmarena1411.91412.52026-09-11
openai/gpt-5.1lmarena1422.21422.62026-09-11
openai/gpt-5-nanolmarena1320.31319.62026-09-11
openai/gpt-5-minilmarena1373.41372.92026-09-11
openai/gpt-5lmarena1405.61405.92026-09-11
openai/gpt-4o-mini-2024-07-18lmarena1286.61286.22026-09-11
openai/gpt-4o-2024-08-06lmarena1282.71282.32026-09-11
openai/gpt-4o-2024-05-13lmarena1300.61300.22026-09-11
openai/gpt-4.1-nanolmarena1284.81284.62026-09-11
openai/gpt-4.1-minilmarena1340.51340.12026-09-11
openai/gpt-4.1lmarena1382.31382.72026-09-11
openai/gpt-4-turbolmarena1271.81271.52026-09-11
moonshotai/kimi-k3price.balanced64.682026-09-11
moonshotai/kimi-k3lmarena1476.314762026-09-11
moonshotai/kimi-k2.6lmarena14551455.12026-09-11
moonshotai/kimi-k2.5lmarena1445.21445.82026-09-11
mistralai/mistral-small-24b-instruct-2501lmarena1233.61233.42026-09-11
mistralai/mistral-medium-3-5lmarena1420.914212026-09-11
mistralai/mistral-large-2407lmarena1266.31265.92026-09-11
minimax/minimax-m3lmarena1434.81435.22026-09-11
minimax/minimax-m2.7lmarena1405.31404.62026-09-11
minimax/minimax-m2.5lmarena13591359.22026-09-11
minimax/minimax-m2lmarena1342.11341.92026-09-11
minimax/minimax-m1lmarena1341.91342.52026-09-11
microsoft/phi-4lmarena1216.81216.52026-09-11
meta/muse-spark-1.1lmarena1478.31479.42026-09-11
meta-llama/llama-3.3-70b-instructlmarena1274.91273.92026-09-11
meta-llama/llama-3.2-3b-instructlmarena1109.71109.32026-09-11
meta-llama/llama-3.2-1b-instructlmarena1054.81054.42026-09-11
meta-llama/llama-3.1-8b-instructlmarena1186.71186.42026-09-11
meta-llama/llama-3.1-70b-instructlmarena12611260.72026-09-11
inception/mercury-2lmarena1357.81357.62026-09-11
ibm-granite/granite-4.2-8bprice.balanced0.11250.10752026-09-11
ibm-granite/granite-4.2-8blmarena—1318.42026-09-11
google/gemma-4-31b-itlmarena1441.71441.82026-09-11
google/gemma-4-26b-a4b-itprice.balanced0.13750.08652026-09-11
google/gemma-4-26b-a4b-itlmarena1434.61434.92026-09-11
google/gemma-3-4b-itlmarena1290.81290.62026-09-11
google/gemma-3-27b-itlmarena1358.31357.52026-09-11
google/gemma-3-12b-itlmarena1334.213342026-09-11
google/gemma-2-27b-itlmarena1231.51231.22026-09-11
google/gemini-3.8-flashlmarena—1495.22026-09-11
google/gemini-3.7-flashlmarena1490.21491.42026-09-11
google/gemini-3.6-flashlmarena1476.51476.32026-09-11
google/gemini-3.5-flash-litelmarena1436.51436.42026-09-11
google/gemini-3.5-flashlmarena1475.21476.82026-09-11
google/gemini-3.1-pro-previewlmarena1479.61480.12026-09-11
google/gemini-3.1-flash-lite-previewlmarena1414.81415.32026-09-11
google/gemini-2.5-prolmarena1457.31457.82026-09-11
google/gemini-2.5-flashlmarena1417.31417.22026-09-11
deepseek/deepseek-v4-proprice.balanced1.19151.18472026-09-11
deepseek/deepseek-v4-prolmarena1439.21439.82026-09-11
deepseek/deepseek-v4-flashprice.balanced0.10690.10712026-09-11
deepseek/deepseek-v4-flashlmarena1431.61431.92026-09-11
deepseek/deepseek-v3.2lmarena1424.61424.82026-09-11
deepseek/deepseek-v3.1-terminuslmarena1416.61416.72026-09-11
deepseek/deepseek-r1-0528lmarena1427.91427.42026-09-11
deepseek/deepseek-r1lmarena1372.71372.52026-09-11
deepseek/deepseek-chat-v3.1lmarena1419.11419.72026-09-11
deepseek/deepseek-chat-v3-0324price.balanced0.50250.43752026-09-11
deepseek/deepseek-chat-v3-0324lmarena1375.11374.92026-09-11
deepseek/deepseek-chatprice.balanced0.46250.45022026-09-11
deepseek/deepseek-chatlmarena1332.61332.32026-09-11
cohere/command-r-plus-08-2024lmarena12291228.72026-09-11
cohere/command-r-08-2024lmarena1187.51187.12026-09-11
arcee-ai/trinity-large-thinkinglmarena1341.913422026-09-11
anthropic/claude-sonnet-5lmarena1442.31443.22026-09-11
anthropic/claude-sonnet-4.6lmarena1457.614582026-09-11
anthropic/claude-sonnet-4.5lmarena14381438.32026-09-11
anthropic/claude-sonnet-4lmarena1338.21339.12026-09-11
anthropic/claude-opus-5lmarena1504.215052026-09-11
anthropic/claude-opus-4.8lmarena1460.91461.72026-09-11
anthropic/claude-opus-4.7lmarena1490.11490.32026-09-11
anthropic/claude-opus-4.6lmarena1502.71503.12026-09-11
anthropic/claude-opus-4.5lmarena1449.91450.42026-09-11
anthropic/claude-opus-4.1lmarena1417.614182026-09-11
anthropic/claude-opus-4lmarena1364.31365.62026-09-11
anthropic/claude-haiku-4.5lmarena1394.91395.12026-09-11
anthropic/claude-fable-5.1lmarena—1513.72026-09-11
anthropic/claude-fable-5lmarena1494.71494.12026-09-11
anthropic/claude-3-haikulmarena1194.81194.52026-09-11
amazon/nova-2-lite-v1lmarena1362.21362.52026-09-11
z-ai/glm-5.3-flash:batchprice.balanced0.23750.11872026-09-08
z-ai/glm-5.3-flashprice.balanced0.11870.11282026-09-08
z-ai/glm-4.7-flashprice.balanced0.1450.14542026-09-08
qwen/qwen3.5-35b-a3bprice.balanced0.24750.54692026-09-08
qwen/qwen3-235b-a22b-2507price.balanced0.15310.2052026-09-08
qwen/qwen3-14bprice.balanced0.150.39812026-09-08
moonshotai/kimi-k2.7-codeprice.balanced1.3451.40752026-09-08
meta/muse-glimmer-30b:batchprice.balanced0.63750.31872026-09-08
deepseek/deepseek-v4-pro-0813:batchprice.balanced1.980.992026-09-08
deepseek/deepseek-v4-proprice.balanced1.25241.19152026-09-08
deepseek/deepseek-v4-flash-0731:batchprice.balanced0.1750.1652026-09-08
deepseek/deepseek-v4-flashprice.balanced0.10970.10692026-09-08
deepseek/deepseek-chat-v3.1price.balanced0.8250.4252026-09-08
deepseek/deepseek-chat-v3-0324price.balanced0.43750.50252026-09-08
undi95/remm-slerp-l2-13bprice.balanced0.50.4252026-09-04
qwen/qwen3.6-27bprice.balanced1.350.7252026-09-04
qwen/qwen3.5-35b-a3bprice.balanced0.50.24752026-09-04
nvidia/nemotron-3-ultra-550b-a55bprice.balanced1.051.252026-09-04
deepseek/deepseek-v4-pro-0813price.balanced0.990.86922026-09-04
deepseek/deepseek-v4-proprice.balanced1.28851.25242026-09-04
deepseek/deepseek-v4-flashprice.balanced0.09920.10972026-09-04
z-ai/glm-4.6price.balanced0.760.96252026-09-03
tencent/hy3price.balanced0.2310.14442026-09-03
qwen/qwen3.8-27bprice.balanced0.95621.0652026-09-03
qwen/qwen3.5-397b-a17bprice.balanced0.87751.28752026-09-03
qwen/qwen2.5-vl-72b-instructprice.balanced0.3750.852026-09-03
nvidia/nemotron-3-ultra-550b-a55bprice.balanced1.251.052026-09-03
meta/muse-glimmer-30bprice.balanced0.5250.52026-09-03
meta-llama/llama-3.3-70b-instructprice.balanced0.710.1552026-09-03
google/gemini-3.7-flash:batchprice.balanced0.3750.752026-09-03
deepseek/deepseek-v4-pro-0813price.balanced1.67310.992026-09-03
deepseek/deepseek-v4-proprice.balanced1.30281.28852026-09-03
deepseek/deepseek-v4-flashprice.balanced0.11080.09922026-09-03
deepseek/deepseek-chat-v3.1price.balanced0.4250.8252026-09-03
deepseek/deepseek-chatprice.balanced0.45020.46252026-09-03
z-ai/glm-5.2price.balanced1.82751.48352026-09-02
z-ai/glm-5.1price.balanced1.9351.48352026-09-02
z-ai/glm-4.6price.balanced0.8750.762026-09-02
thinkingmachines/inklingprice.balanced1.7251.76252026-09-02
tencent/hy3price.balanced0.14440.2312026-09-02
qwen/qwen3.6-35b-a3bprice.balanced0.3550.32026-09-02
qwen/qwen3.6-27bprice.balanced1.041.352026-09-02
qwen/qwen3.5-122b-a10bprice.balanced0.7150.81752026-09-02
qwen/qwen3-vl-30b-a3b-instructprice.balanced0.22750.26252026-09-02
qwen/qwen3-235b-a22b-2507price.balanced0.2050.15312026-09-02
nvidia/nemotron-3-ultra-550b-a55bprice.balanced1.351.252026-09-02
moonshotai/kimi-k2.7-codeprice.balanced1.35251.3452026-09-02
mistralai/devstral-2512price.balanced0.880.82026-09-02
meta/muse-glimmer-30bprice.balanced0.63750.5252026-09-02
meta-llama/llama-4-scoutprice.balanced0.16750.152026-09-02
meta-llama/llama-4-maverickprice.balanced0.350.3242026-09-02
mancer/weaverprice.balanced0.56250.48752026-09-02
google/gemini-3.7-flashprice.balanced0.751.52026-09-02
deepseek/deepseek-v4-pro-0813price.balanced1.981.67312026-09-02
deepseek/deepseek-v4-proprice.balanced1.08751.30282026-09-02
deepseek/deepseek-v4-flash-0731price.balanced0.0750.09382026-09-02
deepseek/deepseek-v4-flashprice.balanced0.10130.11082026-09-02
deepseek/deepseek-v3.2price.balanced0.290.30182026-09-02
deepseek/deepseek-chat-v3.1price.balanced0.8250.4252026-09-02
arcee-ai/trinity-large-thinkingprice.balanced0.37750.38752026-09-02
anthracite-org/magnum-v4-72bprice.balanced3.53.1252026-09-02
qwen/qwen3-235b-a22b-2507lmarena1419.31419.42026-08-30
openai/gpt-5.5lmarena1470.91465.82026-08-30
openai/gpt-5.4lmarena1469.61452.62026-08-30
openai/gpt-5.2lmarena1416.51411.92026-08-30
openai/gpt-5.1lmarena1441.41422.22026-08-30
google/gemini-3.5-flashlmarena1482.61475.22026-08-30
google/gemini-3.1-pro-previewlmarena1479.51479.62026-08-30
deepseek/deepseek-v3.2-explmarena1424.41422.62026-08-30
deepseek/deepseek-v3.1-terminuslmarena1419.61416.62026-08-30
deepseek/deepseek-chat-v3-0324lmarena13751375.12026-08-30
anthropic/claude-opus-4.5lmarena1449.81449.92026-08-30
allenai/olmo-3-32b-thinklmarena1298.51298.62026-08-30
x-ai/grok-4.20lmarena1444.4—2026-08-28
qwen/qwen3.8-maxlmarena—1481.92026-08-28
qwen/qwen3.7-maxlmarena1474.3—2026-08-28
qwen/qwen3.6-max-previewlmarena—1446.32026-08-28
qwen/qwen3-vl-235b-a22b-thinkinglmarena—1400.62026-08-28
qwen/qwen3-next-80b-a3b-thinkinglmarena—1367.52026-08-28
qwen/qwen3-maxlmarena1439.11412.72026-08-28
openai/o4-minilmarena—1352.52026-08-28
openai/o3-mini-highlmarena—1336.62026-08-28
openai/o3-minilmarena1336.61319.22026-08-28
openai/o3lmarena—1409.22026-08-28
openai/o1lmarena1352.91365.92026-08-28
openai/gpt-4.1-nanolmarena—1284.82026-08-28
openai/gpt-4.1-minilmarena—1340.52026-08-28
openai/gpt-4.1lmarena—1382.32026-08-28
openai/gpt-4-turbolmarena—1271.82026-08-28
moonshotai/kimi-k2-0905lmarena1379.1—2026-08-28
moonshotai/kimi-k2lmarena1370.9—2026-08-28
minimax/minimax-m2.1lmarena1391.1—2026-08-28
google/gemini-3.1-pro-previewlmarena—1479.52026-08-28
google/gemini-3.1-flash-lite-previewlmarena—1414.82026-08-28
google/gemini-3.1-flash-litelmarena1414.8—2026-08-28
deepseek/deepseek-v4-prolmarena1450.91439.22026-08-28
cohere/command-r-08-2024lmarena—1187.52026-08-28
arcee-ai/trinity-large-thinkinglmarena—1341.92026-08-28
anthropic/claude-sonnet-4.5lmarena—14382026-08-28
anthropic/claude-sonnet-4lmarena—1338.22026-08-28
anthropic/claude-opus-4.5lmarena—1449.82026-08-28
anthropic/claude-opus-4.1lmarena—1417.62026-08-28
anthropic/claude-opus-4lmarena—1364.32026-08-28
anthropic/claude-haiku-4.5lmarena—1394.92026-08-28
anthropic/claude-3-haikulmarena—1194.82026-08-28
z-ai/glm-5.2price.balanced1.48351.82752026-08-26
z-ai/glm-5.1price.balanced1.48351.9352026-08-26
tencent/hy3price.balanced0.2310.14442026-08-26
qwen/qwen3.8-27bprice.balanced1.050.95622026-08-26
qwen/qwen3.6-27bprice.balanced1.351.042026-08-26
qwen/qwen3-30b-a3bprice.balanced0.22750.2152026-08-26
qwen/qwen2.5-vl-72b-instructprice.balanced0.850.3752026-08-26
moonshotai/kimi-k2.6price.balanced0.97611.71252026-08-26
minimax/minimax-m2.7price.balanced0.420.5252026-08-26
meta-llama/llama-4-scoutprice.balanced0.150.16752026-08-26
meta-llama/llama-3.3-70b-instructprice.balanced0.1550.712026-08-26
google/gemma-4-31b-itprice.balanced0.160.15252026-08-26
deepseek/deepseek-v4-pro-0813price.balanced1.6831.982026-08-26
deepseek/deepseek-v4-proprice.balanced0.49611.08752026-08-26
deepseek/deepseek-v4-flash-0731price.balanced0.1050.0752026-08-26
deepseek/deepseek-v4-flashprice.balanced0.06110.10132026-08-26
Evidence & Ask