OpenRouter or direct
For 143 of the 320 models carrying published offers, the only seller is the lab that made the model. Routing cannot quote you a price the lab did not set, so on those models the question is never "which is cheaper per token" — it is whether one integration is worth a fee. On 46 models it is a real market: two or more independent sellers, at equal-or-better declared precision, with a gap between them.
Offers as of 2026-08-28. Prices are blended per million tokens at each workload's input:output ratio. Counts are over distinct sellers, never offer rows.
What this page cannot tell you
This repository holds no direct list price from any provider. Every offer in data/offers.json arrived from OpenRouter's own endpoints API. So we can measure how
many sellers the marketplace lists for a model and what it charges for each — and we cannot put
a direct price beside it, because no vendor published one to us.
That is why there is no "OpenRouter $X vs direct $Y" table below. Such a table would need a number we would have had to invent, and an invented price is the one bug this project treats as unrecoverable. What follows is the structure of the decision, the counts that constrain it, and the formula you can fill in with your own two prices.
A second limit: an offer row carries an input rate, an output rate and sometimes a cached-input rate. It carries no context-tier ladder and no separate reasoning rate. Every price here is therefore a floor on what a row costs, not a final bill. 1053 offer rows across 320 models collapse to 911 distinct seller quotes.
The break-even
Routing is cheaper than going direct only when the rate it finds, plus the marketplace's fee, beats the direct rate by more than the cost of running one more integration yourself.
- S
- your monthly volume, in millions of tokens at your input:output mix
- pdirect
- the provider's own blended rate, $/M, for that mix
- proute
- the cheapest marketplace rate at equal-or-better declared precision, $/M
- f
- the marketplace's fee on money you load, as a fraction
- C
- monthly cost of one more direct integration: a second key to rotate, a second invoice, a second quota, a second thing to monitor
Routing wins while S · proute · (1 + f) < S · pdirect + C
Break-even volume S* = C ÷ (pdirect − proute(1 + f))
Only the denominator matters. If proute ≥ pdirect, it is never positive and no volume makes routing cheaper on price — the break-even does not exist rather than being large. That is not a corner case here. OpenRouter states it passes provider pricing through without markup, so on the 143 models where the lab is the only seller, proute and pdirect are the same number and the denominator reduces to −pdirect·f.
Read the other direction, the formula says what routing is actually selling: not a lower rate, but a lower C. High volume on few models pushes toward direct. Low volume across many models pushes toward routing. The crossover is set by your engineering time, which is the one term on this page nobody else can measure for you.
Illustrative Every figure in this paragraph is made up to show the shape of the arithmetic; none is measured. Take f = 0.055, a direct blended rate of $1.20/M, and a model the marketplace resells at the same $1.20/M. Routing costs $1.266/M. At 500 M tokens a month that is $33 a month of fee. If running the second integration yourself costs less than $33 a month of attention, direct wins; if it costs more, routing wins — and it wins by the same margin at any volume, because the fee and the saving both scale with S.
Five scenarios, and how many models are in each
The split is on independent sellers: distinct sellers other than the lab named in the model slug. That is the count that decides whether routing has anyone else to ask.
| Scenario | Models | What routing can do about price | Where to look |
|---|---|---|---|
| Lab is the only seller | 143 | Nothing. The quote is the lab's own, so routing adds f and subtracts nothing. Heavy single-model spend belongs direct unless C is large. | /check/ |
| One independent seller | 76 | Produce exactly one alternative quote. Whether it beats direct depends on a direct price this repository does not hold — you have to look it up. | /spreads/ |
| Several sellers, a real gap | 46 | Find a genuinely different price. Median gap 1.25× between cheapest and dearest independent seller at equal-or-better declared precision; widest 7.56×. This is where the marketplace earns its fee. | /spreads/ |
| Several sellers, one precision | 42 | Offer cheaper quotes that are not like-for-like. Only one seller declares the best precision on offer, so the saving underneath it is a trade, not a discount. | /spreads/ |
| Several sellers, same price | 13 | Quote the same number twice. Nothing to shop for. | /spreads/ |
These five partition the population exactly: 143 + 76 + 46
+ 42 + 13 = 320. Precision ranking treats an
undeclared quantisation as its own tier — unknown is never counted as equal to a
declared bf16.
Identifying the lab behind a seller is name-based, with a short hand-maintained alias list for the cases where the strings differ (qwen, bytedance-seed). A miss counts a lab as independent, which inflates the independent-seller count — so 143 is a floor on how many models have nobody but the lab quoting them, never a ceiling.
The workload changes the size of the gap, not whether one exists
The 143 / 76 / 101 split is identical under all 6 workloads. That is expected and worth stating: how many sellers exist is a fact about the market, not about your token mix. What the mix moves is how much the gap between them is worth.
| Workload | Output share | Cache hit rate | Models with a gap | Median gap | Widest gap |
|---|---|---|---|---|---|
| Balanced | 25% | 0% | 46 | 1.25× | 7.56× |
| Summarisation | 5% | 30% | 46 | 1.41× | 13.81× |
| Chat | 40% | 50% | 46 | 1.27× | 7.13× |
| Code generation | 60% | 20% | 45 | 1.3× | 5.46× |
| Agentic loop | 15% | 70% | 46 | 1.5× | 13.5× |
| Classification | 1% | 10% | 45 | 1.39× | 13.13× |
Output share is output tokens as a fraction of the total; cache hit rate is the share of input billed at the cached rate. An output-heavy workload pushes the comparison onto output rates, where sellers of the same open-weights model diverge most. A cache-heavy one flatters whoever publishes a cached-input rate at all.
Why "how many sellers" is three different numbers
Google: Gemini 3.7 Flash is the clearest case on this snapshot. It carries 6 offer rows. Counted three ways, it gives three answers:
| Counted over | Count | Spread | Verdict |
|---|---|---|---|
| Offer rows | 6 | 7.2× | Wrong. This figure shipped once and is in the correction log. |
| Distinct sellers | 2 | 2× | Right for "is this worth shopping". Still not this page's question. |
| Independent sellers | 0 | — | The answer here. Nobody but the lab quotes it, so there is no spread to have. |
The rows are one vendor's flex, standard and priority service tiers. A rate card is not a market. On this snapshot 48 of 320 models have a row-level spread wider than their seller-level spread for the same reason, which is why every count on this page is over distinct sellers.
The mirror-image trap is precision. DeepSeek: DeepSeek V4 Flash 0731 has 29 independent sellers and a 13.2× spread between them — but exactly one of them declares the best precision on offer. The cheap end is a different product. Counting that spread as a saving is how a page tells you to downgrade without saying so.
What routing is worth, stated honestly
A gateway is not a discount, and OpenRouter does not claim it is. Its FAQ says it "pass[es] through the pricing of the underlying model providers without any markup, so you pay the same rate as you would directly with the provider", and that it charges a fee when you purchase credits. Both halves of that sentence are load-bearing: the token rate is not the product, and the fee is real.
What you get for the fee is genuine, and none of it is a lower rate:
- One integration instead of N. That is the C term, and for a team evaluating several models a month it is the whole argument.
- A second seller to fail over to, on the 101 models where more than one independent seller exists.
- Trying a model without opening an account with its lab. The option has value even when the rate does not.
- The fee is charged on credit you load, not per request, so it is a fixed fraction rather than something that compounds with volume.
Rates read from OpenRouter's FAQ on 2026-08-27: 5.5% with a $0.80 minimum on card credit purchases, 5% on crypto, and a plan-dependent BYOK fee above a free allowance. These are OpenRouter's published figures, not an Undominated measurement. Nothing in this repository re-fetches or verifies them, so treat them as a pointer to the source rather than a current fact — a stale fee is a wrong number, and the formula above is written in terms of f precisely so that it survives the rate changing. Note the minimum: on a credit purchase below about $14.55 the $0.80 floor costs more than 5.5% would.
We take no fee, commission or referral from any link on this page, and we do not sell inference. See independence.
Answer it for your own bill instead
Everything above is about the market. Your case turns on two prices and one number only you have, so the useful next step is measurement, not more reading.
- Bill audit — drop an OpenRouter activity export. It is parsed in your browser and never uploaded. It tells you which models you actually pay for, at what mix, and which of them something better and cheaper beats.
- Spreads — which models are worth shopping across distinct sellers at equal-or-better precision, and which apparent bargains are precision trades.
- Check a model — whether one model is on the frontier or beaten, with the switching-cost payback.
- Methodology — how effective price, precision ranking and capability-safe dominance are computed, including what is missing.
Offer data as of 2026-08-28, from OpenRouter's public model-endpoints API. Recomputed on every
build by scripts/compute-openrouter-scenarios.mjs; if a count here disagrees with
that script, the script is right.