MODEL PROOF
Kimi K3
BETTER VALUE OPTION
A stored alternative has an equal-or-higher measured score and an equal-or-lower price.
Serving precision differs between offers.
Muse Spark 1.3 · Δ score 17.4 points · 67% lower measured price · $2.00 / 1M tokens
Gemini 3.8 Flash is a further stored option with these losses: 944K → 66K max output
- Independent LMArena score
- 1,472.3 ± 5.22 · 20,987 votes
- Context window
- 1.05M
- Maximum output
- 943.72K
- Input modalities
- text, image, video
- Output modalities
- text
- Published input price
- $3.00 / 1M tokens
- Published output price
- $15.00 / 1M tokens
- Pricing kind
- fixed
Inspect complete billing conditions and endpoint terms below.
Every price dimension · USD per million tokens
| Input | $3.00/M |
|---|---|
| Output | $15.00/M |
| Cached input | $0.300/M |
| Cache write | Unknown/M |
| Cache write 1h | Unknown/M |
| Reasoning | Unknown/M |
Who sells it
The same weights, different shops. Cheapest is not like-for-like when serving precision differs.
Headline cheapest is a lower precision. Like-for-like at the best declared precision is $2.85/M from DeepInfra.
20 provider offers · rates, limits and conditions
Rates are USD per million tokens. Endpoint terms can differ even when the quoted price and precision match. A listed parameter is a provider declaration, not a task-success test.
| Seller | Input | Output | Precision | Uptime (1d) |
|---|---|---|---|---|
| Sail Researchsail-research/fp4 Endpoint terms
| $1.50 | $10.76 | fp4 | 99.84% |
| InferenceNetinference-net/fp4 Endpoint terms
| $1.95 | $9.75 | fp4 | 98.59% |
| Relacerelace/fp4 Endpoint terms
| $1.95 | $9.75 | fp4 | 98.71% |
| Phalaphala Endpoint terms
| $2.10 | $10.50 | undeclared | 99.47% |
| Waferwafer Endpoint terms
| $2.49 | $14.50 | undeclared | 99.92% |
| Morphmorph Endpoint terms
| $2.50 | $14.00 | undeclared | 100.00% |
| Makoramakora Endpoint terms
| $2.55 | $12.75 | undeclared | 98.21% |
| DigitalOceandigitalocean Endpoint terms
| $2.55 | $12.95 | undeclared | 99.64% |
| DeepInfradeepinfra/bf16 Endpoint terms
| $2.85 | $14.25 | bf16 | 94.48% |
| BaseTenbaseten/fp8 Endpoint terms
| $3.00 | $15.00 | fp8 | 99.14% |
| Parasailparasail/fp4 Endpoint terms
| $3.00 | $15.00 | fp4 | 98.24% |
| Chuteschutes/mxfp4 Endpoint terms
| $3.00 | $15.00 | mxfp4 | 99.07% |
| Modalmodal/mxfp4 Endpoint terms
| $3.00 | $15.00 | mxfp4 | 99.59% |
| Togethertogether Endpoint terms
| $3.00 | $15.00 | undeclared | 99.54% |
| Fireworksfireworks Endpoint terms
| $3.00 | $15.00 | undeclared | 99.84% |
| Moonshot AImoonshotai/mxfp4 Endpoint terms
| $3.00 | $15.00 | mxfp4 | 99.83% |
| Fireworksfireworks/us Endpoint terms
| $3.30 | $16.50 | undeclared | 99.05% |
| Alibabaalibaba Endpoint terms
| $3.45 | $17.25 | undeclared | 98.86% |
| Fireworksfireworks/fast Endpoint terms
| $4.50 | $22.50 | undeclared | 99.70% |
| Morphmorph/fast Endpoint terms
| $6.00 | $22.50 | undeclared | 100.00% |
Independent scores
| LMArena Elo | 1472.3max |
|---|
LMArena Elo under CC BY 4.0. Absence of another board is not a score of zero.
Best rankings by task
- codecategories #1of 108 1379
- uicomponent #1of 104 1370
- website #2of 114 1349
- gamedev #2of 107 1397
- fullstack #2of 41 1325
- htmlslides #2of 23 1259
- 3d #3of 100 1422
- mobileapps #3of 40 1282
Capability
| Context window | 1M |
|---|---|
| Max output | 944K |
| Input modes | text, image, video |
| Tool use | yes |
| Reasoning | optional |
| Open weights | yes |
Provenance
| Price source | openrouter.ai |
|---|---|
| Fetched | 2026-09-23 |
| Quality data | verified |
| Cross-checked | vendor page |
Our take
editorial — not a measurementKimi is a candidate for teams comparing hosted multimodal use with self-hosting. The accepted record lists open weights, but hardware and licence requirements still need review. Use the current proof and benchmark evidence for measured standing. Other models’ withheld latency figures cannot establish a speed comparison; measure responsiveness on your intended deployment.
Strengths
- Open weights are recorded
- Text, image and video input are listed
Weaknesses
- Measure first-token latency and sustained generation separately on your deployment
- The context window and output ceiling are different limits
Reach for it when
- Hosted-versus-self-hosted evaluations
- Multimodal workflows with an explicit latency test
Avoid it if
- The decision depends on a latency advantage you have not measured
- You have not checked serving capacity and licence terms
Sources: openrouter.ai
Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic workflows, and is particularly strong at...
Questions this page answers
What does Kimi K3 cost?
$3.00/M in, $15.00/M out in this catalogue, as of the fetch date on this page. That is the API row, not a subscription.
Does Kimi K3 have an independent quality score?
Kimi K3 has an LMArena score in this catalogue.
What beats Kimi K3?
Muse Spark 1.3 has an equal-or-higher measured score and an equal-or-lower price.
Does this page use Artificial Analysis scores for Kimi K3?
No. Artificial Analysis figures are not published here. Quality on this page is LMArena Elo where a score exists; otherwise the row is unrated.
Is the cheapest Kimi K3 endpoint the same product?
Headline cheapest is a lower precision. Like-for-like at the best declared precision is $2.85/M from DeepInfra.
MODEL MONUMENT