---
title: "MoonshotAI: Kimi K3 — price, capability and what beats it · Undominated.ai"
canonical: https://undominated.ai/models/moonshotai__kimi-k3/
description: "MoonshotAI: Kimi K3: $0.680/M in, $13.00/M out. Independent capability scores, and alternatives compared under recorded workload and capability requirements."
---

# MoonshotAI: Kimi K3 — price, capability and what beats it · Undominated.ai

> MoonshotAI: Kimi K3: $0.680/M in, $13.00/M out. Independent capability scores, and alternatives compared under recorded workload and capability requirements.

[Check base rates for this workload](/check/moonshotai__kimi-k3/) [Cheaper alternatives](/alternatives/moonshotai__kimi-k3/) [Family](/families/kimi/) [Compare models](https://undominated.ai/compare/?models=moonshotai%2Fkimi-k3) Comparison context

Choose up to three standard models in Compare.

Requirements and prompt length apply in Compare. This page keeps its stated price basis and workload controls.

MODEL PROOF

# Kimi K3

Moonshot AI · 16 Jul 2026 · kimi-k3

 Balanced Summarise Chat Code gen Agentic

BETTER VALUE OPTION

## A stored alternative has an equal-or-higher measured score and an equal-or-lower price.

 **$3.76** / 1M tokens Balanced · 3 tokens in per 1 out

Serving precision differs between offers.

[Muse Spark 1.3](/models/meta__muse-spark-1.3/) · Δ score 14 points · 47% lower measured price · $2.00 / 1M tokens

[MiMo-V2.6-Pro](/models/xiaomi__mimo-v2.6-pro/) is a further stored option with these losses: 944K → 131K max output

Evidence as of 06 Oct 2026

 [openrouter/models](https://openrouter.ai/api/v1/models) [Vendor cross-check](https://openrouter.ai/api/v1/models) [LMArena · cc-by-4.0](https://huggingface.co/datasets/lmarena-ai/leaderboard-dataset) Independent LMArena score 1,475.9 ± 4.76 · 28,280 votes
 Context window 1.05M
 Maximum output 943.72K
 Input modalities text, image, video
 Output modalities text
 Published input price $0.680 / 1M tokens
 Published output price $13.00 / 1M tokens
 Pricing kind fixed

Inspect [complete billing conditions](#billing-details) and endpoint terms below.

Price from Relace: the cheapest offer that is serving, at standard delivery and at the maker’s precision or better, at its standard rate.

 Every price dimension · USD per million tokens

| Input | $0.680 /M |
| --- | --- |
| Output | $13.00 /M |
| Cached input | $0.450 /M |
| Cache write | Unknown /M |
| Cache write 1h | Unknown /M |
| Reasoning | Unknown /M |

### Who sells it

The same weights, different shops. Cheapest is not like-for-like when serving precision differs.

Headline cheapest is a lower precision. Like-for-like at the best declared precision is $2.50/M from Morph.

 24 provider offers · rates, limits and conditions

Rates are USD per million tokens. Endpoint terms can differ even when the quoted price and precision match. A listed parameter is a provider declaration, not a task-success test.

| Seller | Input | Output | Precision | Uptime (1d) |
| --- | --- | --- | --- | --- |
| Seller Relace relace/fp4 Endpoint terms Cached input /M $0.450 Context limit 1,048,576 tokens Output limit 943,718 tokens Tools Listed by endpoint Reasoning Listed by endpoint Promotional discount None reported Uptime · last 30 minutes 99.60% | Input $0.680 | Output $13.00 | Precision fp4 | Uptime (1d) 99.57% |
| Seller InferenceNet inference-net/fp4 Endpoint terms Cached input /M $0.230 Context limit 1,048,576 tokens Output limit 943,718 tokens Tools Listed by endpoint Reasoning Listed by endpoint Promotional discount None reported Uptime · last 30 minutes 99.58% | Input $0.690 | Output $14.00 | Precision fp4 | Uptime (1d) 99.91% |
| Seller Sail Research sail-research/fp4 Endpoint terms Cached input /M $0.300 Context limit 1,048,576 tokens Output limit 943,718 tokens Tools Listed by endpoint Reasoning Listed by endpoint Promotional discount None reported Uptime · last 30 minutes 99.97% | Input $0.840 | Output $13.50 | Precision fp4 | Uptime (1d) 99.93% |
| Seller Wafer wafer Endpoint terms Cached input /M $0.400 Context limit 1,048,576 tokens Output limit 943,718 tokens Tools Listed by endpoint Reasoning Listed by endpoint Promotional discount None reported Uptime · last 30 minutes 99.36% | Input $1.00 | Output $14.00 | Precision undeclared | Uptime (1d) 99.55% |
| Seller AkashML akashml/fp4 Endpoint terms Cached input /M $1.20 Context limit 1,048,576 tokens Output limit 943,718 tokens Tools Listed by endpoint Reasoning Listed by endpoint Promotional discount None reported Uptime · last 30 minutes Unknown | Input $1.20 | Output $14.00 | Precision fp4 | Uptime (1d) 98.56% |
| Seller Makora makora Endpoint terms Cached input /M $0.204 Context limit 1,048,576 tokens Output limit 943,718 tokens Tools Listed by endpoint Reasoning Listed by endpoint Promotional discount None reported Uptime · last 30 minutes 99.19% | Input $1.53 | Output $12.75 | Precision undeclared | Uptime (1d) 97.54% |
| Seller Phala phala Endpoint terms Cached input /M $0.195 Context limit 1,048,576 tokens Output limit 943,718 tokens Tools Listed by endpoint Reasoning Listed by endpoint Promotional discount 35% · already included in these rates Uptime · last 30 minutes 99.32% | Input $1.95 | Output $9.75 | Precision undeclared | Uptime (1d) 97.77% |
| Seller Decart decart/mxfp4 Endpoint terms Cached input /M $0.201 Context limit 1,048,576 tokens Output limit 943,718 tokens Tools Listed by endpoint Reasoning Listed by endpoint Promotional discount 33% · already included in these rates Uptime · last 30 minutes 93.41% | Input $2.01 | Output $10.05 | Precision mxfp4 | Uptime (1d) 99.17% |
| Seller Morph morph/fp8 Endpoint terms Cached input /M $0.278 Context limit 1,048,576 tokens Output limit 943,718 tokens Tools Listed by endpoint Reasoning Listed by endpoint Promotional discount None reported Uptime · last 30 minutes 100.00% | Input $2.50 | Output $14.00 | Precision fp8 | Uptime (1d) 99.75% |
| Seller DigitalOcean digitalocean Endpoint terms Cached input /M $0.255 Context limit 1,048,576 tokens Output limit 943,718 tokens Tools Listed by endpoint Reasoning Listed by endpoint Promotional discount None reported Uptime · last 30 minutes Unknown | Input $2.55 | Output $12.95 | Precision undeclared | Uptime (1d) 99.88% |
| Seller Together together Endpoint terms Cached input /M $0.270 Context limit 1,048,576 tokens Output limit 943,718 tokens Tools Listed by endpoint Reasoning Listed by endpoint Promotional discount None reported Uptime · last 30 minutes 99.35% | Input $2.70 | Output $13.50 | Precision undeclared | Uptime (1d) 99.36% |
| Seller Wafer wafer/us Endpoint terms Cached input /M $0.300 Context limit 1,048,576 tokens Output limit 943,718 tokens Tools Listed by endpoint Reasoning Listed by endpoint Promotional discount None reported Uptime · last 30 minutes 98.39% | Input $2.80 | Output $14.00 | Precision undeclared | Uptime (1d) 99.16% |
| Seller DeepInfra deepinfra/mxfp4 Endpoint terms Cached input /M $0.285 Context limit 1,048,576 tokens Output limit 943,718 tokens Tools Listed by endpoint Reasoning Listed by endpoint Promotional discount None reported Uptime · last 30 minutes 99.80% | Input $2.85 | Output $14.25 | Precision mxfp4 | Uptime (1d) 99.27% |
| Seller BaseTen baseten/fp8 Endpoint terms Cached input /M $0.300 Context limit 1,048,576 tokens Output limit 262,144 tokens Tools Listed by endpoint Reasoning Listed by endpoint Promotional discount None reported Uptime · last 30 minutes 98.57% | Input $3.00 | Output $15.00 | Precision fp8 | Uptime (1d) 97.93% |
| Seller Parasail parasail/fp4 Endpoint terms Cached input /M $0.300 Context limit 1,048,576 tokens Output limit 943,718 tokens Tools Listed by endpoint Reasoning Listed by endpoint Promotional discount None reported Uptime · last 30 minutes 93.69% | Input $3.00 | Output $15.00 | Precision fp4 | Uptime (1d) 98.53% |
| Seller Amazon Bedrock amazon-bedrock/us-east-2 Endpoint terms Cached input /M $0.300 Context limit 1,048,576 tokens Output limit 128,000 tokens Tools Listed by endpoint Reasoning Listed by endpoint Promotional discount None reported Uptime · last 30 minutes 99.10% | Input $3.00 | Output $15.00 | Precision undeclared | Uptime (1d) 93.42% |
| Seller Chutes chutes/mxfp4 Endpoint terms Cached input /M $0.300 Context limit 1,048,576 tokens Output limit 65,535 tokens Tools Not listed by endpoint Reasoning Listed by endpoint Promotional discount None reported Uptime · last 30 minutes Unknown | Input $3.00 | Output $15.00 | Precision mxfp4 | Uptime (1d) 95.89% |
| Seller Modal modal/mxfp4 Endpoint terms Cached input /M $0.300 Context limit 1,048,576 tokens Output limit 943,718 tokens Tools Listed by endpoint Reasoning Listed by endpoint Promotional discount None reported Uptime · last 30 minutes 99.57% | Input $3.00 | Output $15.00 | Precision mxfp4 | Uptime (1d) 98.78% |
| Seller Fireworks fireworks Endpoint terms Cached input /M $0.300 Context limit 1,048,576 tokens Output limit 943,718 tokens Tools Listed by endpoint Reasoning Listed by endpoint Promotional discount None reported Uptime · last 30 minutes 99.94% | Input $3.00 | Output $15.00 | Precision undeclared | Uptime (1d) 99.11% |
| Seller Moonshot AI moonshotai/mxfp4 Endpoint terms Cached input /M $0.300 Context limit 1,048,576 tokens Output limit 943,718 tokens Tools Listed by endpoint Reasoning Listed by endpoint Promotional discount None reported Uptime · last 30 minutes 100.00% | Input $3.00 | Output $15.00 | Precision mxfp4 | Uptime (1d) 99.99% |
| Seller Alibaba alibaba Endpoint terms Cached input /M $0.345 Context limit 1,048,576 tokens Output limit 131,072 tokens Tools Listed by endpoint Reasoning Listed by endpoint Promotional discount None reported Uptime · last 30 minutes Unknown | Input $3.45 | Output $17.25 | Precision undeclared | Uptime (1d) 99.20% |
| Seller InferenceNet inference-net/fast Endpoint terms Cached input /M $0.450 Context limit 250,000 tokens Output limit 225,000 tokens Tools Listed by endpoint Reasoning Listed by endpoint Promotional discount None reported Uptime · last 30 minutes 99.67% | Input $3.50 | Output $15.00 | Precision fp4 | Uptime (1d) 99.86% |
| Seller Fireworks fireworks/us Endpoint terms Cached input /M $0.450 Context limit 1,048,576 tokens Output limit 943,718 tokens Tools Listed by endpoint Reasoning Listed by endpoint Promotional discount None reported Uptime · last 30 minutes 100.00% | Input $4.50 | Output $22.50 | Precision undeclared | Uptime (1d) 99.67% |
| Seller Fireworks fireworks/fast Endpoint terms Cached input /M $0.450 Context limit 1,048,576 tokens Output limit 943,718 tokens Tools Not listed by endpoint Reasoning Listed by endpoint Promotional discount None reported Uptime · last 30 minutes 99.22% | Input $4.50 | Output $22.50 | Precision undeclared | Uptime (1d) 97.89% |

### Independent scores

| LMArena Elo | 1475.9 max |
| --- | --- |

LMArena Elo under CC BY 4.0. Absence of another board is not a score of zero.

#### Best rankings by task

 - Websites #2 1344
- Code categories #2 1370
- Full-stack apps #2 1313
- Htmlslides #2 1256
- Game development #3 1379
- Data visualisation #3 1356
- UI components #3 1362
- Agenticgamedev #3 1252

Per-task Elo and rank from [Design Arena](https://designarena.ai), via OpenRouter’s model feed. The rank is the feed’s, and it counts models OpenRouter does not list.

### Capability

| Context window | 1M |
| --- | --- |
| Max output | 943K |
| Input modes | text, image, video |
| Tool use | yes |
| Reasoning | optional |
| Open weights | yes |

### Provenance

| Price source | [openrouter.ai](https://openrouter.ai/api/v1/models) |
| --- | --- |
| Fetched | 2026-10-06 |
| Quality data | verified |
| Cross-checked | [vendor page](https://openrouter.ai/api/v1/models) |

## Our take

 editorial — not a measurement

Kimi is a candidate for teams comparing hosted multimodal use with self-hosting. The accepted record lists open weights, but hardware and licence requirements still need review. Use the current proof and benchmark evidence for measured standing. Other models’ withheld latency figures cannot establish a speed comparison; measure responsiveness on your intended deployment.

### Strengths

 - Open weights are recorded
- Text, image and video input are listed

### Weaknesses

 - Measure first-token latency and sustained generation separately on your deployment
- The context window and output ceiling are different limits

### Reach for it when

 - Hosted-versus-self-hosted evaluations
- Multimodal workflows with an explicit latency test

### Avoid it if

 - The decision depends on a latency advantage you have not measured
- You have not checked serving capacity and licence terms

Sources: [openrouter.ai](https://openrouter.ai/api/v1/models)

Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic workflows, and is particularly strong at...

## Questions this page answers

 What does Kimi K3 cost?

$0.680/M in, $13.00/M out in this catalogue, as of the fetch date on this page. That is the API row, not a subscription.

 Does Kimi K3 have an independent quality score?

Kimi K3 has an LMArena score in this catalogue.

 What beats Kimi K3?

Muse Spark 1.3 has an equal-or-higher measured score and an equal-or-lower price.

 Does this page use Artificial Analysis scores for Kimi K3?

No. Artificial Analysis figures are not published here. Quality on this page is LMArena Elo where a score exists; otherwise the row is unrated.

 Is the cheapest Kimi K3 endpoint the same product?

Headline cheapest is a lower precision. Like-for-like at the best declared precision is $2.50/M from Morph.

MODEL MONUMENT

## Save or share this proof

 Download proof card · 1200 × 630 Preview and more sharing options

MoonshotAI: Kimi K3. Balanced workload. Better value option available. Evidence as of 06 Oct 2026.

 Copy model link Download square card · 1080 × 1080 Copy landscape image Share landscape card

## Related

 - Both better and cheaper [Muse Spark 1.3](/models/meta__muse-spark-1.3/)
- Higher score and cheaper, with a named trade [MiMo-V2.6-Pro](/models/xiaomi__mimo-v2.6-pro/)
- Side by side [DeepSeek V4.1 Flash](/compare/deepseek__deepseek-v4.1-flash--vs--moonshotai__kimi-k3/)
- Nearby on the capability axis [Gemini 3.5 Flash](/models/google__gemini-3.5-flash/)
- Same vendor [Kimi K2.6](/models/moonshotai__kimi-k2.6/)
- Moonshot AI models [Moonshot AI](/providers/moonshot-ai/)
- Where this model appears [wire](/wire/)
- Side by side [Kimi K2.6](/compare/moonshotai__kimi-k2.6--vs--moonshotai__kimi-k3/)
- Nearby on the capability axis [GLM 5.3](/models/z-ai__glm-5.3/)
- Same vendor [Kimi K2.5](/models/moonshotai__kimi-k2.5/)
- Where this model appears [self-host](/self-host/)
- Nearby on the capability axis [Muse Spark 1.1](/models/meta__muse-spark-1.1/)

## Badge

The verdict as one SVG, for a README or a docs page. It states this model's dominance status at the balanced workload on the LMArena lens, with the date it was computed, and it is rebuilt with the catalogue — so it changes when the verdict changes, including to one you would rather it did not.

 Markdown Copy
 [![Kimi K3 — Undominated.ai dominance verdict](https://undominated.ai/badge/moonshotai__kimi-k3.svg)](https://undominated.ai/models/moonshotai__kimi-k3/)
 HTML Copy
 <a href="https://undominated.ai/models/moonshotai__kimi-k3/"><img src="https://undominated.ai/badge/moonshotai__kimi-k3.svg" alt="Kimi K3 — Undominated.ai dominance verdict" height="36"></a>

Direct file: [https://undominated.ai/badge/moonshotai__kimi-k3.svg](https://undominated.ai/badge/moonshotai__kimi-k3.svg) — a static SVG, written by scripts/build-badges.mjs on every build, so the copy you embed is never older than the last deploy.

## Continue your investigation

 - [Explore model families](/families/)
- [Compare a shortlist](/compare/)
- [Check a model](/check/)
- [Understand the evidence](/methodology/)
