MODEL PROOF

Kimi K3

BETTER VALUE OPTION

A stored alternative has an equal-or-higher measured score and an equal-or-lower price.

$6.00 / 1M tokens Balanced · 3 tokens in per 1 out

Serving precision differs between offers.

Muse Spark 1.3 · Δ score 17.4 points · 67% lower measured price · $2.00 / 1M tokens

Gemini 3.8 Flash is a further stored option with these losses: 944K → 66K max output

Independent LMArena score
1,472.3 ± 5.22 · 20,987 votes
Context window
1.05M
Maximum output
943.72K
Input modalities
text, image, video
Output modalities
text
Published input price
$3.00 / 1M tokens
Published output price
$15.00 / 1M tokens
Pricing kind
fixed

Inspect complete billing conditions and endpoint terms below.

Every price dimension · USD per million tokens
Input$3.00/M
Output$15.00/M
Cached input$0.300/M
Cache writeUnknown/M
Cache write 1hUnknown/M
ReasoningUnknown/M

Who sells it

The same weights, different shops. Cheapest is not like-for-like when serving precision differs.

Headline cheapest is a lower precision. Like-for-like at the best declared precision is $2.85/M from DeepInfra.

20 provider offers · rates, limits and conditions

Rates are USD per million tokens. Endpoint terms can differ even when the quoted price and precision match. A listed parameter is a provider declaration, not a task-success test.

SellerInputOutputPrecisionUptime (1d)
Sail Researchsail-research/fp4
Endpoint terms
Cached input /M
$0.300
Context limit
1,048,576 tokens
Output limit
943,718 tokens
Tools
Listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
99.86%
$1.50 $10.76 fp4 99.84%
InferenceNetinference-net/fp4
Endpoint terms
Cached input /M
$0.195
Context limit
1,048,576 tokens
Output limit
943,718 tokens
Tools
Listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
35% · already included in these rates
Uptime · last 30 minutes
98.85%
$1.95 $9.75 fp4 98.59%
Relacerelace/fp4
Endpoint terms
Cached input /M
$0.195
Context limit
1,048,576 tokens
Output limit
943,718 tokens
Tools
Listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
99.55%
$1.95 $9.75 fp4 98.71%
Phalaphala
Endpoint terms
Cached input /M
$0.210
Context limit
1,048,576 tokens
Output limit
943,718 tokens
Tools
Listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
30% · already included in these rates
Uptime · last 30 minutes
99.35%
$2.10 $10.50 undeclared 99.47%
Waferwafer
Endpoint terms
Cached input /M
$0.250
Context limit
1,048,576 tokens
Output limit
943,718 tokens
Tools
Listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
99.84%
$2.49 $14.50 undeclared 99.92%
Morphmorph
Endpoint terms
Cached input /M
$0.290
Context limit
1,048,576 tokens
Output limit
943,718 tokens
Tools
Listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
Unknown
$2.50 $14.00 undeclared 100.00%
Makoramakora
Endpoint terms
Cached input /M
$0.256
Context limit
1,048,576 tokens
Output limit
943,718 tokens
Tools
Listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
99.74%
$2.55 $12.75 undeclared 98.21%
DigitalOceandigitalocean
Endpoint terms
Cached input /M
$0.255
Context limit
1,048,576 tokens
Output limit
943,718 tokens
Tools
Listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
99.90%
$2.55 $12.95 undeclared 99.64%
DeepInfradeepinfra/bf16
Endpoint terms
Cached input /M
$0.285
Context limit
1,048,576 tokens
Output limit
16,384 tokens
Tools
Listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
99.77%
$2.85 $14.25 bf16 94.48%
BaseTenbaseten/fp8
Endpoint terms
Cached input /M
$0.300
Context limit
1,048,576 tokens
Output limit
262,144 tokens
Tools
Listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
99.62%
$3.00 $15.00 fp8 99.14%
Parasailparasail/fp4
Endpoint terms
Cached input /M
$0.300
Context limit
1,048,576 tokens
Output limit
943,718 tokens
Tools
Listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
96.74%
$3.00 $15.00 fp4 98.24%
Chuteschutes/mxfp4
Endpoint terms
Cached input /M
$0.300
Context limit
1,048,576 tokens
Output limit
65,535 tokens
Tools
Not listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
Unknown
$3.00 $15.00 mxfp4 99.07%
Modalmodal/mxfp4
Endpoint terms
Cached input /M
$0.300
Context limit
1,048,576 tokens
Output limit
943,718 tokens
Tools
Listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
99.34%
$3.00 $15.00 mxfp4 99.59%
Togethertogether
Endpoint terms
Cached input /M
$0.300
Context limit
1,048,576 tokens
Output limit
943,718 tokens
Tools
Listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
99.97%
$3.00 $15.00 undeclared 99.54%
Fireworksfireworks
Endpoint terms
Cached input /M
$0.300
Context limit
1,048,576 tokens
Output limit
943,718 tokens
Tools
Listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
100.00%
$3.00 $15.00 undeclared 99.84%
Moonshot AImoonshotai/mxfp4
Endpoint terms
Cached input /M
$0.300
Context limit
1,048,576 tokens
Output limit
943,718 tokens
Tools
Listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
99.68%
$3.00 $15.00 mxfp4 99.83%
Fireworksfireworks/us
Endpoint terms
Cached input /M
$0.330
Context limit
1,048,576 tokens
Output limit
943,718 tokens
Tools
Listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
99.79%
$3.30 $16.50 undeclared 99.05%
Alibabaalibaba
Endpoint terms
Cached input /M
$0.345
Context limit
1,048,576 tokens
Output limit
943,718 tokens
Tools
Listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
99.57%
$3.45 $17.25 undeclared 98.86%
Fireworksfireworks/fast
Endpoint terms
Cached input /M
$0.450
Context limit
1,048,576 tokens
Output limit
943,718 tokens
Tools
Not listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
100.00%
$4.50 $22.50 undeclared 99.70%
Morphmorph/fast
Endpoint terms
Cached input /M
$0.600
Context limit
1,048,576 tokens
Output limit
943,718 tokens
Tools
Listed by endpoint
Reasoning
Listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
Unknown
$6.00 $22.50 undeclared 100.00%

Independent scores

LMArena Elo1472.3max

LMArena Elo under CC BY 4.0. Absence of another board is not a score of zero.

Best rankings by task

  • codecategories #1of 108 1379
  • uicomponent #1of 104 1370
  • website #2of 114 1349
  • gamedev #2of 107 1397
  • fullstack #2of 41 1325
  • htmlslides #2of 23 1259
  • 3d #3of 100 1422
  • mobileapps #3of 40 1282

Capability

Context window1M
Max output944K
Input modestext, image, video
Tool useyes
Reasoningoptional
Open weightsyes

Provenance

Price sourceopenrouter.ai
Fetched2026-09-23
Quality dataverified
Cross-checkedvendor page

Our take

editorial — not a measurement

Kimi is a candidate for teams comparing hosted multimodal use with self-hosting. The accepted record lists open weights, but hardware and licence requirements still need review. Use the current proof and benchmark evidence for measured standing. Other models’ withheld latency figures cannot establish a speed comparison; measure responsiveness on your intended deployment.

Strengths

  • Open weights are recorded
  • Text, image and video input are listed

Weaknesses

  • Measure first-token latency and sustained generation separately on your deployment
  • The context window and output ceiling are different limits

Reach for it when

  • Hosted-versus-self-hosted evaluations
  • Multimodal workflows with an explicit latency test

Avoid it if

  • The decision depends on a latency advantage you have not measured
  • You have not checked serving capacity and licence terms

Sources: openrouter.ai

Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic workflows, and is particularly strong at...

Questions this page answers

What does Kimi K3 cost?

$3.00/M in, $15.00/M out in this catalogue, as of the fetch date on this page. That is the API row, not a subscription.

Does Kimi K3 have an independent quality score?

Kimi K3 has an LMArena score in this catalogue.

What beats Kimi K3?

Muse Spark 1.3 has an equal-or-higher measured score and an equal-or-lower price.

Does this page use Artificial Analysis scores for Kimi K3?

No. Artificial Analysis figures are not published here. Quality on this page is LMArena Elo where a score exists; otherwise the row is unrated.

Is the cheapest Kimi K3 endpoint the same product?

Headline cheapest is a lower precision. Like-for-like at the best declared precision is $2.85/M from DeepInfra.

MODEL MONUMENT

Save or share this proof

Badge

The verdict as one SVG, for a README or a docs page. It states this model's dominance status at the balanced workload on the LMArena lens, with the date it was computed, and it is rebuilt with the catalogue — so it changes when the verdict changes, including to one you would rather it did not.

Kimi K3 — Undominated.ai dominance verdict

Markdown
[![Kimi K3 — Undominated.ai dominance verdict](https://undominated.ai/badge/moonshotai__kimi-k3.svg)](https://undominated.ai/models/moonshotai__kimi-k3/)
HTML
<a href="https://undominated.ai/models/moonshotai__kimi-k3/"><img src="https://undominated.ai/badge/moonshotai__kimi-k3.svg" alt="Kimi K3 — Undominated.ai dominance verdict" height="36"></a>

Direct file: https://undominated.ai/badge/moonshotai__kimi-k3.svg — a static SVG, written by scripts/build-badges.mjs on every build, so the copy you embed is never older than the last deploy.

Evidence & Ask