Kimi K3

Moonshot AI released 2026-07-16 modified-mit

$6.00 per million tokens, balanced

Grok 4.6 scores higher and costs less — but you would give something up.

+1.2 on the capability index and 50% cheaper ($3.00/M against $6.00/M). What you lose:

  • 1.0M → 500K context
  • no video input

Capability scores do not measure context length, output ceiling or which inputs a model accepts, so a higher score does not mean a drop-in replacement.

Our take

editorial — not a measurement

The strongest open-weight model in this dataset on measured intelligence — AA index 60, matching GLM 5.3 and within three points of Opus 5, with an LMArena score of 1489 at max effort. The number that stands out is latency: 3.3 seconds to first token, more than 10x faster than Opus 5 (~41s) and 30x faster than GPT-5.6 Sol (~105s). If you need near-frontier quality in an interactive product, nothing else in the sampled set is close on responsiveness. $3/$15 is not cheap for open weights, but you can self-host and escape that entirely. 1M context, 1M max output.

Strengths

  • AA intelligence 60 and LMArena 1489 — frontier-adjacent
  • 3.3s time-to-first-token, by far the fastest in the sampled set
  • 1M context and 1M max output tokens
  • Open weights — self-hostable

Weaknesses

  • $3/$15 hosted is expensive relative to other open-weight options
  • 35 output tokens/sec is slow once generation starts
  • Modified-MIT licence not independently verified

Reach for it when

  • Interactive products needing near-frontier quality
  • Self-hosted frontier-adjacent deployment
  • Latency-critical agentic loops

Avoid it if

  • Sustained generation throughput matters more than first-token latency
  • You want the cheapest hosted tokens

Sources: openrouter.aiartificialanalysis.aiarena.ai

Every price dimension

Input$3.00/M
Output$15.00/M
Cached input$0.300/M

Independent scores

Intelligence59.7
Coding76.2
Agentic54.3
LMArena Elo1476.3max

Sources are listed separately rather than averaged. Across the 62 models both have scored they correlate at r = 0.806 — close agreement overall, but four models rank very differently between them, and a blended score would hide exactly those.

Best rankings by task

  • website #1of 151 1372
  • codecategories #1of 145 1407
  • gamedev #1of 144 1431
  • dataviz #1of 142 1376
  • uicomponent #1of 140 1377
  • 3d #1of 136 1445
  • mobileapps #1of 56 1290
  • svg #2of 98 1356

Capability

Context window1.0M
Max output1.0M
Input modestext, image, video
Tool useyes
Reasoningoptional
Open weightsyes

Provenance

Price sourceopenrouter.ai
Fetched2026-08-24
Quality dataverified
Cross-checkedvendor page

Kimi K3 is a 2.8T parameter open-weight multimodal reasoning model from Moonshot AI. It is suited for complex coding, knowledge work, and long-horizon agentic workflows, and is particularly strong at...

Markdown for LLMs