Grok 4.20

xAI released 2026-03-31 proprietary

$1.56 per million tokens, balanced

No independent quality score.

No benchmark we track has measured this model. That is not the same as measuring it and finding it wanting — we simply cannot rank it, so we do not.

Our take

editorial — not a measurement

The context-window outlier: 2M tokens at $1.25/$2.50, which is the cheapest long-context option from any major proprietary vendor. Same 200k tier threshold as the rest of the Grok line, doubling to $2.50/$5 above it — so the 2M window is only cheap in headline terms, and a genuinely 2M-token prompt bills entirely at the upper tier. Still, for bulk long-document processing the arithmetic is hard to beat. A multi-agent variant exists at the same rate.

Strengths

  • 2M token context window — largest here
  • $1.25/$2.50 short-context, very cheap
  • Multi-agent variant at identical pricing
  • Published tier threshold

Weaknesses

  • Everything past 200k bills at $2.50/$5, so the big window is not cheap in practice
  • Below the 4.5/4.6 line on capability
  • No batch discount published

Reach for it when

  • Bulk long-document ingestion
  • Cheap very-long-context retrieval

Avoid it if

  • You need frontier reasoning quality
  • You assumed the headline rate applies at 2M tokens

Sources: docs.x.aiarena.ai

Every price dimension

Input$1.25/M
Output$2.50/M
Cached input$0.200/M
Web search$0.005/call

Past 200,000 tokens the price changes. Input goes to $2.50/M (2×) and output to $5.00/M. The headline rate does not apply to a long-context workload.

Independent scores

LMArena Elo1444.4beta1

Sources are listed separately rather than averaged. Across the 62 models both have scored they correlate at r = 0.806 — close agreement overall, but four models rank very differently between them, and a blended score would hide exactly those.

Best rankings by task

  • htmlslides #12of 34 1180
  • asciiart #20of 83 1205
  • webapps #23of 56 1181
  • godotgamedev #26of 43 1089
  • svg #30of 98 1203
  • fullstack #30of 59 1086

Capability

Context window2M
Input modestext, image, file
Tool useyes
Reasoningoptional
Knowledge cutoff2025-09-01
Open weightsno

Provenance

Price sourceopenrouter.ai
Fetched2026-08-24
Quality dataunrated
Cross-checkedvendor page

Grok 4.20 is a reasoning model from SpaceXAI with industry-leading speed and agentic tool calling capabilities. It combines the lowest hallucination rate on the market with strict prompt adherance, delivering...

Markdown for LLMs