Grok 4.6

xAI released 2026-08-12 proprietary

$3.00 per million tokens, balanced

Nothing is both better and cheaper.

Under a balanced workload, no other model in the catalogue scores higher and costs less. This model is on the value frontier.

Our take

editorial — not a measurement

Genuinely competitive at the frontier and the cheapest way to reach AA index 61 — it matches GPT-5.6 Sol's intelligence score at $2/$6 versus Sol's $4/$20, with a blended cost Artificial Analysis puts at $0.84 per million against Sol's $1.23. Output at $6 is the standout number; most frontier models charge 3-5x that. xAI publishes its long-context threshold clearly at 200k, above which it becomes $4/$12. The 500k context is smaller than the 1M-token peers. OpenRouter now labels the vendor 'SpaceXAI'.

  • Unverified in this take: This note cites a time-to-first-token figure that the data layer withheld as implausible (it appears to be total reasoning time mislabelled at source). Treat it as unverified.

Strengths

  • AA intelligence index 61 at a fraction of Sol's price
  • $6 output per MTok — exceptionally cheap at this capability
  • Published 200k tier threshold
  • 67 output tokens/sec, ~44s TTFT — faster than Sol and Opus 5

Weaknesses

  • 500k context, half the 1M-token peers
  • Doubles to $4/$12 past 200k tokens
  • No batch discount published
  • Consumer-side documentation is poor; API docs are the only reliable source

Reach for it when

  • Cost-sensitive frontier reasoning
  • Output-heavy generation where $6/MTok matters
  • Workloads that fit under 200k tokens

Avoid it if

  • You need a 1M-token window
  • You need a published batch discount

Sources: docs.x.aiartificialanalysis.ai

Every price dimension

Input$2.00/M
Output$6.00/M
Cached input$0.500/M
Web search$0.005/call

Past 200,000 tokens the price changes. Input goes to $4.00/M (2×) and output to $12.00/M. The headline rate does not apply to a long-context workload.

Independent scores

Intelligence60.9
Coding76.8
Agentic58.7
LMArena Elo1443.7high

Sources are listed separately rather than averaged. Across the 62 models both have scored they correlate at r = 0.806 — close agreement overall, but four models rank very differently between them, and a blended score would hide exactly those.

Best rankings by task

  • androidnative #1of 54 1311
  • mobileapps #2of 56 1281
  • asciiart #4of 83 1312
  • webapps #5of 56 1282
  • gamedev #6of 144 1340
  • uicomponent #6of 140 1341
  • svg #6of 98 1280
  • fullstack #6of 59 1293

Capability

Context window500K
Input modestext, image, file
Tool useyes
Reasoningalways on
Open weightsno

Provenance

Price sourceopenrouter.ai
Fetched2026-08-24
Quality dataverified
Cross-checkedvendor page

A published time-to-first-token figure for this model was implausible (44s) and has been withheld rather than displayed.

Grok 4.6 is SpaceXAI's smartest model with frontier performance on coding, knowledge work, and STEM.

Markdown for LLMs