Grok 4.6
xAI released 2026-08-12 proprietary
Nothing is both better and cheaper.
Under a balanced workload, no other model in the catalogue scores higher and costs less. This model is on the value frontier.
Our take
editorial — not a measurementGenuinely competitive at the frontier and the cheapest way to reach AA index 61 — it matches GPT-5.6 Sol's intelligence score at $2/$6 versus Sol's $4/$20, with a blended cost Artificial Analysis puts at $0.84 per million against Sol's $1.23. Output at $6 is the standout number; most frontier models charge 3-5x that. xAI publishes its long-context threshold clearly at 200k, above which it becomes $4/$12. The 500k context is smaller than the 1M-token peers. OpenRouter now labels the vendor 'SpaceXAI'.
- Unverified in this take: This note cites a time-to-first-token figure that the data layer withheld as implausible (it appears to be total reasoning time mislabelled at source). Treat it as unverified.
Strengths
- AA intelligence index 61 at a fraction of Sol's price
- $6 output per MTok — exceptionally cheap at this capability
- Published 200k tier threshold
- 67 output tokens/sec, ~44s TTFT — faster than Sol and Opus 5
Weaknesses
- 500k context, half the 1M-token peers
- Doubles to $4/$12 past 200k tokens
- No batch discount published
- Consumer-side documentation is poor; API docs are the only reliable source
Reach for it when
- Cost-sensitive frontier reasoning
- Output-heavy generation where $6/MTok matters
- Workloads that fit under 200k tokens
Avoid it if
- You need a 1M-token window
- You need a published batch discount
Sources: docs.x.aiartificialanalysis.ai
Every price dimension
| Input | $2.00/M |
|---|---|
| Output | $6.00/M |
| Cached input | $0.500/M |
| Web search | $0.005/call |
Past 200,000 tokens the price changes. Input goes to $4.00/M (2×) and output to $12.00/M. The headline rate does not apply to a long-context workload.
Independent scores
| Intelligence | 60.9 |
|---|---|
| Coding | 76.8 |
| Agentic | 58.7 |
| LMArena Elo | 1443.7high |
Sources are listed separately rather than averaged. Across the 62 models both have scored they correlate at r = 0.806 — close agreement overall, but four models rank very differently between them, and a blended score would hide exactly those.
Best rankings by task
- androidnative #1of 54 1311
- mobileapps #2of 56 1281
- asciiart #4of 83 1312
- webapps #5of 56 1282
- gamedev #6of 144 1340
- uicomponent #6of 140 1341
- svg #6of 98 1280
- fullstack #6of 59 1293
Capability
| Context window | 500K |
|---|---|
| Input modes | text, image, file |
| Tool use | yes |
| Reasoning | always on |
| Open weights | no |
Provenance
| Price source | openrouter.ai |
|---|---|
| Fetched | 2026-08-24 |
| Quality data | verified |
| Cross-checked | vendor page |
A published time-to-first-token figure for this model was implausible (44s) and has been withheld rather than displayed.
Grok 4.6 is SpaceXAI's smartest model with frontier performance on coding, knowledge work, and STEM.