Qwen3.8 2.4T A95B
Qwen released 2026-08-12 apache-2.0
GLM 5.3 is both better and cheaper.
It scores +1.8 higher and costs 28% less ($2.15/M against $3.00/M) on this workload — and it does everything this model does.
Grok 4.6 is cheaper still (0% less) but drops 1.0M → 500K context.
Every price dimension
| Input | $2.00/M |
|---|---|
| Output | $6.00/M |
| Cached input | $0.250/M |
Independent scores
| Intelligence | 57.7 |
|---|---|
| Coding | 71.9 |
| Agentic | 57.1 |
Sources are listed separately rather than averaged. Across the 62 models both have scored they correlate at r = 0.806 — close agreement overall, but four models rank very differently between them, and a blended score would hide exactly those.
Capability
| Context window | 1.0M |
|---|---|
| Max output | 131K |
| Input modes | text |
| Tool use | yes |
| Reasoning | always on |
| Open weights | yes |
Provenance
| Price source | openrouter.ai |
|---|---|
| Fetched | 2026-08-24 |
| Quality data | verified |
| Cross-checked | vendor page |
Qwen3.8 2.4T A95B is an open-weight sparse mixture-of-experts model from Qwen and the open-weight variant of [Qwen3.8 Max](/qwen/qwen3.8-max), with 95 billion active parameters out of 2.4 trillion total. It is...