Llama 4 Maverick
Meta released 2025-04-05 llama-4-community
MiMo-V2.5 is both better and cheaper.
It scores +23.5 higher and costs 50% less ($0.175/M against $0.350/M) on this workload — and it does everything this model does.
DeepSeek V4 Flash 0731 is cheaper still (70% less) but drops no image input.
Our take
editorial — not a measurementNow the legacy option in Meta's lineup, superseded by the Muse line, but still the one with weights you can actually download and a licence you can read. $0.20/$0.80 on OpenRouter's default endpoint with a 1M window. The Llama community licence is not OSI-open — it carries field-of-use and scale restrictions — so check it against your deployment before assuming Apache-style freedom. For genuinely unrestricted weights, gpt-oss-120b or the Apache-licensed Qwen and Granite models are cleaner choices.
Strengths
- Genuinely downloadable weights with a published licence
- $0.20/$0.80 at the default hosted endpoint
- 1M context window, vision-capable
- Broad third-party hosting — real price competition
Weaknesses
- Llama community licence has field-of-use and scale restrictions
- Superseded by the Muse line on capability
- 16k max output is restrictive
- No caching support at the listed endpoint
Reach for it when
- Self-hosted deployment within licence terms
- Vision tasks at low cost
- Fine-tuning bases
Avoid it if
- The Llama licence restrictions bind you
- You need long outputs
Sources: openrouter.ai
Every price dimension
| Input | $0.200/M |
|---|---|
| Output | $0.800/M |
Independent scores
| Intelligence | 14.5 |
|---|---|
| Coding | 16.3 |
| Agentic | 1.2 |
Sources are listed separately rather than averaged. Across the 62 models both have scored they correlate at r = 0.806 — close agreement overall, but four models rank very differently between them, and a blended score would hide exactly those.
Capability
| Context window | 1.0M |
|---|---|
| Max output | 16K |
| Input modes | text, image |
| Tool use | yes |
| Reasoning | no |
| Knowledge cutoff | 2024-08-31 |
| Open weights | yes |
Provenance
| Price source | openrouter.ai |
|---|---|
| Fetched | 2026-08-24 |
| Quality data | verified |
| Cross-checked | vendor page |
Llama 4 Maverick 17B Instruct (128E) is a high-capacity multimodal language model from Meta, built on a mixture-of-experts (MoE) architecture with 128 experts and 17 billion active parameters per forward...