Llama 4 Maverick

Meta released 2025-04-05 llama-4-community

$0.350 per million tokens, balanced

MiMo-V2.5 is both better and cheaper.

It scores +23.5 higher and costs 50% less ($0.175/M against $0.350/M) on this workload — and it does everything this model does.

DeepSeek V4 Flash 0731 is cheaper still (70% less) but drops no image input.

Our take

editorial — not a measurement

Now the legacy option in Meta's lineup, superseded by the Muse line, but still the one with weights you can actually download and a licence you can read. $0.20/$0.80 on OpenRouter's default endpoint with a 1M window. The Llama community licence is not OSI-open — it carries field-of-use and scale restrictions — so check it against your deployment before assuming Apache-style freedom. For genuinely unrestricted weights, gpt-oss-120b or the Apache-licensed Qwen and Granite models are cleaner choices.

Strengths

  • Genuinely downloadable weights with a published licence
  • $0.20/$0.80 at the default hosted endpoint
  • 1M context window, vision-capable
  • Broad third-party hosting — real price competition

Weaknesses

  • Llama community licence has field-of-use and scale restrictions
  • Superseded by the Muse line on capability
  • 16k max output is restrictive
  • No caching support at the listed endpoint

Reach for it when

  • Self-hosted deployment within licence terms
  • Vision tasks at low cost
  • Fine-tuning bases

Avoid it if

  • The Llama licence restrictions bind you
  • You need long outputs

Sources: openrouter.ai

Every price dimension

Input$0.200/M
Output$0.800/M

Independent scores

Intelligence14.5
Coding16.3
Agentic1.2

Sources are listed separately rather than averaged. Across the 62 models both have scored they correlate at r = 0.806 — close agreement overall, but four models rank very differently between them, and a blended score would hide exactly those.

Capability

Context window1.0M
Max output16K
Input modestext, image
Tool useyes
Reasoningno
Knowledge cutoff2024-08-31
Open weightsyes

Provenance

Price sourceopenrouter.ai
Fetched2026-08-24
Quality dataverified
Cross-checkedvendor page

Llama 4 Maverick 17B Instruct (128E) is a high-capacity multimodal language model from Meta, built on a mixture-of-experts (MoE) architecture with 128 experts and 17 billion active parameters per forward...

Markdown for LLMs