The value frontier

Every model, plotted by what it costs against how good it is. The staircase is the frontier: nothing beats those models on both axes at once. Everything below and to the right is a strictly worse deal — there is something both smarter and cheaper.

Changing the workload below moves the frontier less than you would expect: 91% of frontier models are on it under every workload (only Solar Pro 4 moves). What does change is the price — the same model swings up to across these shapes. So the control is not about which models to consider. It is about what you would actually pay for them.

10 undominated of 106 96 worse deals
Workload
Measured on
$0.030$0.100$0.300$1.00$3.00$10.00102030405060effective $ per million tokens — log scaleGeneral capability
  1. 1 Ling-3.0-flashinclusionAI 37.8 $0.032/M cheapest rung
  2. 2 Solar Pro 4Upstage 41.6 $0.052/M +3.8 for 1.7× the price
  3. 3 DeepSeek V4 Flash 0423DeepSeek 42.1 $0.061/M +0.5 for 1.2× the price
  4. 4 DeepSeek V4 Flash 0731DeepSeek 51.8 $0.105/M +9.7 for 1.7× the price
  5. 5 GPT-5.6 LunaOpenAI 52.3 $0.450/M +0.5 for 4.3× the price
  6. 6 Gemini 3.7 FlashGoogle 56.0 $0.750/M +3.7 for 1.7× the price
  7. 7 Muse Spark 1.2Meta 56.8 $2.00/M +0.8 for 2.7× the price
  8. 8 GLM 5.3Z.ai 59.5 $2.15/M +2.7 for 1.1× the price
  9. 9 Grok 4.6xAI 60.9 $3.00/M +1.4 for 1.4× the price
  10. 10 Claude Opus 5Anthropic 63.1 $10.00/M +2.2 for 3.3× the price
Markdown for LLMs