La frontiera del valore

Ogni modello, tracciato in base a quanto costa rispetto a quanto è buono. La scala a gradini è la frontiera: nessun modello batte quelli su entrambi gli assi contemporaneamente. Tutto ciò che sta sotto e a destra è un affare nettamente peggiore — esiste qualcosa insieme più intelligente e più economico.

Changing the workload below moves the frontier less than you would expect: 91% of frontier models are on it under every workload (only Solar Pro 4 moves). What does change is the price — the same model swings up to across these shapes. So the control is not about which models to consider. It is about what you would actually pay for them.

10 non dominati su 106 96 affari peggiori
Carico di lavoro
Misurato su
$0.030$0.100$0.300$1.00$3.00$10.00102030405060$ effettivi per milione di token — scala logaritmicacapacità Generale
  1. 1 Ling-3.0-flashinclusionAI 37.8 $0.032/M gradino più economico
  2. 2 Solar Pro 4Upstage 41.6 $0.052/M +3.8 per 1.7× il prezzo
  3. 3 DeepSeek V4 Flash 0423DeepSeek 42.1 $0.061/M +0.5 per 1.2× il prezzo
  4. 4 DeepSeek V4 Flash 0731DeepSeek 51.8 $0.105/M +9.7 per 1.7× il prezzo
  5. 5 GPT-5.6 LunaOpenAI 52.3 $0.450/M +0.5 per 4.3× il prezzo
  6. 6 Gemini 3.7 FlashGoogle 56.0 $0.750/M +3.7 per 1.7× il prezzo
  7. 7 Muse Spark 1.2Meta 56.8 $2.00/M +0.8 per 2.7× il prezzo
  8. 8 GLM 5.3Z.ai 59.5 $2.15/M +2.7 per 1.1× il prezzo
  9. 9 Grok 4.6xAI 60.9 $3.00/M +1.4 per 1.4× il prezzo
  10. 10 Claude Opus 5Anthropic 63.1 $10.00/M +2.2 per 3.3× il prezzo
Markdown for LLMs