バリューフロンティア

すべてのモデルを、価格と性能の関係でプロットしています。階段状の線がフロンティアです。両方の軸で同時にそれらのモデルを上回るものはありません。その下および右側にあるものは、明確に割の悪い選択肢です。より賢く、しかも安いものが存在します。

Changing the workload below moves the frontier less than you would expect: 91% of frontier models are on it under every workload (only Solar Pro 4 moves). What does change is the price — the same model swings up to across these shapes. So the control is not about which models to consider. It is about what you would actually pay for them.

10 106 件中の非劣位 割の悪い選択肢 96 件
ワークロード
測定基準
$0.030$0.100$0.300$1.00$3.00$10.00102030405060100万トークンあたりの実効 $ — 対数スケール総合 の能力
  1. 1 Ling-3.0-flashinclusionAI 37.8 $0.032/M 最安の段
  2. 2 Solar Pro 4Upstage 41.6 $0.052/M 価格 1.7× で +3.8
  3. 3 DeepSeek V4 Flash 0423DeepSeek 42.1 $0.061/M 価格 1.2× で +0.5
  4. 4 DeepSeek V4 Flash 0731DeepSeek 51.8 $0.105/M 価格 1.7× で +9.7
  5. 5 GPT-5.6 LunaOpenAI 52.3 $0.450/M 価格 4.3× で +0.5
  6. 6 Gemini 3.7 FlashGoogle 56.0 $0.750/M 価格 1.7× で +3.7
  7. 7 Muse Spark 1.2Meta 56.8 $2.00/M 価格 2.7× で +0.8
  8. 8 GLM 5.3Z.ai 59.5 $2.15/M 価格 1.1× で +2.7
  9. 9 Grok 4.6xAI 60.9 $3.00/M 価格 1.4× で +1.4
  10. 10 Claude Opus 5Anthropic 63.1 $10.00/M 価格 3.3× で +2.2
Markdown for LLMs