Is R1 a good deal?

Whether anything in this catalogue beats R1 on both quality and price, and what you give up if it does. A computation on the current catalogue, not an opinion.

13 undominated of 136 · Sep 23, 2026

As of Sep 23, 2026, R1 is dominated for Balanced on LMArena. GLM 5.3 Flash scores 99.3 higher and costs 79% less, with a covering envelope. 13 of 136 rated, priced standard models are undominated.

Inspect model evidence Compare differences & requirements

Current model

R1

Balanced

3 tokens in per 1 out

GLM 5.3 Flash is both better and cheaper than R1: 99.3 points higher on LMArena and 79% less per million tokens, $0.91 cheaper at this mix.

R1

LMArena Elo · higher is better

Scale starts at 1040 Elo

1372.6

Effective $/M · Balanced · lower is better

$1.15/M

GLM 5.3 Flash

LMArena Elo · higher is better

Scale starts at 1040 Elo

1471.9

Effective $/M · Balanced · lower is better

$0.237/M

R1 takes text, returns up to 16,000 tokens from a 64,000-token context, and is listed by 1 seller.

Compared against 136 rated, priced models on this lens: 39 models dominate it and give up nothing, 5 more dominate it but give something up. R1 scores 1372.6 at $1.15 per million tokens for this mix.

Envelope-safe replacements

Each row scores at least as high, costs no more, and covers this model’s context, output, modalities, tools, and reasoning. A cheaper narrower model is not listed here.

ModelLMArenaEffective $/MYou save
GLM 5.3 Flash1471.9 +99.3$0.237/M79%
GLM 5.21466.9 +94.3$0.998/M13%
MiMo-V2.5-Pro1464.8 +92.2$0.544/M53%
Qwen3.7 Plus1454.2 +81.6$0.560/M51%
GLM 51446.3 +73.7$0.930/M19%
Kimi K2.51445.6 +73.0$0.900/M22%
Gemma 4 31B1441.7 +69.1$0.153/M87%
Hy31440.6 +68.0$0.144/M87%
GLM 4.61440.5 +67.9$0.760/M34%
Qwen3.8 27B1439.3 +66.7$1.06/M7%
Qwen3.6 Plus1436.7 +64.1$0.731/M36%
GLM 4.71435.9 +63.3$0.738/M36%
Gemini 3.5 Flash Lite1435.5 +62.9$0.850/M26%
Gemma 4 26B A4B 1434.5 +61.9$0.143/M88%
MiniMax M31433.5 +60.9$0.525/M54%
DeepSeek V4 Flash 04231431.8 +59.2$0.103/M91%
GLM 4.51430.2 +57.6$1.00/M13%
GPT-5.6 Luna1429.9 +57.3$0.450/M61%
R1 05281427.5 +54.9$0.912/M21%
MiMo-V2.51427.4 +54.8$0.175/M85%
DeepSeek V3.21424.8 +52.2$0.302/M74%
DeepSeek V3.2 Exp1422.6 +50.0$0.305/M73%
DeepSeek V3.11419.7 +47.1$0.425/M63%
Qwen3.5-122B-A10B1417.9 +45.3$0.715/M38%
Gemini 2.5 Flash1417.3 +44.7$0.850/M26%
DeepSeek V3.1 Terminus1416.9 +44.3$0.453/M61%
Gemini 3.1 Flash Lite Preview1415.3 +42.7$0.563/M51%
Qwen3 235B A22B Thinking 25071415.1 +42.5$0.747/M35%
Inkling Small1412.3 +39.7$0.637/M45%
Qwen3.5-27B1408.1 +35.5$0.536/M53%
MiniMax M2.71404.6 +32.0$0.525/M54%
Step 3.5 Flash1403.7 +31.1$0.150/M87%
Qwen3.5-Flash1397.7 +25.1$0.114/M90%
Qwen3.5-35B-A3B1395.5 +22.9$0.547/M52%
GLM 4.5 Air1383.7 +11.1$0.310/M73%
Solar Pro 41377.3 +4.7$0.158/M86%
GLM 4.6V1376.5 +3.9$0.450/M61%
GPT-5 Mini1373.0 +0.4$0.688/M40%
GPT-5.4 Nano1372.9 +0.3$0.463/M60%

Higher score, lower price, named losses

Not a drop-in. The loss is why these are not a recommendation.

ModelLMArenaEffective $/MYou give up
Qwen3 VL 235B A22B Instruct1420.6$0.632/Mreasoning
Qwen3 235B A22B Instruct 25071419.8$0.153/Mreasoning
Qwen3 Next 80B A3B Instruct1417.6$0.343/Mreasoning
Qwen3 30B A3B Instruct 25071383.6$0.084/Mreasoning
DeepSeek V3 03241375.0$0.438/Mreasoning
GLM 5.3 Flash1471.9 · $0.237/MR11372.6 · $1.15/MLMArena EloEffective $/M · Balanced
136 rated, priced standard models. Chartreuse is the frontier. Cobalt is the model you named, when it is not on the staircase.

Open the frontier

The link keeps the model and mix, using the current catalogue. Monthly spend and switching cost stay in this browser tab and are left out of the link.

Saved decision references

Save the model, workload, catalogue date and capability-preservation rule in this browser. Spend, switching cost and bill contents are not saved. No account or notifications.

Constraint: replacements must preserve the model’s capabilities; any losses remain named trade-offs.

Historical decisions cannot be fully replayed from saved references: past prices, scores and capability evidence are not stored.

Evidence & Ask