Is Qwen3.5-Flash a good deal?

Whether anything in this catalogue beats Qwen3.5-Flash on both quality and price, and what you give up if it does. A computation on the current catalogue, not an opinion.

13 undominated of 136 · Sep 23, 2026

As of Sep 23, 2026, DeepSeek V4 Flash 0423 scores higher and costs less than Qwen3.5-Flash for Balanced on LMArena, but it is not a drop-in. You would give up: image, video.

Inspect model evidence Compare differences & requirements

Current model

Qwen3.5-Flash

Balanced

3 tokens in per 1 out

DeepSeek V4 Flash 0423 scores 34.1 points higher and costs 9% less, but you would give up: image, video.

Qwen3.5-Flash

LMArena Elo · higher is better

Scale starts at 1040 Elo

1397.7

Effective $/M · Balanced · lower is better

$0.114/M

DeepSeek V4 Flash 0423

LMArena Elo · higher is better

Scale starts at 1040 Elo

1431.8

Effective $/M · Balanced · lower is better

$0.103/M

Qwen3.5-Flash takes text, image, video, returns up to 65,536 tokens from a 1,000,000-token context, and is listed by 1 seller.

Compared against 136 rated, priced models on this lens: nothing dominates it outright, 1 more dominates it but gives something up. Qwen3.5-Flash scores 1397.7 at $0.11 per million tokens for this mix.

Higher score, lower price, named losses

Not a drop-in. The loss is why these are not a recommendation.

ModelLMArenaEffective $/MYou give up
DeepSeek V4 Flash 04231431.8$0.103/Mimage, video
Qwen3.5-Flash1397.7 · $0.114/MLMArena EloEffective $/M · Balanced
136 rated, priced standard models. Chartreuse is the frontier. Cobalt is the model you named, when it is not on the staircase.

Open the frontier

The link keeps the model and mix, using the current catalogue. Monthly spend and switching cost stay in this browser tab and are left out of the link.

Saved decision references

Save the model, workload, catalogue date and capability-preservation rule in this browser. Spend, switching cost and bill contents are not saved. No account or notifications.

Constraint: replacements must preserve the model’s capabilities; any losses remain named trade-offs.

Historical decisions cannot be fully replayed from saved references: past prices, scores and capability evidence are not stored.

Evidence & Ask