Open weights against closed

Rated models only, on the LMArena general board. "Open weights" is the catalogue's own flag: the model's weights are available. It is not "open source": licence and commercial-use terms still vary, and the flag says nothing about training data or code.

Best against best

Best open-weight

Kimi K3

Score
1475.9 ± 4.76
Significance rank
11
Balanced $/M
$3.76

Best closed

Claude Opus 5.5

Score
1511.7 ± 9.05
Significance rank
1
Balanced $/M
$8.00

The best closed model scores 35.8 points above the best open-weight one. That is more than the two half-widths together (13.81), so the leaderboard separates them.

Half-widths are LMArena's published intervals. Two scores are separable when they differ by more than the two half-widths together, the same rule as the leaderboard's significance rank.

On the value frontier

Of the 12 models on the leaderboard's value frontier, 2 are open-weight, 2 closed, and 8 have no recorded weights status. The frontier is the whole board's, at the balanced workload, not one drawn inside each group.

Price at matched score floors

The cheapest rated model at or above each LMArena floor, per group, balanced workload, USD per million tokens. The two picks need not score alike above the floor.
FloorCheapest open-weightCheapest closedClosed ÷ openCheapest, no recorded status
≥ 1350gpt-oss-120b 1365.4 · $0.065GPT-6 Luna 1391.5 · $0.2003.1×Gemma 3 27B 1357.8 · $0.100
≥ 1400DeepSeek V4 Flash 0423 1432.1 · $0.113GPT-5.6 Luna 1431.0 · $0.4504.0×Gemma 4 26B A4B 1433.9 · $0.127
≥ 1450GLM 5.3 1471.4 · $1.00Gemini 3.6 Flash 1479.5 · $1.501.5×MiMo-V2.6-Flash 1456.4 · $0.175

What the flag is

The weights status is not read from the price feed. Open is set model by model, by hand. Closed is set the same way for 33 of the 58 closed models, and for 25 it follows from the maker’s own pages, for the lines of models it publishes no weights for: Anthropic’s models, OpenAI’s except its open gpt-oss family, and Google’s Gemini line but not Gemma. The pages are quoted in data/maker-weights.json. Anthropic says so in a statement; OpenAI’s and Google’s are terms that forbid extracting the model, weaker evidence than a statement. A model whose record does not set it has no status, which is why the third group is large. The price feed names a Hugging Face repository for 61 of those 76. A named repository is evidence, not a reviewed status, so they stay unknown here rather than being guessed into either column. The open-weights list has every flagged model, rated or not.

Evidence & Ask