Open weights against closed
Rated models only, on the LMArena general board. "Open weights" is the catalogue's own flag: the model's weights are available. It is not "open source": licence and commercial-use terms still vary, and the flag says nothing about training data or code.
Best against best
The best closed model scores 35.8 points above the best open-weight one. That is more than the two half-widths together (13.81), so the leaderboard separates them.
Half-widths are LMArena's published intervals. Two scores are separable when they differ by more than the two half-widths together, the same rule as the leaderboard's significance rank.
On the value frontier
Of the 12 models on the leaderboard's value frontier, 2 are open-weight, 2 closed, and 8 have no recorded weights status. The frontier is the whole board's, at the balanced workload, not one drawn inside each group.
Price at matched score floors
| Floor | Cheapest open-weight | Cheapest closed | Closed ÷ open | Cheapest, no recorded status |
|---|---|---|---|---|
| ≥ 1350 | gpt-oss-120b 1365.4 · $0.065 | GPT-6 Luna 1391.5 · $0.200 | 3.1× | Gemma 3 27B 1357.8 · $0.100 |
| ≥ 1400 | DeepSeek V4 Flash 0423 1432.1 · $0.113 | GPT-5.6 Luna 1431.0 · $0.450 | 4.0× | Gemma 4 26B A4B 1433.9 · $0.127 |
| ≥ 1450 | GLM 5.3 1471.4 · $1.00 | Gemini 3.6 Flash 1479.5 · $1.50 | 1.5× | MiMo-V2.6-Flash 1456.4 · $0.175 |
What the flag is
The weights status is not read from the price feed. Open is set model by model, by hand. Closed is
set the same way for 33 of the 58
closed models, and for 25 it follows from the maker’s
own pages, for the lines of models it publishes no weights for: Anthropic’s models, OpenAI’s except
its open gpt-oss family, and Google’s Gemini line but not Gemma. The pages are quoted in data/maker-weights.json. Anthropic says so in a statement; OpenAI’s and Google’s are
terms that forbid extracting the model, weaker evidence than a statement. A model whose record does
not set it has no status, which is why the third group is large. The price feed names a Hugging Face repository for 61 of those 76. A named repository is evidence, not a reviewed status, so they stay
unknown here rather than being guessed into either column. The open-weights list has every flagged model, rated or not.