MODEL PROOF
GLM 4.7 Flash
BETTER VALUE OPTION
A stored alternative has an equal-or-higher measured score and an equal-or-lower price.
Serving precision differs between offers.
Gemma 4 26B A4B · Δ score 83 points · 13% lower measured price · $0.127 / 1M tokens
Qwen3.5-Flash is a further stored option with these losses: 118K → 66K max output
- Independent LMArena score
- 1,350.9 ± 5.68 · 11,977 votes
- Context window
- 200K
- Maximum output
- 117.96K
- Input modalities
- text
- Output modalities
- text
- Published input price
- $0.060 / 1M tokens
- Published output price
- $0.400 / 1M tokens
- Pricing kind
- fixed
Inspect complete billing conditions and endpoint terms below.
Price from Venice: the cheapest offer that is serving, at standard delivery and at a declared precision that is not 4-bit (fp8), at its standard rate. The maker’s own precision is not known.
Every price dimension · USD per million tokens
| Input | $0.060/M |
|---|---|
| Output | $0.400/M |
| Cached input | $0.010/M |
| Cache write | Unknown/M |
| Cache write 1h | Unknown/M |
| Reasoning | Unknown/M |
Who sells it
The same weights, different shops. Cheapest is not like-for-like when serving precision differs.
Headline cheapest is a lower precision. Like-for-like at the best declared precision is $0.060/M from Venice.
3 provider offers · rates, limits and conditions
Rates are USD per million tokens. Endpoint terms can differ even when the quoted price and precision match. A listed parameter is a provider declaration, not a task-success test.
| Seller | Input | Output | Precision | Uptime (1d) |
|---|---|---|---|---|
| Venicevenice/fp8 Endpoint terms
| $0.060 | $0.400 | fp8 | 98.13% |
| Cloudflarecloudflare Endpoint terms
| $0.061 | $0.400 | undeclared | 99.32% |
| Novitanovita/bf16 Endpoint terms
| $0.070 | $0.400 | bf16 | 87.74% |
Independent scores
| LMArena Elo | 1350.9default |
|---|
LMArena Elo under CC BY 4.0. Absence of another board is not a score of zero.
Best rankings by task
- UI components #49 1214
- Websites #57 1201
- Code categories #59 1183
- Game development #72 1148
- 3D #73 1128
- SVG #75 1037
- Data visualisation #86 1130
Per-task Elo and rank from Design Arena, via OpenRouter’s model feed. The rank is the feed’s, and it counts models OpenRouter does not list.
Capability
| Context window | 200K |
|---|---|
| Max output | 117K |
| Input modes | text |
| Tool use | yes |
| Reasoning | optional |
| Open weights | Unknown |
Provenance
| Price source | openrouter.ai |
|---|---|
| Fetched | 2026-10-06 |
| Quality data | verified |
As a 30B-class SOTA model, GLM-4.7-Flash offers a new option that balances performance and efficiency. It is further optimized for agentic coding use cases, strengthening coding capabilities, long-horizon task planning,...
Questions this page answers
What does GLM 4.7 Flash cost?
$0.060/M in, $0.400/M out in this catalogue, as of the fetch date on this page. That is the API row, not a subscription.
Does GLM 4.7 Flash have an independent quality score?
GLM 4.7 Flash has an LMArena score in this catalogue.
What beats GLM 4.7 Flash?
Gemma 4 26B A4B has an equal-or-higher measured score and an equal-or-lower price.
Does this page use Artificial Analysis scores for GLM 4.7 Flash?
No. Artificial Analysis figures are not published here. Quality on this page is LMArena Elo where a score exists; otherwise the row is unrated.
Is the cheapest GLM 4.7 Flash endpoint the same product?
Headline cheapest is a lower precision. Like-for-like at the best declared precision is $0.060/M from Venice.
MODEL MONUMENT