MODEL PROOF
gpt-oss-120b
FRONTIER
Nothing is both better and cheaper under this workload.
Serving precision differs between offers.
- Independent LMArena score
- 1,365.4 ± 4.37 · 30,018 votes
- Context window
- 131.07K
- Maximum output
- 117.96K
- Input modalities
- text
- Output modalities
- text
- Published input price
- $0.030 / 1M tokens
- Published output price
- $0.170 / 1M tokens
- Pricing kind
- fixed
Inspect complete billing conditions and endpoint terms below.
Price from CoreWeave: the cheapest offer that is serving, at standard delivery and at the maker’s precision or better, at its standard rate.
Every price dimension · USD per million tokens
| Input | $0.030/M |
|---|---|
| Output | $0.170/M |
| Cached input | $0.030/M |
| Cache write | Unknown/M |
| Cache write 1h | Unknown/M |
| Reasoning | Unknown/M |
Who sells it
The same weights, different shops. Cheapest is not like-for-like when serving precision differs.
Headline cheapest is a lower precision. Like-for-like at the best declared precision is $0.030/M from DekaLLM.
23 provider offers · rates, limits and conditions
Rates are USD per million tokens. Endpoint terms can differ even when the quoted price and precision match. A listed parameter is a provider declaration, not a task-success test.
| Seller | Input | Output | Precision | Uptime (1d) |
|---|---|---|---|---|
| DekaLLMdekallm/bf16 Endpoint terms
| $0.030 | $0.180 | bf16 | 99.54% |
| CoreWeavecoreweave/fp4 Endpoint terms
| $0.030 | $0.170 | fp4 | 98.88% |
| DeepInfradeepinfra/bf16 Endpoint terms
| $0.037 | $0.170 | bf16 | 98.74% |
| AkashMLakashml/bf16 Endpoint terms
| $0.037 | $0.187 | bf16 | 99.95% |
| Mancer 2mancer/fp8 Endpoint terms
| $0.045 | $0.250 | fp8 | 97.70% |
| Crusoecrusoe/bf16 Endpoint terms
| $0.050 | $0.250 | bf16 | 99.99% |
| Novitanovita/fp4 Endpoint terms
| $0.050 | $0.250 | fp4 | 98.73% |
| DigitalOceandigitalocean Endpoint terms
| $0.060 | $0.420 | undeclared | 99.98% |
| Googlegoogle-vertex/global Endpoint terms
| $0.090 | $0.360 | undeclared | 66.28% |
| BaseTenbaseten/fp4 Endpoint terms
| $0.100 | $0.500 | fp4 | 99.96% |
| BaseTenbaseten/fp4 Endpoint terms
| $0.100 | $0.500 | fp4 | 99.99% |
| Parasailparasail/fp4 Endpoint terms
| $0.100 | $0.750 | fp4 | 99.97% |
| SambaNovasambanova Endpoint terms
| $0.140 | $0.950 | undeclared | 99.81% |
| DeepInfradeepinfra/turbo Endpoint terms
| $0.150 | $0.600 | bf16 | 99.98% |
| SiliconFlowsiliconflow/fp8 Endpoint terms
| $0.150 | $0.600 | fp8 | 91.35% |
| Nebiusnebius/fp4 Endpoint terms
| $0.150 | $0.600 | fp4 | 97.80% |
| Amazon Bedrockamazon-bedrock/eu-west-1 Endpoint terms
| $0.150 | $0.600 | undeclared | 99.97% |
| Amazon Bedrockamazon-bedrock Endpoint terms
| $0.150 | $0.600 | undeclared | 99.41% |
| Phalaphala Endpoint terms
| $0.150 | $0.600 | undeclared | 99.31% |
| Togethertogether Endpoint terms
| $0.150 | $0.600 | undeclared | 87.70% |
| Groqgroq Endpoint terms
| $0.150 | $0.600 | undeclared | 99.39% |
| Maramara Endpoint terms
| $0.150 | $0.750 | undeclared | 96.50% |
| Cerebrascerebras/fp16 Endpoint terms
| $0.350 | $0.750 | fp16 | 99.98% |
Independent scores
| LMArena Elo | 1365.4default |
|---|
LMArena Elo under CC BY 4.0. Absence of another board is not a score of zero.
Best rankings by task
- Game development #110 1005
- Data visualisation #110 995
- 3D #111 909
- UI components #116 932
- Code categories #120 967
- Websites #122 974
Per-task Elo and rank from Design Arena, via OpenRouter’s model feed. The rank is the feed’s, and it counts models OpenRouter does not list.
Capability
| Context window | 131K |
|---|---|
| Max output | 117K |
| Input modes | text |
| Tool use | yes |
| Reasoning | always on |
| Knowledge cutoff | 2024-06-30 |
| Open weights | yes |
Provenance
| Price source | openrouter.ai |
|---|---|
| Fetched | 2026-10-06 |
| Quality data | verified |
| Cross-checked | vendor page |
This is the gpt-oss-120b API row. A search that names Claude Sonnet 4.6 as the other side is answered on the compare page, not by blending the two SKUs here.
Our take
editorial — not a measurementThe practical reason to consider this model is deployment control: the accepted record identifies open weights and an Apache licence. Review the actual licence and serving requirements before deployment. Hosted offers are seller-specific; self-hosting replaces token charges with hardware, operations and capacity costs rather than eliminating cost. Use the current proof for measured capability and the offer table for hosted prices.
Strengths
- Open weights and a licence identifier are recorded
- Text input, tool use and reasoning support are listed
Weaknesses
- A hosted endpoint can impose its own limits
- Self-hosted capacity and operating costs need a separate estimate
Reach for it when
- Deployment-control and self-hosting evaluations
- Text workloads with explicit capacity planning
Avoid it if
- You require image or audio input
- You cannot operate the serving infrastructure you plan to use
Sources: openrouter.ai
gpt-oss-120b is an open-weight, 117B-parameter Mixture-of-Experts (MoE) language model from OpenAI designed for high-reasoning, agentic, and general-purpose production use cases. It activates 5.1B parameters per forward pass and is optimized...
Questions this page answers
What does gpt-oss-120b cost?
$0.030/M in, $0.170/M out in this catalogue, as of the fetch date on this page. That is the API row, not a subscription.
Does gpt-oss-120b have an independent quality score?
gpt-oss-120b has an LMArena score in this catalogue.
What beats gpt-oss-120b?
Nothing is both better and cheaper than gpt-oss-120b.
Does this page use Artificial Analysis scores for gpt-oss-120b?
No. Artificial Analysis figures are not published here. Quality on this page is LMArena Elo where a score exists; otherwise the row is unrated.
Is the cheapest gpt-oss-120b endpoint the same product?
Headline cheapest is a lower precision. Like-for-like at the best declared precision is $0.030/M from DekaLLM.
MODEL MONUMENT