MODEL PROOF
Llama 3.1 8B Instruct
FRONTIER
Nothing is both better and cheaper under this workload.
Serving precision differs between offers.
- Independent LMArena score
- 1,186.5 ± 4.1 · 49,605 votes
- Context window
- 131.07K
- Maximum output
- 117.96K
- Input modalities
- text
- Output modalities
- text
- Published input price
- $0.020 / 1M tokens
- Published output price
- $0.040 / 1M tokens
- Pricing kind
- fixed
Inspect complete billing conditions and endpoint terms below.
Price from DeepInfra: the cheapest offer that is serving, at standard delivery and at a declared precision that is not 4-bit (fp8), at its standard rate. The maker’s own precision is not known.
Every price dimension · USD per million tokens
| Input | $0.020/M |
|---|---|
| Output | $0.040/M |
| Cached input | Unknown/M |
| Cache write | Unknown/M |
| Cache write 1h | Unknown/M |
| Reasoning | Unknown/M |
Who sells it
The same weights, different shops. Cheapest is not like-for-like when serving precision differs.
Headline cheapest is a lower precision. Like-for-like at the best declared precision is $0.220/M from CoreWeave.
5 provider offers · rates, limits and conditions
Rates are USD per million tokens. Endpoint terms can differ even when the quoted price and precision match. A listed parameter is a provider declaration, not a task-success test.
| Seller | Input | Output | Precision | Uptime (1d) |
|---|---|---|---|---|
| DeepInfradeepinfra/fp8 Endpoint terms
| $0.020 | $0.040 | fp8 | 99.91% |
| Novitanovita/fp8 Endpoint terms
| $0.020 | $0.050 | fp8 | 98.92% |
| Groqgroq Endpoint terms
| $0.050 | $0.080 | undeclared | 99.67% |
| Cloudflarecloudflare/fp8 Endpoint terms
| $0.152 | $0.287 | fp8 | 99.26% |
| CoreWeavecoreweave/bf16 Endpoint terms
| $0.220 | $0.220 | bf16 | 100.00% |
Independent scores
| LMArena Elo | 1186.5default |
|---|
LMArena Elo under CC BY 4.0. Absence of another board is not a score of zero.
Capability
| Context window | 131K |
|---|---|
| Max output | 117K |
| Input modes | text |
| Tool use | yes |
| Reasoning | no |
| Knowledge cutoff | 2023-12-31 |
| Open weights | Unknown |
Provenance
| Price source | openrouter.ai |
|---|---|
| Fetched | 2026-10-06 |
| Quality data | verified |
Meta's latest class of model (Llama 3.1) launched with a variety of sizes & flavors. This 8B instruct-tuned version is fast and efficient. It has demonstrated strong performance compared to...
Questions this page answers
What does Llama 3.1 8B Instruct cost?
$0.020/M in, $0.040/M out in this catalogue, as of the fetch date on this page. That is the API row, not a subscription.
Does Llama 3.1 8B Instruct have an independent quality score?
Llama 3.1 8B Instruct has an LMArena score in this catalogue.
What beats Llama 3.1 8B Instruct?
Nothing is both better and cheaper than Llama 3.1 8B Instruct.
Does this page use Artificial Analysis scores for Llama 3.1 8B Instruct?
No. Artificial Analysis figures are not published here. Quality on this page is LMArena Elo where a score exists; otherwise the row is unrated.
Is the cheapest Llama 3.1 8B Instruct endpoint the same product?
Headline cheapest is a lower precision. Like-for-like at the best declared precision is $0.220/M from CoreWeave.
MODEL MONUMENT