---
title: "Meta: Llama 3.3 70B Instruct — price, capability and what beats it · Undominated.ai"
canonical: https://undominated.ai/models/meta-llama__llama-3.3-70b-instruct/
description: "Meta: Llama 3.3 70B Instruct: $0.135/M in, $0.400/M out. Independent capability scores, and alternatives compared under recorded workload and capability requirements."
---

# Meta: Llama 3.3 70B Instruct — price, capability and what beats it · Undominated.ai

> Meta: Llama 3.3 70B Instruct: $0.135/M in, $0.400/M out. Independent capability scores, and alternatives compared under recorded workload and capability requirements.

[Check base rates for this workload](/check/meta-llama__llama-3.3-70b-instruct/) [Cheaper alternatives](/alternatives/meta-llama__llama-3.3-70b-instruct/) [Family](/families/llama/) [Compare models](https://undominated.ai/compare/?models=meta-llama%2Fllama-3.3-70b-instruct) Comparison context

Choose up to three standard models in Compare.

Requirements and prompt length apply in Compare. This page keeps its stated price basis and workload controls.

MODEL PROOF

# Llama 3.3 70B Instruct

Meta · 06 Dec 2024

 Balanced Summarise Chat Code gen Agentic

BETTER VALUE OPTION

## A stored alternative has an equal-or-higher measured score and an equal-or-lower price.

 **$0.201** / 1M tokens Balanced · 3 tokens in per 1 out

Serving precision differs between offers.

[DeepSeek V4.1 Flash](/models/deepseek__deepseek-v4.1-flash/) · Δ score 188.4 points · 9% lower measured price · $0.182 / 1M tokens

[Gemma 3 4B](/models/google__gemma-3-4b-it/) is a further stored option with these losses: no tool use

Evidence as of 06 Oct 2026

 [openrouter/models](https://openrouter.ai/api/v1/models) [LMArena · cc-by-4.0](https://huggingface.co/datasets/lmarena-ai/leaderboard-dataset) Independent LMArena score 1,274.1 ± 3.49 · 54,366 votes
 Context window 131.07K
 Maximum output 16.38K
 Input modalities text
 Output modalities text
 Published input price $0.135 / 1M tokens
 Published output price $0.400 / 1M tokens
 Pricing kind fixed

Inspect [complete billing conditions](#billing-details) and endpoint terms below.

Price from Novita: the cheapest offer that is serving, at standard delivery and at a declared precision that is not 4-bit (bf16), at its standard rate. The maker’s own precision is not known.

Cheaper right now: $0.155/M at DeepInfra — delivery tier (turbo) It does not pass the like-for-like test, so it never sets a rank.

 Every price dimension · USD per million tokens

| Input | $0.135 /M |
| --- | --- |
| Output | $0.400 /M |
| Cached input | Unknown /M |
| Cache write | Unknown /M |
| Cache write 1h | Unknown /M |
| Reasoning | Unknown /M |

### Who sells it

The same weights, different shops. Cheapest is not like-for-like when serving precision differs.

Headline cheapest is a lower precision. Like-for-like at the best declared precision is $0.135/M from Novita.

 11 provider offers · rates, limits and conditions

Rates are USD per million tokens. Endpoint terms can differ even when the quoted price and precision match. A listed parameter is a provider declaration, not a task-success test.

| Seller | Input | Output | Precision | Uptime (1d) |
| --- | --- | --- | --- | --- |
| Seller DeepInfra deepinfra/turbo Endpoint terms Cached input /M Unknown Context limit 131,072 tokens Output limit 16,384 tokens Tools Listed by endpoint Reasoning Not listed by endpoint Promotional discount None reported Uptime · last 30 minutes 95.94% | Input $0.100 | Output $0.320 | Precision fp8 | Uptime (1d) 98.28% |
| Seller Novita novita/bf16 Endpoint terms Cached input /M Unknown Context limit 12,288 tokens Output limit 11,059 tokens Tools Listed by endpoint Reasoning Not listed by endpoint Promotional discount None reported Uptime · last 30 minutes 99.62% | Input $0.135 | Output $0.400 | Precision bf16 | Uptime (1d) 97.61% |
| Seller AkashML akashml/fp8 Endpoint terms Cached input /M $0.100 Context limit 131,072 tokens Output limit 128,000 tokens Tools Listed by endpoint Reasoning Not listed by endpoint Promotional discount None reported Uptime · last 30 minutes 99.54% | Input $0.200 | Output $0.520 | Precision fp8 | Uptime (1d) 99.23% |
| Seller Parasail parasail/fp8 Endpoint terms Cached input /M $0.110 Context limit 131,072 tokens Output limit 16,384 tokens Tools Not listed by endpoint Reasoning Not listed by endpoint Promotional discount None reported Uptime · last 30 minutes 99.42% | Input $0.220 | Output $0.500 | Precision fp8 | Uptime (1d) 99.66% |
| Seller Cloudflare cloudflare/fp8 Endpoint terms Cached input /M Unknown Context limit 24,000 tokens Output limit 21,600 tokens Tools Not listed by endpoint Reasoning Not listed by endpoint Promotional discount None reported Uptime · last 30 minutes 98.82% | Input $0.293 | Output $2.25 | Precision fp8 | Uptime (1d) 98.72% |
| Seller SambaNova sambanova-turbo Endpoint terms Cached input /M Unknown Context limit 131,072 tokens Output limit 3,072 tokens Tools Not listed by endpoint Reasoning Not listed by endpoint Promotional discount 25% · already included in these rates Uptime · last 30 minutes 99.62% | Input $0.450 | Output $0.900 | Precision undeclared | Uptime (1d) 99.14% |
| Seller Groq groq Endpoint terms Cached input /M $0.295 Context limit 131,072 tokens Output limit 32,768 tokens Tools Listed by endpoint Reasoning Not listed by endpoint Promotional discount None reported Uptime · last 30 minutes 99.77% | Input $0.590 | Output $0.790 | Precision undeclared | Uptime (1d) 99.86% |
| Seller CoreWeave coreweave/fp16 Endpoint terms Cached input /M $0.710 Context limit 128,000 tokens Output limit 115,200 tokens Tools Listed by endpoint Reasoning Not listed by endpoint Promotional discount None reported Uptime · last 30 minutes 100.00% | Input $0.710 | Output $0.710 | Precision fp16 | Uptime (1d) 97.36% |
| Seller Google google-vertex/us-central1 Endpoint terms Cached input /M Unknown Context limit 128,000 tokens Output limit 8,192 tokens Tools Listed by endpoint Reasoning Not listed by endpoint Promotional discount None reported Uptime · last 30 minutes Unknown | Input $0.720 | Output $0.720 | Precision undeclared | Uptime (1d) — |
| Seller Google google-vertex Endpoint terms Cached input /M Unknown Context limit 128,000 tokens Output limit 115,200 tokens Tools Not listed by endpoint Reasoning Not listed by endpoint Promotional discount None reported Uptime · last 30 minutes Unknown | Input $0.720 | Output $0.720 | Precision undeclared | Uptime (1d) — |
| Seller Together together Endpoint terms Cached input /M Unknown Context limit 131,072 tokens Output limit 2,048 tokens Tools Listed by endpoint Reasoning Not listed by endpoint Promotional discount None reported Uptime · last 30 minutes 99.55% | Input $1.04 | Output $1.04 | Precision undeclared | Uptime (1d) 93.11% |

### Independent scores

| LMArena Elo | 1274.1 default |
| --- | --- |

LMArena Elo under CC BY 4.0. Absence of another board is not a score of zero.

### Capability

| Context window | 131K |
| --- | --- |
| Max output | 16K |
| Input modes | text |
| Tool use | yes |
| Reasoning | no |
| Knowledge cutoff | 2023-12-31 |
| Open weights | Unknown |

### Provenance

| Price source | [openrouter.ai](https://openrouter.ai/api/v1/models) |
| --- | --- |
| Fetched | 2026-10-06 |
| Quality data | verified |

The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out). The Llama 3.3 instruction tuned text only model...

## Questions this page answers

 What does Llama 3.3 70B Instruct cost?

$0.135/M in, $0.400/M out in this catalogue, as of the fetch date on this page. That is the API row, not a subscription.

 Does Llama 3.3 70B Instruct have an independent quality score?

Llama 3.3 70B Instruct has an LMArena score in this catalogue.

 What beats Llama 3.3 70B Instruct?

DeepSeek V4.1 Flash has an equal-or-higher measured score and an equal-or-lower price.

 Does this page use Artificial Analysis scores for Llama 3.3 70B Instruct?

No. Artificial Analysis figures are not published here. Quality on this page is LMArena Elo where a score exists; otherwise the row is unrated.

 Is the cheapest Llama 3.3 70B Instruct endpoint the same product?

Headline cheapest is a lower precision. Like-for-like at the best declared precision is $0.135/M from Novita.

MODEL MONUMENT

## Save or share this proof

 Download proof card · 1200 × 630 Preview and more sharing options

Meta: Llama 3.3 70B Instruct. Balanced workload. Better value option available. Evidence as of 06 Oct 2026.

 Copy model link Download square card · 1080 × 1080 Copy landscape image Share landscape card

## Related

 - Both better and cheaper [DeepSeek V4.1 Flash](/models/deepseek__deepseek-v4.1-flash/)
- Higher score and cheaper, with a named trade [Gemma 3 4B](/models/google__gemma-3-4b-it/)
- Nearby on the capability axis [GPT-4o (2024-08-06)](/models/openai__gpt-4o-2024-08-06/)
- Same vendor [Llama 3.1 70B Instruct](/models/meta-llama__llama-3.1-70b-instruct/)
- Meta models [Meta](/providers/meta/)
- Where this model appears [wire](/wire/)
- Nearby on the capability axis [GPT-4 Turbo](/models/openai__gpt-4-turbo/)
- Same vendor [Llama 3.1 8B Instruct](/models/meta-llama__llama-3.1-8b-instruct/)
- Where this model appears [spreads](/spreads/)
- Nearby on the capability axis [Qwen2.5 72B Instruct](/models/qwen__qwen-2.5-72b-instruct/)
- Same vendor [Llama 3.2 3B Instruct](/models/meta-llama__llama-3.2-3b-instruct/)
- Where this model appears [traps](/traps/)

## Badge

The verdict as one SVG, for a README or a docs page. It states this model's dominance status at the balanced workload on the LMArena lens, with the date it was computed, and it is rebuilt with the catalogue — so it changes when the verdict changes, including to one you would rather it did not.

 Markdown Copy
 [![Llama 3.3 70B Instruct — Undominated.ai dominance verdict](https://undominated.ai/badge/meta-llama__llama-3.3-70b-instruct.svg)](https://undominated.ai/models/meta-llama__llama-3.3-70b-instruct/)
 HTML Copy
 <a href="https://undominated.ai/models/meta-llama__llama-3.3-70b-instruct/"><img src="https://undominated.ai/badge/meta-llama__llama-3.3-70b-instruct.svg" alt="Llama 3.3 70B Instruct — Undominated.ai dominance verdict" height="36"></a>

Direct file: [https://undominated.ai/badge/meta-llama__llama-3.3-70b-instruct.svg](https://undominated.ai/badge/meta-llama__llama-3.3-70b-instruct.svg) — a static SVG, written by scripts/build-badges.mjs on every build, so the copy you embed is never older than the last deploy.

## Continue your investigation

 - [Explore model families](/families/)
- [Compare a shortlist](/compare/)
- [Check a model](/check/)
- [Understand the evidence](/methodology/)
