MODEL PROOF

Llama 3.3 70B Instruct

BETTER VALUE OPTION

A stored alternative has an equal-or-higher measured score and an equal-or-lower price.

$0.201 / 1M tokens Balanced · 3 tokens in per 1 out

Serving precision differs between offers.

DeepSeek V4.1 Flash · Δ score 188.4 points · 9% lower measured price · $0.182 / 1M tokens

Gemma 3 4B is a further stored option with these losses: no tool use

Independent LMArena score
1,274.1 ± 3.49 · 54,366 votes
Context window
131.07K
Maximum output
16.38K
Input modalities
text
Output modalities
text
Published input price
$0.135 / 1M tokens
Published output price
$0.400 / 1M tokens
Pricing kind
fixed

Inspect complete billing conditions and endpoint terms below.

Price from Novita: the cheapest offer that is serving, at standard delivery and at a declared precision that is not 4-bit (bf16), at its standard rate. The maker’s own precision is not known.

Cheaper right now: $0.155/M at DeepInfra — delivery tier (turbo) It does not pass the like-for-like test, so it never sets a rank.

Every price dimension · USD per million tokens
Input$0.135/M
Output$0.400/M
Cached inputUnknown/M
Cache writeUnknown/M
Cache write 1hUnknown/M
ReasoningUnknown/M

Who sells it

The same weights, different shops. Cheapest is not like-for-like when serving precision differs.

Headline cheapest is a lower precision. Like-for-like at the best declared precision is $0.135/M from Novita.

11 provider offers · rates, limits and conditions

Rates are USD per million tokens. Endpoint terms can differ even when the quoted price and precision match. A listed parameter is a provider declaration, not a task-success test.

SellerInputOutputPrecisionUptime (1d)
DeepInfradeepinfra/turbo
Endpoint terms
Cached input /M
Unknown
Context limit
131,072 tokens
Output limit
16,384 tokens
Tools
Listed by endpoint
Reasoning
Not listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
95.94%
$0.100 $0.320 fp8 98.28%
Novitanovita/bf16
Endpoint terms
Cached input /M
Unknown
Context limit
12,288 tokens
Output limit
11,059 tokens
Tools
Listed by endpoint
Reasoning
Not listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
99.62%
$0.135 $0.400 bf16 97.61%
AkashMLakashml/fp8
Endpoint terms
Cached input /M
$0.100
Context limit
131,072 tokens
Output limit
128,000 tokens
Tools
Listed by endpoint
Reasoning
Not listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
99.54%
$0.200 $0.520 fp8 99.23%
Parasailparasail/fp8
Endpoint terms
Cached input /M
$0.110
Context limit
131,072 tokens
Output limit
16,384 tokens
Tools
Not listed by endpoint
Reasoning
Not listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
99.42%
$0.220 $0.500 fp8 99.66%
Cloudflarecloudflare/fp8
Endpoint terms
Cached input /M
Unknown
Context limit
24,000 tokens
Output limit
21,600 tokens
Tools
Not listed by endpoint
Reasoning
Not listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
98.82%
$0.293 $2.25 fp8 98.72%
SambaNovasambanova-turbo
Endpoint terms
Cached input /M
Unknown
Context limit
131,072 tokens
Output limit
3,072 tokens
Tools
Not listed by endpoint
Reasoning
Not listed by endpoint
Promotional discount
25% · already included in these rates
Uptime · last 30 minutes
99.62%
$0.450 $0.900 undeclared 99.14%
Groqgroq
Endpoint terms
Cached input /M
$0.295
Context limit
131,072 tokens
Output limit
32,768 tokens
Tools
Listed by endpoint
Reasoning
Not listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
99.77%
$0.590 $0.790 undeclared 99.86%
CoreWeavecoreweave/fp16
Endpoint terms
Cached input /M
$0.710
Context limit
128,000 tokens
Output limit
115,200 tokens
Tools
Listed by endpoint
Reasoning
Not listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
100.00%
$0.710 $0.710 fp16 97.36%
Googlegoogle-vertex/us-central1
Endpoint terms
Cached input /M
Unknown
Context limit
128,000 tokens
Output limit
8,192 tokens
Tools
Listed by endpoint
Reasoning
Not listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
Unknown
$0.720 $0.720 undeclared —
Googlegoogle-vertex
Endpoint terms
Cached input /M
Unknown
Context limit
128,000 tokens
Output limit
115,200 tokens
Tools
Not listed by endpoint
Reasoning
Not listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
Unknown
$0.720 $0.720 undeclared —
Togethertogether
Endpoint terms
Cached input /M
Unknown
Context limit
131,072 tokens
Output limit
2,048 tokens
Tools
Listed by endpoint
Reasoning
Not listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
99.55%
$1.04 $1.04 undeclared 93.11%

Independent scores

LMArena Elo1274.1default

LMArena Elo under CC BY 4.0. Absence of another board is not a score of zero.

Capability

Context window131K
Max output16K
Input modestext
Tool useyes
Reasoningno
Knowledge cutoff2023-12-31
Open weightsUnknown

Provenance

Price sourceopenrouter.ai
Fetched2026-10-06
Quality dataverified

The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out). The Llama 3.3 instruction tuned text only model...

Questions this page answers

What does Llama 3.3 70B Instruct cost?

$0.135/M in, $0.400/M out in this catalogue, as of the fetch date on this page. That is the API row, not a subscription.

Does Llama 3.3 70B Instruct have an independent quality score?

Llama 3.3 70B Instruct has an LMArena score in this catalogue.

What beats Llama 3.3 70B Instruct?

DeepSeek V4.1 Flash has an equal-or-higher measured score and an equal-or-lower price.

Does this page use Artificial Analysis scores for Llama 3.3 70B Instruct?

No. Artificial Analysis figures are not published here. Quality on this page is LMArena Elo where a score exists; otherwise the row is unrated.

Is the cheapest Llama 3.3 70B Instruct endpoint the same product?

Headline cheapest is a lower precision. Like-for-like at the best declared precision is $0.135/M from Novita.

MODEL MONUMENT

Save or share this proof

Badge

The verdict as one SVG, for a README or a docs page. It states this model's dominance status at the balanced workload on the LMArena lens, with the date it was computed, and it is rebuilt with the catalogue — so it changes when the verdict changes, including to one you would rather it did not.

Llama 3.3 70B Instruct — Undominated.ai dominance verdict

Markdown
[![Llama 3.3 70B Instruct — Undominated.ai dominance verdict](https://undominated.ai/badge/meta-llama__llama-3.3-70b-instruct.svg)](https://undominated.ai/models/meta-llama__llama-3.3-70b-instruct/)
HTML
<a href="https://undominated.ai/models/meta-llama__llama-3.3-70b-instruct/"><img src="https://undominated.ai/badge/meta-llama__llama-3.3-70b-instruct.svg" alt="Llama 3.3 70B Instruct — Undominated.ai dominance verdict" height="36"></a>

Direct file: https://undominated.ai/badge/meta-llama__llama-3.3-70b-instruct.svg — a static SVG, written by scripts/build-badges.mjs on every build, so the copy you embed is never older than the last deploy.

Evidence & Ask