Price from Novita: the cheapest offer that is serving, at standard delivery and at a declared precision that is not 4-bit (bf16), at its standard rate. The maker’s own precision is not known.
Cheaper right now: $0.155/M at DeepInfra — delivery tier (turbo) It does not pass the like-for-like test, so it never sets a rank.
Every price dimension · USD per million tokens
Input
$0.135/M
Output
$0.400/M
Cached input
Unknown/M
Cache write
Unknown/M
Cache write 1h
Unknown/M
Reasoning
Unknown/M
Who sells it
The same weights, different shops. Cheapest is not like-for-like when serving precision differs.
Headline cheapest is a lower precision. Like-for-like at the best declared precision is $0.135/M from Novita.
11 provider offers · rates, limits and conditions
Rates are USD per million tokens. Endpoint terms can differ even when the quoted price and precision match. A listed parameter is a provider declaration, not a task-success test.
Seller
Input
Output
Precision
Uptime (1d)
Seller
DeepInfradeepinfra/turbo Endpoint terms
Cached input /M
Unknown
Context limit
131,072 tokens
Output limit
16,384 tokens
Tools
Listed by endpoint
Reasoning
Not listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
95.94%
Input$0.100
Output$0.320
Precisionfp8
Uptime (1d)98.28%
Seller
Novitanovita/bf16 Endpoint terms
Cached input /M
Unknown
Context limit
12,288 tokens
Output limit
11,059 tokens
Tools
Listed by endpoint
Reasoning
Not listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
99.62%
Input$0.135
Output$0.400
Precisionbf16
Uptime (1d)97.61%
Seller
AkashMLakashml/fp8 Endpoint terms
Cached input /M
$0.100
Context limit
131,072 tokens
Output limit
128,000 tokens
Tools
Listed by endpoint
Reasoning
Not listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
99.54%
Input$0.200
Output$0.520
Precisionfp8
Uptime (1d)99.23%
Seller
Parasailparasail/fp8 Endpoint terms
Cached input /M
$0.110
Context limit
131,072 tokens
Output limit
16,384 tokens
Tools
Not listed by endpoint
Reasoning
Not listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
99.42%
Input$0.220
Output$0.500
Precisionfp8
Uptime (1d)99.66%
Seller
Cloudflarecloudflare/fp8 Endpoint terms
Cached input /M
Unknown
Context limit
24,000 tokens
Output limit
21,600 tokens
Tools
Not listed by endpoint
Reasoning
Not listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
98.82%
Input$0.293
Output$2.25
Precisionfp8
Uptime (1d)98.72%
Seller
SambaNovasambanova-turbo Endpoint terms
Cached input /M
Unknown
Context limit
131,072 tokens
Output limit
3,072 tokens
Tools
Not listed by endpoint
Reasoning
Not listed by endpoint
Promotional discount
25% · already included in these rates
Uptime · last 30 minutes
99.62%
Input$0.450
Output$0.900
Precisionundeclared
Uptime (1d)99.14%
Seller
Groqgroq Endpoint terms
Cached input /M
$0.295
Context limit
131,072 tokens
Output limit
32,768 tokens
Tools
Listed by endpoint
Reasoning
Not listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
99.77%
Input$0.590
Output$0.790
Precisionundeclared
Uptime (1d)99.86%
Seller
CoreWeavecoreweave/fp16 Endpoint terms
Cached input /M
$0.710
Context limit
128,000 tokens
Output limit
115,200 tokens
Tools
Listed by endpoint
Reasoning
Not listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
100.00%
Input$0.710
Output$0.710
Precisionfp16
Uptime (1d)97.36%
Seller
Googlegoogle-vertex/us-central1 Endpoint terms
Cached input /M
Unknown
Context limit
128,000 tokens
Output limit
8,192 tokens
Tools
Listed by endpoint
Reasoning
Not listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
Unknown
Input$0.720
Output$0.720
Precisionundeclared
Uptime (1d)—
Seller
Googlegoogle-vertex Endpoint terms
Cached input /M
Unknown
Context limit
128,000 tokens
Output limit
115,200 tokens
Tools
Not listed by endpoint
Reasoning
Not listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
Unknown
Input$0.720
Output$0.720
Precisionundeclared
Uptime (1d)—
Seller
Togethertogether Endpoint terms
Cached input /M
Unknown
Context limit
131,072 tokens
Output limit
2,048 tokens
Tools
Listed by endpoint
Reasoning
Not listed by endpoint
Promotional discount
None reported
Uptime · last 30 minutes
99.55%
Input$1.04
Output$1.04
Precisionundeclared
Uptime (1d)93.11%
Independent scores
LMArena Elo
1274.1default
LMArena Elo under CC BY 4.0. Absence of another board is not a score of zero.
The Meta Llama 3.3 multilingual large language model (LLM) is a pretrained and instruction tuned generative model in 70B (text in/text out). The Llama 3.3 instruction tuned text only model...
Questions this page answers
What does Llama 3.3 70B Instruct cost?
$0.135/M in, $0.400/M out in this catalogue, as of the fetch date on this page. That is the API row, not a subscription.
Does Llama 3.3 70B Instruct have an independent quality score?
Llama 3.3 70B Instruct has an LMArena score in this catalogue.
What beats Llama 3.3 70B Instruct?
DeepSeek V4.1 Flash has an equal-or-higher measured score and an equal-or-lower price.
Does this page use Artificial Analysis scores for Llama 3.3 70B Instruct?
No. Artificial Analysis figures are not published here. Quality on this page is LMArena Elo where a score exists; otherwise the row is unrated.
Is the cheapest Llama 3.3 70B Instruct endpoint the same product?
Headline cheapest is a lower precision. Like-for-like at the best declared precision is $0.135/M from Novita.
MODEL MONUMENT
Save or share this proof
Preview and more sharing options
Meta: Llama 3.3 70B Instruct. Balanced workload. Better value option available. Evidence as of 06 Oct 2026.
The verdict as one SVG, for a README or a docs page. It states this model's dominance
status at the balanced workload on the LMArena lens, with the date it was computed, and
it is rebuilt with the catalogue — so it changes when the verdict changes, including to
one you would rather it did not.