Model comparison · accepted catalogue facts 2026-09-23

Gemini 3.1 Pro Preview vs Gemini 3.8 Flash

A quality score or a computable price is missing. This pair has no better-and-cheaper verdict.

Editorial comparison: Stable Flash versus preview Pro Inclusion does not recommend either model.

Gemini 3.1 Pro Preview

LMArena general · Elo · higher is better

Scale starts at 1040 Elo

1480.1

Balanced effective USD / 1M tokens · lower is better

Unknown

Gemini 3.8 Flash

LMArena general · Elo · higher is better

Scale starts at 1040 Elo

1494.7

Balanced effective USD / 1M tokens · lower is better

Unknown

Compare · model decision

Choose two or three models

Inspect the differences, estimate your workload and keep the evidence. All standard catalogue models are available, including unrated ones.

Catalogue facts
2026-09-23

2 / 3 selected · 0 requirements
Open public comparison link
Public workload and requirements

These settings are included in the share link. The fixed-request scenario and private evaluation data below are excluded.

A ↔ B

A quality score or a computable price is missing. This pair has no better-and-cheaper verdict.

Estimated blend: 75% input, 25% output; 0% cached input. Base context tier. Prompt length selects the billing tier. An unlisted cache-read rate uses that tier’s full input rate in this blend; missing required input or output rates prevent an estimate. The fixed-request estimate below validates actual context and output limits and requires a published cache-read rate when cache hits are assumed.

What changes between these models

Quality first, then estimated price. LMArena scores are human-preference measurements, not a benchmark of your application. A missing score is unrated.

  • Gemini 3.1 Pro Preview · Meets the stated model requirements
  • Gemini 3.8 Flash · Meets the stated model requirements
Material differences and unknown evidence. Deltas compare each model with model A. Unknown facts remain visible.
Measured factA Gemini 3.1 Pro PreviewB Gemini 3.8 Flash
LMArena general score (Elo)1,480.11,494.7Δ vs A: +14.6 Elo
Score interval half-width (± Elo)3.148.53Δ vs A: +5.39 Elo
Benchmark reasoning effortdefaulthigh
LMArena document score (Elo)1,443.9Unknown
Estimated effective USD / 1M tokensUnknownUnknown
Base input USD / 1M tokens$2$0.75Δ vs A: -1.25 USD/M
Base output USD / 1M tokens$12$3.75Δ vs A: -8.25 USD/M
Base cached-input USD / 1M tokens$0.2$0.075Δ vs A: -0.125 USD/M
Input modalitiesaudio, file, image, text, videotext, image, video, file, audio
Open weightsNot supportedUnknown
Recorded licenceproprietaryUnknown
Recorded retirement dateUnknownUnknown
8 shared measured facts
Shared measured facts. Deltas compare each model with model A. Unknown facts remain visible.
Measured factA Gemini 3.1 Pro PreviewB Gemini 3.8 Flash
Context window (tokens)1,048,5761,048,576Δ vs A: 0 tokens
Output limit (tokens)65,53665,536Δ vs A: 0 tokens
Tool useSupportedSupportedΔ vs A: 0
ReasoningSupportedSupportedΔ vs A: 0
Model providerGoogleGoogle
Distinct sellers in the catalogue11Δ vs A: 0
Serving precision differs across offersNot supportedNot supportedΔ vs A: 0
Retirement evidenceNone announced; not a guaranteeNone announced; not a guarantee

Provider counts and precision flags describe catalogue offers. They do not establish the precision, availability, region, tool behaviour or limits of your chosen endpoint. Open each model’s provider table before switching.

Complete price ladders and billing conditions

Gemini 3.1 Pro Preview

Estimate unavailable: Separately billed reasoning tokens are not part of this token blend.

Separate reasoning: $12/1M tokens; reasoning-token counts are not modelled here. Rates below are USD per 1M tokens. A threshold selects that whole input/output rate pair.

Input-token thresholdInputOutputCached input
Base, up to and including first threshold$2$12$0.2
> 200,000$4$18$0.4

Gemini 3.8 Flash

Estimate unavailable: Separately billed reasoning tokens are not part of this token blend.

Separate reasoning: $3.75/1M tokens; reasoning-token counts are not modelled here. Rates below are USD per 1M tokens. A threshold selects that whole input/output rate pair.

Input-token thresholdInputOutputCached input
Base, up to and including first threshold$0.75$3.75$0.075
Migration preflight · compare each candidate with model A

Keeping Gemini 3.1 Pro Preview is a valid decision. No pairwise comparison certifies a drop-in replacement.

Gemini 3.1 Pro Preview → Gemini 3.8 Flash

No known model-capability loss

  • Provider endpoint, serving precision, availability and task success need your own verification. Catalogue model capabilities do not certify an endpoint.
Check alternatives to Gemini 3.1 Pro Preview

Estimate a fixed request shape

Private to this tab unless you save locally or explicitly export private details. These are estimates, not observed invoices. Cross-model token counts and task success are assumptions.

0.1 extra attempts means 10 extra full attempts per 100 requests. Cache writes, tool charges, taxes and unlisted charges are excluded.

  • Gemini 3.1 Pro PreviewEstimated total: unknown Separately billed reasoning tokens are not part of this request shape.
  • Gemini 3.8 FlashEstimated total: unknown Separately billed reasoning tokens are not part of this request shape.
Prompt and output sensitivity · sampled scenarios

Every point recalculates the entire request shape. Tier-boundary samples include one token below, at and above each threshold. Ordering between samples is not guaranteed. Lowest known cost excludes unknown estimates and is not a replacement recommendation.

Vary input; hold output, requests, cache and retries fixed.
Input tokensGemini 3.1 Pro PreviewGemini 3.8 FlashCoverage
2,000unknownSeparately billed reasoning tokens are not part of this request shape.unknownSeparately billed reasoning tokens are not part of this request shape.Incomplete
4,096unknownSeparately billed reasoning tokens are not part of this request shape.unknownSeparately billed reasoning tokens are not part of this request shape.Incomplete
32,768unknownSeparately billed reasoning tokens are not part of this request shape.unknownSeparately billed reasoning tokens are not part of this request shape.Incomplete
131,072unknownSeparately billed reasoning tokens are not part of this request shape.unknownSeparately billed reasoning tokens are not part of this request shape.Incomplete
199,999unknownSeparately billed reasoning tokens are not part of this request shape.unknownSeparately billed reasoning tokens are not part of this request shape.Incomplete
200,000unknownSeparately billed reasoning tokens are not part of this request shape.unknownSeparately billed reasoning tokens are not part of this request shape.Incomplete
200,001unknownSeparately billed reasoning tokens are not part of this request shape.unknownSeparately billed reasoning tokens are not part of this request shape.Incomplete
Vary output; hold input, requests, cache and retries fixed.
Output tokensGemini 3.1 Pro PreviewGemini 3.8 Flash
128unknownSeparately billed reasoning tokens are not part of this request shape.unknownSeparately billed reasoning tokens are not part of this request shape.
2,048unknownSeparately billed reasoning tokens are not part of this request shape.unknownSeparately billed reasoning tokens are not part of this request shape.
8,192unknownSeparately billed reasoning tokens are not part of this request shape.unknownSeparately billed reasoning tokens are not part of this request shape.

Bring your observed evaluation results

JSON is parsed locally. Prompts, filenames, individual responses and extra fields are not retained. Import only aggregate measurements for selected models. Task/version labels do not establish that two evaluations used the same dataset.

Expected format
[{"model":"provider/exact-model-id","task":"support-answer","version":"rubric-1","attempts":100,"successes":84,"costUsd":2.4,"latencyMs":950}]

Use the exact catalogue ID. Replace the example values with observed measurements. Omit costUsd or latencyMs when they are unknown. No public score is inferred from these observations.

Keep the decision and its evidence

A portable record retains selected facts, complete tier ladders, public requirements and the publication manifest. It also embeds the full permitted public catalogue to verify the input artifact checksum. Files are larger for this reason (up to 2 MB); local storage is limited to 12 records. Team review is local to this browser; there is no hosted collaboration.

Saving is opt-in and includes private review fields. Shared comparison links never include them. Version 1 Check references remain separate and are not overwritten or guessed into this format.

Saved local decisions · 0

No version 2 decision records are stored in this browser.

Questions this page answers

Is Gemini 3.1 Pro Preview a better-and-cheaper replacement for Gemini 3.8 Flash?

A quality score or a computable price is missing. This pair has no better-and-cheaper verdict.

What is included in this comparison?

Accepted model capabilities, context and output limits, complete price-tier ladders, independent LMArena measurements and missing evidence. Provider endpoint precision and application success still need verification.

Can I compare a different workload or keep a decision record?

Yes. Change the public workload, prompt length and requirements in the comparison workspace. Fixed-request scenarios, aggregate evaluations and review notes stay local unless explicitly exported.

Evidence & Ask