Gemini 3.1 Flash Lite

Google released 2026-05-07 proprietary

$0.563 per million tokens, balanced

No independent quality score.

No benchmark we track has measured this model. That is not the same as measuring it and finding it wanting — we simply cannot rank it, so we do not.

Every price dimension

Input$0.250/M
Output$1.50/M
Cached input$0.025/M
Cache write$0.083/M
Reasoning$1.50/M
Web search$0.014/call
Batch discount50%

Reasoning tokens are billed separately at $1.50/M, on top of output. On a reasoning-heavy workload this can be 40% of the bill and it does not appear in the advertised price.

Independent scores

LMArena Elo1414.8preview

Sources are listed separately rather than averaged. Across the 62 models both have scored they correlate at r = 0.806 — close agreement overall, but four models rank very differently between them, and a blended score would hide exactly those.

Capability

Context window1.0M
Max output66K
Input modestext, image, video, file, audio
Tool useyes
Reasoningoptional
Open weightsno

Provenance

Price sourceopenrouter.ai
Fetched2026-08-24
Quality dataunrated
Cross-checkedvendor page

Gemini 3.1 Flash Lite is Google’s GA high-efficiency multimodal model optimized for low-latency, high-volume workloads. It supports text, image, video, audio, and PDF inputs, and is designed for lightweight agentic...

Markdown for LLMs