Context
1M
Max output
66K
Serving providers
3
Cheapest input
$0.30 /1M
Prices are on-demand list rates. Enterprise commitments (committed-use discounts, provisioned throughput, negotiated rates) can be materially lower.
Benchmarks
independent evaluations · source: Artificial AnalysisSpeed & latency
Output speed
300 t/s
tokens / second
Time to first token
18.15s
latency
Headline indices
Individual benchmarks · click a row for effort variants
| Benchmark | Best score | Effort | Cost / task | |
|---|---|---|---|---|
| CritPt | 0% | default | — | |
| GDPval | 1140 | default | — | |
| GPQA Diamond | 84% | default | — | |
| Humanity's Last Exam | 18% | default | — | |
| Long-Context Reasoning | 62% | default | — | |
| MMMU-Pro | 79% | default | — | |
| Omniscience | 6.9 | default | — | |
| SciCode | 41% | default | — | |
| TerminalBench v2.1 | 54% | default | — | |
| τ-bench Banking | 16% | default | — |
Pricing by serving provider
standard tier · per 1M tokens · click a provider to expand| Serving provider | Input /1M | Output /1M | Endpoints | |||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Vertex AICheapest | $0.30 | $2.50 | 1 | |||||||||||||
| ||||||||||||||||
| GoogleCheapest | $0.30 | $2.50 | 1 | |||||||||||||
| ||||||||||||||||
| OpenRouterCheapest | $0.30 | $2.50 | 1 | |||||||||||||
| ||||||||||||||||
Price history
input + output $/1M since we started tracking Input Output
Price history is accruing. We record a point each day a price changes; the full backfill lands shortly.
Cost calculator
estimate your monthly spend on this model$260
estimated / month
Model IDs
copy the exact identifier for your platformgemini-3.5-flash-litegemini/gemini-3.5-flash-litegoogle/gemini-3.5-flash-lite
Frequently asked questions
Gemini 3.5 Flash-Lite pricing, context and availabilityHow much does Gemini 3.5 Flash-Lite cost?
Gemini 3.5 Flash-Lite costs $0.30 per 1M input tokens and $2.50 per 1M output tokens at its cheapest provider via Vertex AI.
What is the context window of Gemini 3.5 Flash-Lite?
Gemini 3.5 Flash-Lite accepts up to 1M tokens of context and can return up to 66K output tokens.
Which providers serve Gemini 3.5 Flash-Lite?
Gemini 3.5 Flash-Lite is available from 3 serving providers, each with its own pricing and model ID. Vertex AI is currently the cheapest.
What can Gemini 3.5 Flash-Lite do?
Gemini 3.5 Flash-Lite supports image input (vision), tool use, reasoning, audio input, prompt caching, structured output and web search.
Other Google models
compare pricing across the Google lineupGemma 3 27B$0.00/1M · 8 providersGemini 2.5 Flash$0.15/1M · 8 providersGemini 2.5 Pro$1.25/1M · 7 providersGemma 3 12B$0.05/1M · 5 providersGemini 3 Flash Preview$0.50/1M · 5 providersGemini 3 Pro Preview$2.00/1M · 5 providersGemini 2.5 Flash Lite$0.07/1M · 4 providersGemini 2.0 Flash 001$0.10/1M · 4 providers
Pricing verified against LiteLLM + OpenRouter + provider pages · refreshed daily
Every weekday
AI moves fast. Here's your debrief.
News, analysis, tools, and more.
For people who build with AI