This model was retired on 2027-05-19. The pricing below is the last-known rate, kept for migration reference.
Gemini 3.5 FlashDeprecated
VisionTool useReasoningAudioPrompt cachingStructured outputWeb search
Context
1M
Max output
66K
Serving providers
5
Cheapest input
$1.50 /1M
Prices are on-demand list rates. Enterprise commitments (committed-use discounts, provisioned throughput, negotiated rates) can be materially lower.
Benchmarks
independent evaluations · source: ARC Prize and Artificial AnalysisSpeed & latency
Output speed
204 t/s
tokens / second
Time to first token
19.87s
latency
Headline indices
Intelligence Index
52.0
of 100
Coding Index
70.1
of 100
Agentic Index
39.7
of 100
Individual benchmarks · click a row for effort variants
| Benchmark | Best score | Effort | Cost / task | |||||||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Aa Analyst Agent | 0.5 | default | — | |||||||||||||||||
| AA-Briefcase | 18% | — | — | |||||||||||||||||
| APEX Agents | 47% | default | — | |||||||||||||||||
| ARC-AGI-2 | 72% | high | $0.850 | |||||||||||||||||
| ||||||||||||||||||||
| Automation Bench | 0.4 | default | — | |||||||||||||||||
| CritPt | 13% | default | — | |||||||||||||||||
| ||||||||||||||||||||
| GDPval | 1343.4 | default | — | |||||||||||||||||
| GPQA Diamond | 92% | default | — | |||||||||||||||||
| ||||||||||||||||||||
| Harvey Lab | 0.8 | default | — | |||||||||||||||||
| Humanity's Last Exam | 43% | default | — | |||||||||||||||||
| ||||||||||||||||||||
| IFBench | 76% | default | — | |||||||||||||||||
| ||||||||||||||||||||
| IT-Bench SRE | 40% | default | — | |||||||||||||||||
| Long-Context Reasoning | 81% | default | — | |||||||||||||||||
| ||||||||||||||||||||
| Mlcr Overall | 0.2 | default | — | |||||||||||||||||
| MMMU-Pro | 84% | default | — | |||||||||||||||||
| ||||||||||||||||||||
| Omniscience | 21.2 | default | — | |||||||||||||||||
| ||||||||||||||||||||
| Omniscience Accuracy | 0.5 | default | — | |||||||||||||||||
| ||||||||||||||||||||
| Omniscience Non Hallucination | 0.4 | medium | — | |||||||||||||||||
| ||||||||||||||||||||
| SciCode | 53% | default | — | |||||||||||||||||
| ||||||||||||||||||||
| TerminalBench Hard | 46% | minimal | — | |||||||||||||||||
| ||||||||||||||||||||
| TerminalBench v2.1 | 79% | default | — | |||||||||||||||||
| τ-bench Banking | 32% | default | — | |||||||||||||||||
| τ²-bench | 96% | medium | — | |||||||||||||||||
| ||||||||||||||||||||
Pricing detail
cost beyond the standard rate · source: models.devAudio
$1.50 / n/a
input / output · per 1M
Pricing by serving provider
last-known · per 1M tokens · click a provider to expand| Serving provider | Input /1M | Output /1M | Endpoints | |||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Vertex AI | $1.50 | $9.00 | 1 | |||||||||||||
| ||||||||||||||||
| $1.50 | $9.00 | 1 | ||||||||||||||
| ||||||||||||||||
| Vertex AI | $1.50 | $9.00 | 1 | |||||||||||||
| ||||||||||||||||
| DeepInfra | $1.50 | $9.00 | 1 | |||||||||||||
| ||||||||||||||||
| OpenRouter | $1.50 | $9.00 | 1 | |||||||||||||
| ||||||||||||||||
Price history
input + output $/1M since we started tracking Input Output
Cost calculator
estimate your monthly spend on this model$1,020
estimated / month
Model IDs
copy the exact identifier for your platformvertex_ai/gemini-3.5-flashgemini/gemini-3.5-flashgemini-3.5-flashdeepinfra/google/gemini-3.5-flashgoogle/gemini-3.5-flash
Frequently asked questions
Gemini 3.5 Flash pricing, context and availabilityIs Gemini 3.5 Flash still available?
Gemini 3.5 Flash was retired on 2027-05-19. The pricing on this page is the last-known rate, kept for migration reference.
How much did Gemini 3.5 Flash cost?
Gemini 3.5 Flash's last-known pricing, before it was retired on 2027-05-19, was $1.50 per 1M input tokens and $9.00 per 1M output tokens.
What was the context window of Gemini 3.5 Flash?
Gemini 3.5 Flash had a 1M token context window and could return up to 66K output tokens.
What could Gemini 3.5 Flash do?
Gemini 3.5 Flash supported image input (vision), tool use, reasoning, audio input, prompt caching, structured output and web search.
Other Google models
compare pricing across the Google lineupPopular comparisons
head-to-head pages featuring Gemini 3.5 Flash Pricing verified against LiteLLM + OpenRouter + provider pages · refreshed daily
Every weekday
AI moves fast. Here's your debrief.
News, analysis, tools, and more.
For people who build with AI