Context
1M
Max output
66K
Serving providers
1
Cheapest input
$0.63 /1M
Prices are on-demand list rates. Enterprise commitments (committed-use discounts, provisioned throughput, negotiated rates) can be materially lower.
Benchmarks
independent evaluations · source: Artificial AnalysisSpeed & latency
Output speed
187 t/s
tokens / second
Time to first token
0.99s
latency
Headline indices
Individual benchmarks · click a row for effort variants
| Benchmark | Best score | Effort | Cost / task | |||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| APEX Agents | 28% | reasoning: true | — | |||||||||||||
| CritPt | 9% | reasoning: true | — | |||||||||||||
| ||||||||||||||||
| GPQA Diamond | 90% | reasoning: true | — | |||||||||||||
| ||||||||||||||||
| Humanity's Last Exam | 37% | reasoning: true | — | |||||||||||||
| ||||||||||||||||
| IFBench | 78% | reasoning: true | — | |||||||||||||
| ||||||||||||||||
| LiveCodeBench | 91% | reasoning: true | — | |||||||||||||
| ||||||||||||||||
| Long-Context Reasoning | 73% | reasoning: true | — | |||||||||||||
| ||||||||||||||||
| Mlcr Overall | 0.1 | default | — | |||||||||||||
| ||||||||||||||||
| MMLU-Pro | 89% | reasoning: true | — | |||||||||||||
| ||||||||||||||||
| MMMU-Pro | 80% | reasoning: true | — | |||||||||||||
| ||||||||||||||||
| Omniscience | 10.1 | reasoning: true | — | |||||||||||||
| ||||||||||||||||
| Omniscience Accuracy | 0.5 | reasoning: true | — | |||||||||||||
| ||||||||||||||||
| Omniscience Non Hallucination | 0.1 | default | — | |||||||||||||
| ||||||||||||||||
| SciCode | 51% | reasoning: true | — | |||||||||||||
| ||||||||||||||||
| TerminalBench Hard | 39% | reasoning: true | — | |||||||||||||
| ||||||||||||||||
| τ-bench Banking | 21% | reasoning: true | — | |||||||||||||
| τ²-bench | 80% | reasoning: true | — | |||||||||||||
| ||||||||||||||||
Pricing detail
cost beyond the standard rate · source: models.devAudio
$1.00 / n/a
input / output · per 1M
Pricing by serving provider
standard tier · per 1M tokens · click a provider to expand| Serving provider | Input /1M | Output /1M | Endpoints | |||||||||||||
|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|---|
| DatabricksCheapest | $0.63 | $3.75 | 1 | |||||||||||||
| ||||||||||||||||
Price history
input + output $/1M since we started tracking Input Output
Cost calculator
estimate your monthly spend on this model$425
estimated / month
Model IDs
copy the exact identifier for your platformdatabricks/databricks-gemini-3-flash
Frequently asked questions
Gemini 3 Flash pricing, context and availabilityHow much does Gemini 3 Flash cost?
Gemini 3 Flash costs $0.63 per 1M input tokens and $3.75 per 1M output tokens at its cheapest provider via Databricks.
What is the context window of Gemini 3 Flash?
Gemini 3 Flash accepts up to 1M tokens of context and can return up to 66K output tokens.
What can Gemini 3 Flash do?
Gemini 3 Flash supports tool use and prompt caching.
Other Google models
compare pricing across the Google lineup Pricing verified against LiteLLM + OpenRouter + provider pages · refreshed daily
Every weekday
AI moves fast. Here's your debrief.
News, analysis, tools, and more.
For people who build with AI