DeepInfra LLM API pricing
Every model DeepInfra serves, with its pricing per 1M tokens and the exact model ID to copy. Sorted cheapest input first. Verified against LiteLLM and OpenRouter daily.
Models served
135
on this provider
Cheapest input
$0.02
per 1M tokens
Cheapest output
$0.02
per 1M tokens
Max context
10.5M
tokens
| Model ID | |||||
|---|---|---|---|---|---|
| Mistral Nemo Instruct 2407 | Mistral | 131K | $0.02 | $0.03 | deepinfra/mistralai/Mistral-Nemo-Instruct-2407 |
| Meta Llama 3.1 8B Instruct Turbo | Meta | 131K | $0.02 | $0.04 | deepinfra/meta-llama/Meta-Llama-3.1-8B-Instruct-Turbo |
| Llama 3.2 3B Instruct | Meta | 131K | $0.02 | $0.02 | deepinfra/meta-llama/Llama-3.2-3B-Instruct |
| Gemma 4 E4b It | 131K | $0.02 | $0.10 | deepinfra/google/gemma-4-E4B-it | |
| gpt-oss-20b | OpenAI | 131K | $0.03 | $0.14 | deepinfra/openai/gpt-oss-20b |
| Meta Llama 3.1 8B Instruct | Meta | 131K | $0.03 | $0.05 | deepinfra/meta-llama/Meta-Llama-3.1-8B-Instruct |
| Meta Llama 3 8B Instruct | Meta | 8K | $0.03 | $0.06 | deepinfra/meta-llama/Meta-Llama-3-8B-Instruct |
| gpt-oss-120b | OpenAI | 131K | $0.04 | $0.17 | deepinfra/openai/gpt-oss-120b |
| L3 8B Lunaris V1 Turbo | Other | 8K | $0.04 | $0.05 | deepinfra/Sao10K/L3-8B-Lunaris-v1-Turbo |
| Nvidia Nemotron Nano 9B | NVIDIA | 131K | $0.04 | $0.16 | deepinfra/nvidia/NVIDIA-Nemotron-Nano-9B-v2 |
| Qwen2 5 7B Instruct | Alibaba | 33K | $0.04 | $0.10 | deepinfra/Qwen/Qwen2.5-7B-Instruct |
| Llama 3.2 11B Vision Instruct | Meta | 131K | $0.05 | $0.05 | deepinfra/meta-llama/Llama-3.2-11B-Vision-Instruct |
| Gemma 3 12B | 131K | $0.05 | $0.15 | deepinfra/google/gemma-3-12b-it | |
| Gemma 3 4B | 131K | $0.05 | $0.10 | deepinfra/google/gemma-3-4b-it | |
| Mistral Small 3 | Mistral | 33K | $0.05 | $0.08 | deepinfra/mistralai/Mistral-Small-24B-Instruct-2501 |
| Nemotron 3 Nano 30B A3B | NVIDIA | 262K | $0.05 | $0.20 | deepinfra/nvidia/Nemotron-3-Nano-30B-A3B |
| Llama Guard 3 8B | Meta | 131K | $0.06 | $0.06 | deepinfra/meta-llama/Llama-Guard-3-8B |
| GLM 4.7 Flash | Zhipu | 203K | $0.06 | $0.40 | deepinfra/zai-org/GLM-4.7-Flash |
| Ling-3.0-flash | Other | 262K | $0.06 | $0.18 | deepinfra/inclusionAI/Ling-3.0-flash |
| Phi 4 | Microsoft | 16K | $0.07 | $0.14 | deepinfra/microsoft/phi-4 |
| Gemma 4 26B A4B | 262K | $0.07 | $0.34 | deepinfra/google/gemma-4-26B-A4B-it | |
| Mistral Small 3.2 24B Instruct 2506 | Mistral | 128K | $0.07 | $0.20 | deepinfra/mistralai/Mistral-Small-3.2-24B-Instruct-2506 |
| Nvidia Nemotron 3.5 Lightning | NVIDIA | 262K | $0.08 | $0.20 | deepinfra/nvidia/NVIDIA-Nemotron-3.5-Lightning |
| Qwen3 32B | Alibaba | 131K | $0.08 | $0.28 | deepinfra/Qwen/Qwen3-32B |
| Gemma 3 27B | 131K | $0.08 | $0.16 | deepinfra/google/gemma-3-27b-it | |
| DeepSeek V4 Flash 0731 | DeepSeek | 1M | $0.08 | $0.18 | deepinfra/deepseek-ai/DeepSeek-V4-Flash-0731 |
| Nvidia Nemotron 3 Super 120B A12b | NVIDIA | 262K | $0.09 | $0.40 | deepinfra/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B |
| Qwen3 235B A22b Instruct 2507 | Alibaba | 262K | $0.09 | $0.55 | deepinfra/Qwen/Qwen3-235B-A22B-Instruct-2507 |
| Qwen3 Next 80B A3B Instruct | Alibaba | 262K | $0.09 | $1.10 | deepinfra/Qwen/Qwen3-Next-80B-A3B-Instruct |
| Gemma 4 31B It Turbo | 262K | $0.09 | $0.34 | deepinfra/google/gemma-4-31B-it-turbo | |
| DeepSeek V4 Flash 0423 | DeepSeek | 1M | $0.09 | $0.18 | deepinfra/deepseek-ai/DeepSeek-V4-Flash |
| Llama 3.3 70B Instruct Turbo | Meta | 131K | $0.10 | $0.32 | deepinfra/meta-llama/Llama-3.3-70B-Instruct-Turbo |
| Llama 4 Scout 17B 16e Instruct | Meta | 10.5M | $0.10 | $0.30 | deepinfra/meta-llama/Llama-4-Scout-17B-16E-Instruct |
| Llama 3.3 Nemotron Super 49B V1 5 | NVIDIA | 131K | $0.10 | $0.40 | deepinfra/nvidia/Llama-3.3-Nemotron-Super-49B-v1.5 |
| Gemini 2.0 Flash 001 | 1M | $0.10 | $0.40 | deepinfra/google/gemini-2.0-flash-001 | |
| Qwen3.6 35B A3B | Alibaba | 262K | $0.10 | $0.95 | deepinfra/Qwen/Qwen3.6-35B-A3B |
| Seed-2.0-Mini | Other | 256K | $0.10 | $0.40 | deepinfra/ByteDance/Seed-2.0-mini |
| Qwen3.5-9B | Alibaba | 262K | $0.10 | $0.15 | deepinfra/Qwen/Qwen3.5-9B |
| Qwen3 14B | Alibaba | 41K | $0.12 | $0.24 | deepinfra/Qwen/Qwen3-14B |
| Qwen3 30B A3B | Alibaba | 131K | $0.12 | $0.50 | deepinfra/Qwen/Qwen3-30B-A3B |
| Gemma 4 31B | 262K | $0.13 | $0.38 | deepinfra/google/gemma-4-31B-it | |
| Qwen3 Next 80B A3B Thinking | Alibaba | 262K | $0.14 | $1.40 | deepinfra/Qwen/Qwen3-Next-80B-A3B-Thinking |
| Qwen3.5-35B-A3B | Alibaba | 262K | $0.14 | $1.00 | deepinfra/Qwen/Qwen3.5-35B-A3B |
| Hy3 | Other | 262K | $0.14 | $0.58 | deepinfra/tencent/Hy3 |
| QwQ 32B | Alibaba | 131K | $0.15 | $0.40 | deepinfra/Qwen/QwQ-32B |
| GPT OSS 120B Turbo | OpenAI | 131K | $0.15 | $0.60 | deepinfra/openai/gpt-oss-120b-Turbo |
| Qwen3 VL 30B A3B Instruct | Alibaba | 262K | $0.15 | $0.60 | deepinfra/Qwen/Qwen3-VL-30B-A3B-Instruct |
| Llama Guard 4 12B | Meta | 1M | $0.18 | $0.18 | deepinfra/meta-llama/Llama-Guard-4-12B |
| Qwen3 235B A22B | Alibaba | 262K | $0.18 | $0.54 | deepinfra/Qwen/Qwen3-235B-A22B |
| Llama 4 Maverick 17B 128e Instruct FP8 | Meta | 1M | $0.20 | $0.80 | deepinfra/meta-llama/Llama-4-Maverick-17B-128E-Instruct-FP8 |
1–50 of 135
Page 1 of 3
Pricing verified against LiteLLM + OpenRouter + provider pages · refreshed daily
Every weekday
AI moves fast. Here's your debrief.
News, analysis, tools, and more.
For people who build with AI