DeepInfra LLM API pricing

Every model DeepInfra serves, with its pricing per 1M tokens and the exact model ID to copy. Sorted cheapest input first. Verified against LiteLLM and OpenRouter daily.

Models served

135

on this provider

Cheapest input

$0.02

per 1M tokens

Cheapest output

$0.02

per 1M tokens

Max context

10.5M

tokens

Model ID
Mistral Nemo Instruct 2407Mistral131K$0.02$0.03deepinfra/mistralai/Mistral-Nemo-Instruct-2407
Meta Llama 3.1 8B Instruct TurboMeta131K$0.02$0.04deepinfra/meta-llama/Meta-Llama-3.1-8B-Instruct-Turbo
Llama 3.2 3B InstructMeta131K$0.02$0.02deepinfra/meta-llama/Llama-3.2-3B-Instruct
Gemma 4 E4b ItGoogle131K$0.02$0.10deepinfra/google/gemma-4-E4B-it
gpt-oss-20bOpenAI131K$0.03$0.14deepinfra/openai/gpt-oss-20b
Meta Llama 3.1 8B InstructMeta131K$0.03$0.05deepinfra/meta-llama/Meta-Llama-3.1-8B-Instruct
Meta Llama 3 8B InstructMeta8K$0.03$0.06deepinfra/meta-llama/Meta-Llama-3-8B-Instruct
gpt-oss-120bOpenAI131K$0.04$0.17deepinfra/openai/gpt-oss-120b
L3 8B Lunaris V1 TurboOther8K$0.04$0.05deepinfra/Sao10K/L3-8B-Lunaris-v1-Turbo
Nvidia Nemotron Nano 9BNVIDIA131K$0.04$0.16deepinfra/nvidia/NVIDIA-Nemotron-Nano-9B-v2
Qwen2 5 7B InstructAlibaba33K$0.04$0.10deepinfra/Qwen/Qwen2.5-7B-Instruct
Llama 3.2 11B Vision InstructMeta131K$0.05$0.05deepinfra/meta-llama/Llama-3.2-11B-Vision-Instruct
Gemma 3 12BGoogle131K$0.05$0.15deepinfra/google/gemma-3-12b-it
Gemma 3 4BGoogle131K$0.05$0.10deepinfra/google/gemma-3-4b-it
Mistral Small 3Mistral33K$0.05$0.08deepinfra/mistralai/Mistral-Small-24B-Instruct-2501
Nemotron 3 Nano 30B A3BNVIDIA262K$0.05$0.20deepinfra/nvidia/Nemotron-3-Nano-30B-A3B
Llama Guard 3 8BMeta131K$0.06$0.06deepinfra/meta-llama/Llama-Guard-3-8B
GLM 4.7 FlashZhipu203K$0.06$0.40deepinfra/zai-org/GLM-4.7-Flash
Ling-3.0-flashOther262K$0.06$0.18deepinfra/inclusionAI/Ling-3.0-flash
Phi 4Microsoft16K$0.07$0.14deepinfra/microsoft/phi-4
Gemma 4 26B A4B Google262K$0.07$0.34deepinfra/google/gemma-4-26B-A4B-it
Mistral Small 3.2 24B Instruct 2506Mistral128K$0.07$0.20deepinfra/mistralai/Mistral-Small-3.2-24B-Instruct-2506
Nvidia Nemotron 3.5 LightningNVIDIA262K$0.08$0.20deepinfra/nvidia/NVIDIA-Nemotron-3.5-Lightning
Qwen3 32BAlibaba131K$0.08$0.28deepinfra/Qwen/Qwen3-32B
Gemma 3 27BGoogle131K$0.08$0.16deepinfra/google/gemma-3-27b-it
DeepSeek V4 Flash 0731DeepSeek1M$0.08$0.18deepinfra/deepseek-ai/DeepSeek-V4-Flash-0731
Nvidia Nemotron 3 Super 120B A12bNVIDIA262K$0.09$0.40deepinfra/nvidia/NVIDIA-Nemotron-3-Super-120B-A12B
Qwen3 235B A22b Instruct 2507Alibaba262K$0.09$0.55deepinfra/Qwen/Qwen3-235B-A22B-Instruct-2507
Qwen3 Next 80B A3B InstructAlibaba262K$0.09$1.10deepinfra/Qwen/Qwen3-Next-80B-A3B-Instruct
Gemma 4 31B It TurboGoogle262K$0.09$0.34deepinfra/google/gemma-4-31B-it-turbo
DeepSeek V4 Flash 0423DeepSeek1M$0.09$0.18deepinfra/deepseek-ai/DeepSeek-V4-Flash
Llama 3.3 70B Instruct TurboMeta131K$0.10$0.32deepinfra/meta-llama/Llama-3.3-70B-Instruct-Turbo
Llama 4 Scout 17B 16e InstructMeta10.5M$0.10$0.30deepinfra/meta-llama/Llama-4-Scout-17B-16E-Instruct
Llama 3.3 Nemotron Super 49B V1 5NVIDIA131K$0.10$0.40deepinfra/nvidia/Llama-3.3-Nemotron-Super-49B-v1.5
Gemini 2.0 Flash 001Google1M$0.10$0.40deepinfra/google/gemini-2.0-flash-001
Qwen3.6 35B A3BAlibaba262K$0.10$0.95deepinfra/Qwen/Qwen3.6-35B-A3B
Seed-2.0-MiniOther256K$0.10$0.40deepinfra/ByteDance/Seed-2.0-mini
Qwen3.5-9BAlibaba262K$0.10$0.15deepinfra/Qwen/Qwen3.5-9B
Qwen3 14BAlibaba41K$0.12$0.24deepinfra/Qwen/Qwen3-14B
Qwen3 30B A3BAlibaba131K$0.12$0.50deepinfra/Qwen/Qwen3-30B-A3B
Gemma 4 31BGoogle262K$0.13$0.38deepinfra/google/gemma-4-31B-it
Qwen3 Next 80B A3B ThinkingAlibaba262K$0.14$1.40deepinfra/Qwen/Qwen3-Next-80B-A3B-Thinking
Qwen3.5-35B-A3BAlibaba262K$0.14$1.00deepinfra/Qwen/Qwen3.5-35B-A3B
Hy3Other262K$0.14$0.58deepinfra/tencent/Hy3
QwQ 32BAlibaba131K$0.15$0.40deepinfra/Qwen/QwQ-32B
GPT OSS 120B TurboOpenAI131K$0.15$0.60deepinfra/openai/gpt-oss-120b-Turbo
Qwen3 VL 30B A3B InstructAlibaba262K$0.15$0.60deepinfra/Qwen/Qwen3-VL-30B-A3B-Instruct
Llama Guard 4 12BMeta1M$0.18$0.18deepinfra/meta-llama/Llama-Guard-4-12B
Qwen3 235B A22BAlibaba262K$0.18$0.54deepinfra/Qwen/Qwen3-235B-A22B
Llama 4 Maverick 17B 128e Instruct FP8Meta1M$0.20$0.80deepinfra/meta-llama/Llama-4-Maverick-17B-128E-Instruct-FP8
150 of 135
Page 1 of 3
Pricing verified against LiteLLM + OpenRouter + provider pages · refreshed daily

Every weekday

AI moves fast. Here's your debrief.

News, analysis, tools, and more.

For people who build with AI