IFBench benchmark

321 AI models ranked on IFBench. Each row shows the model’s score next to its cheapest API price per 1M tokens where it has one, so you can weigh quality against cost.

Top model today

Grok 4.3 (xAI) leads at 83%, from $1.25 per 1M input.

Data via arXiv:2507.02833

Models ranked

321

on this benchmark

Top score

83%

current leader

Type

Single eval

benchmark

Updated

29 Aug 2026

last source fetch

Cost vs IFBench

200 priced models
Most attractive quadrant (cheap + high)
= reasoning model (extended thinking)

Each labelled name links to that model’s page. Curated to the cost-efficiency frontier, top scorers and cheapest so labels stay readable; the full field is in the leaderboard table below (and via “select all”). Reasoning models (extended “thinking” before answering) carry a ring around the dot.

Models ranked on IFBench

highest score first · 321 models
Maker
#
1Grok 4.3xAI83%medium$1.25$2.50
2Grok 4.20 0309Spacexai83%n/an/a
3MiniMax M3MiniMax83%$0.23$0.96
4Nvidia Nemotron 3 Ultra 550B A55bNVIDIA81%$0.50$2.20
5Grok 4.20xAI81%$1.25$2.50
6Qwen3.7 MaxAlibaba81%$1.25$3.75
7Nemotron Cascade 2 30B A3BNVIDIA80%n/an/a
8MiMo-V2.5-ProXiaomi80%$0.43$0.87
9Nova 2.0 Pro PreviewAmazon80%lown/an/a
10DeepSeek V4 Flash 0423DeepSeek79%$0.09$0.17
11DeepSeek V4 Flash (max)DeepSeek79%n/an/a
12Qwen3.5 397B A17BAlibaba79%$0.39$2.34
13Qwen3.7 PlusAlibaba78%$0.32$1.28
14Gemini 3 FlashGoogle78%reasoning: true$0.63$3.75
15GPT-5.2-CodexOpenAI78%$1.75$14.00
16Gemini 3.1 Flash Lite PreviewGoogle77%$0.25$1.50
17Gemini 3.1 Pro PreviewGoogle77%$2.00$12.00
18Qwen3.6 Max PreviewAlibaba77%n/an/a
19DeepSeek V4 Pro 0423DeepSeek76%$0.43$0.87
20DeepSeek V4 Pro (max)DeepSeek76%n/an/a
21Gemini 3.5 FlashGoogle76%$1.50$9.00
22GLM 5.1Zhipu76%$1.05$3.50
23Kimi K2.6Moonshot76%$0.65$3.40
24GPT-5.4 NanoOpenAI76%$0.20$1.25
25Muse SparkMeta76%n/an/a
26GPT-5.5OpenAI76%$5.00$30.00
27MiniMax M2.7MiniMax76%$0.25$0.55
28Qwen3.5-122B-A10BAlibaba76%$0.25$1.75
29Qwen3.5-27BAlibaba76%$0.20$1.56
30Gemma 4 31BGoogle76%$0.14$0.40
31GPT-5 MiniOpenAI75%$0.25$2.00
32GPT-5.2OpenAI75%$1.75$14.00
33GPT-5.3-CodexOpenAI75%$1.75$14.00
34Qwen3.6 PlusAlibaba75%$0.33$1.95
35GPT 5 CodexOpenAI74%$1.25$10.00
36GPT-5.4OpenAI74%$2.50$15.00
37Command A+Cohere74%n/an/a
38Gemma 4 12BGoogle74%n/an/a
39GLM 5.2Zhipu73%$0.61$1.98
40GPT-5.4 MiniOpenAI73%$0.75$4.50
41GLM 5 TurboZhipu73%$1.20$4.00
42GPT-5OpenAI73%$1.25$10.00
43GPT-5.1OpenAI73%$1.25$10.00
44GPT-5.6 SolOpenAI73%$2.00$10.00
45Qwen3.5-35B-A3BAlibaba73%$0.14$1.00
46Gemma 4 26B A4bGoogle72%$0.13$0.40
47GLM 5Zhipu72%$0.60$1.92
48MiniMax M2MiniMax72%$0.26$1.02
49MiMo-V2-Flash (Feb 2026)Xiaomi72%n/an/a
50MiniMax M2.5MiniMax72%$0.27$1.08
150 of 321
Page 1 of 7
Prices are on-demand list rates. Enterprise commitments can be materially lower. Cost per task uses best-effort token usage.

Frequently asked questions

What is the IFBench benchmark?

IFBench (Ai2) tests how well models follow new, previously unseen verifiable output constraints such as word-count limits, required keywords or formatting rules that were not part of training. Scoring is fully programmatic: verifier functions check whether each response satisfies the constraint.

Which AI model scores highest on IFBench?

Grok 4.3 (xAI) leads with 83%, from $1.25 per 1M input tokens.

How many models are ranked on IFBench?

321 models carry a IFBench score. Scores come from Artificial Analysis; where a model has a live API price it is verified against LiteLLM and OpenRouter, and models with no current price are listed without one.

Other benchmarks

compare the same models on a different eval
Benchmark scores via Artificial Analysis · pricing verified against LiteLLM + OpenRouter · refreshed daily · updated 29 Aug 2026

Every weekday

AI moves fast. Here's your debrief.

News, analysis, tools, and more.

For people who build with AI