Skip to content
AI Atlas

IFBench — cost vs performance

Best current row per canonical model in the group “accuracy · evaluator=Artificial Analysis” (334 models) against total parameters. The dashed line is the Pareto frontier: no model is both better and cheaper than a point on it.

open / restricted weights closed Pareto frontier (9 models)bubble = context window

Frontier models 9

Pareto frontier
Modelaccuracy on IFBenchRanktotal parametersProviderTrust
MiniMax-M3MiniMax · Open weights82.9%3427BIndependent
MiniMax M2.7MiniMax · Open weights75.7%26228.7BIndependent
Gemma 4 31BGoogle · Open weights75.6%2831.3BIndependent
Gemma 4 26B A4BGoogle · Open weights72.5%4625.8BIndependent
Qwen3.5-9BQwen · Open weights66.7%819.65BIndependent
LFM2.5-8B-A1BLiquid AI · Open weights55.6%1238.47BIndependent
Qwen3.5-4BQwen · Open weights52.0%1414.66BIndependent
LFM2.5-1.2B-InstructLiquid AI · Open weights43.8%1751.17BIndependent
Qwen3-0.6BQwen · Open weights23.3%312751.6MIndependent

Methodology. Points are the best current row per canonical model in comparability group 'accuracy · evaluator=Artificial Analysis'. Price = cheapest current offer across providers (the provider shown). Pareto frontier maximises the score and minimises x; exact ties are all kept. memory_estimate is an estimate (see /methodology). Only the selected comparability group is plotted; points under other configurations are not mixed in. Price = cheapest current offer across providers at the time of the last crawl. Nothing is estimated except where marked.