Skip to content
AI Atlas

Terminal-Bench — cost vs performance

Best current row per canonical model in the group “accuracy · variant=hard · evaluator=Artificial Analysis” (317 models) against cheapest current input price (USD / 1M tokens). The dashed line is the Pareto frontier: no model is both better and cheaper than a point on it.

XOutput priceInput priceParametersContextMemory (est.)ScaleloglinearBubblecontextparamsnone
Group

open / restricted weights closed Pareto frontier (10 models)bubble = context window

Frontier models 10

Pareto frontier
Modelaccuracy on Terminal-BenchRankcheapest current input price (USD / 1M tokens)ProviderTrust
gpt-5.6-solOpenAI · Closed65.9%1$1OpenAI APIIndependent
gpt-5.4-miniOpenAI · Closed52.3%11$0.375OpenRouterIndependent
KAT-Coder-Pro V2Kwaipilot · Closed49.2%14$0.30OpenRouterIndependent
gpt-5.4-nanoOpenAI · Closed42.4%30$0.10OpenRouterIndependent
Gemma 4 31BGoogle · Open weights36.4%51$0.09OpenRouterIndependent
Nemotron 3 SuperNVIDIA · Open weights28.8%88$0.085NVIDIA NIM / build.nvidia.comIndependent
Gemma 4 26B A4BGoogle · Open weights25%100$0.042Google Gemini APIIndependent
gpt-oss-120bOpenAI · Open weights23.5%110$0.037OpenRouterIndependent
gpt-5-nanoOpenAI · Closed17.4%131$0.025OpenAI APIIndependent
Granite 4.0 MicroIBM · Open weights1.52%256$0.017OpenRouterIndependent

Methodology. Points are the best current row per canonical model in comparability group 'accuracy · variant=hard · evaluator=Artificial Analysis'. Price = cheapest current offer across providers (the provider shown). Pareto frontier maximises the score and minimises x; exact ties are all kept. memory_estimate is an estimate (see /methodology). Only the selected comparability group is plotted; points under other configurations are not mixed in. Price = cheapest current offer across providers at the time of the last crawl. Nothing is estimated except where marked.