Skip to content
AI Atlas
BenchmarkActivecategory · instruction-followingfamily · ifeval · variant IFBench

IFBench

precise instruction following with novel constraints

data quality51

Updated 4 h ago · first seen 11 Sept 2026

Metric
accuracy · %
Current results
898
Models
334
Current leader
Grok 4.3 83.3%

Score history · gpt-5.6-sol 10 rows

Score history for gpt-5.6-sol0%20%40%60%80%Sept 26Sept 26Sept 26Sept 26Sept 26Sept 26Sept 26Sept 26Sept 26
  • gpt-5.6-sol
  • 66.53%aa_slug=gpt-5-6-sol-low · evaluator=Artificial Analysis · index_version=4.3 · reasoning_effort=low12 Sept 2026
  • 69.59%aa_slug=gpt-5-6-sol-medium · evaluator=Artificial Analysis · index_version=4.3 · reasoning_effort=medium12 Sept 2026
  • 69.18%aa_slug=gpt-5-6-sol-high · evaluator=Artificial Analysis · index_version=4.3 · reasoning_effort=high12 Sept 2026
  • 72.65%aa_slug=gpt-5-6-sol · evaluator=Artificial Analysis · index_version=4.3 · reasoning_effort=max12 Sept 2026
  • 71.02%aa_slug=gpt-5-6-sol-xhigh · evaluator=Artificial Analysis · index_version=4.3 · reasoning_effort=xhigh12 Sept 2026
  • 66.53%aa_slug=gpt-5-6-sol-low · evaluator=Artificial Analysis · index_version=4.3 · aa_variant_slug=gpt-5-6-sol-low11 Sept 2026
  • 69.59%aa_slug=gpt-5-6-sol-medium · evaluator=Artificial Analysis · index_version=4.3 · aa_variant_slug=gpt-5-6-sol-medium11 Sept 2026
  • 69.18%aa_slug=gpt-5-6-sol-high · evaluator=Artificial Analysis · index_version=4.3 · aa_variant_slug=gpt-5-6-sol-high11 Sept 2026
  • 72.65%aa_slug=gpt-5-6-sol · evaluator=Artificial Analysis · index_version=4.311 Sept 2026
  • 71.02%aa_slug=gpt-5-6-sol-xhigh · evaluator=Artificial Analysis · index_version=4.3 · aa_variant_slug=gpt-5-6-sol-xhigh11 Sept 2026

Back to the leaderboard

Frontier over time · accuracy · evaluator=Artificial Analysis

6 leader changes recorded, all dated 11 Sept 2026 — the frontier line needs at least two distinct dates. The corpus is young: every result was first observed on the same day, so leader changes will separate in time as sources are re-crawled.

  1. 83.3%Grok 4.3 xAI Independent11 Sept 2026
  2. 81.3%Grok 4.3 xAI Independent11 Sept 2026
  3. 81.2%Grok 4.20 xAI Independent11 Sept 2026
  4. 81.0%Grok 4.3 xAI Independent11 Sept 2026
  5. 73.5%gemma-4-12B Google Independent11 Sept 2026
  6. 45.9%grok-3-mini-reasoning SpaceXAI Independent11 Sept 2026

Includes closed rows (history). A point is emitted whenever a result beats every earlier result of the same group, ordered by evaluated_at when the source publishes it, else observed_at.

Leaderboard 334 models · trust independent-evaluator

Select models with +, then Compare.

Leaderboard
#ModelScoreTrustConfigurationvs leaderEvaluatedSourceActions
201Gemini 2.0 FlashClosedGoogle · Gemini 2.0 · best of 2 rows40.2%IndependentreasoningoffPartially comparable0.00 ptobs. 12 Sept 2026artificialanalysis.aiT2History
202Qwen3 VL 8B InstructOpen weightsQwen · Qwen3 · best of 4 rows39.9%IndependentreasoningonPartially comparable-0.34 ptobs. 12 Sept 2026artificialanalysis.aiT2History
203Mistral Medium 3.1ClosedMistral AI · Mistral · best of 2 rows39.8%IndependentreasoningoffPartially comparable-0.40 ptobs. 12 Sept 2026artificialanalysis.aiT2History
204deepseek-r1Open weightsDeepSeek · DeepSeek-R1 · best of 2 rows39.6%IndependentreasoningonPartially comparable-0.61 ptobs. 12 Sept 2026artificialanalysis.aiT2History
205Llama 4 ScoutOpen weightsMeta AI · Llama 4 · best of 2 rows39.5%IndependentreasoningoffPartially comparable-0.68 ptobs. 12 Sept 2026artificialanalysis.aiT2History
206llama-3-3-nemotron-super-49bOpen weightsNVIDIA · Llama 3.3 · best of 4 rows39.5%IndependentreasoningoffPartially comparable-0.74 ptobs. 12 Sept 2026artificialanalysis.aiT2History
207Mistral Medium 3ClosedMistral AI · Mistral · best of 2 rows39.3%IndependentreasoningoffPartially comparable-0.95 ptobs. 12 Sept 2026artificialanalysis.aiT2History
208ernie-4-5-300b-a47bOpen weightsBaidu · ERNIE 4.5 · best of 2 rows39.1%IndependentreasoningoffPartially comparable-1.08 ptobs. 12 Sept 2026artificialanalysis.aiT2History
209Llama-3.1-405BRestricted weightsMeta AI · Llama 3.1 · best of 2 rows39.0%IndependentreasoningoffPartially comparable-1.15 ptobs. 12 Sept 2026artificialanalysis.aiT2History
210deepseek-r1-0120Open weightsDeepSeek · DeepSeek · best of 2 rows39.0%IndependentreasoningonPartially comparable-1.22 ptobs. 12 Sept 2026artificialanalysis.aiT2History
211QwQ-32BOpen weightsAlibaba Group · Qwen · best of 2 rows38.8%IndependentreasoningonPartially comparable-1.42 ptobs. 12 Sept 2026artificialanalysis.aiT2History
212qwen3-235b-a22b-instructOpen weightsAlibaba Group · Qwen3 · best of 4 rows38.7%IndependentreasoningonPartially comparable-1.49 ptobs. 12 Sept 2026artificialanalysis.aiT2History
213granite-4.1-8bOpen weightsIBM · Granite 4.1 · best of 2 rows38.6%IndependentreasoningoffPartially comparable-1.56 ptobs. 12 Sept 2026artificialanalysis.aiT2History
214gpt-4.1-miniClosedOpenAI · GPT 4.1 · best of 2 rows38.3%IndependentreasoningoffPartially comparable-1.90 ptobs. 12 Sept 2026artificialanalysis.aiT2History
215llama-3-1-nemotron-ultra-253b-v1-reasoningOpen weightsNVIDIA · Llama 3.1 · best of 2 rows38.2%IndependentreasoningonPartially comparable-2.04 ptobs. 12 Sept 2026artificialanalysis.aiT2History
216Devstral 2Open weightsMistral AI · Devstral 2 · best of 2 rows38.1%IndependentreasoningoffPartially comparable-2.10 ptobs. 12 Sept 2026artificialanalysis.aiT2History
216nova-proClosedAmazon Web Services · Nova · best of 2 rows38.1%IndependentreasoningoffPartially comparable-2.10 ptobs. 12 Sept 2026artificialanalysis.aiT2History
216olmo-2-32bOpen weightsAllen Institute for AI · OLMo 2 · best of 2 rows38.1%IndependentreasoningoffPartially comparable-2.10 ptobs. 12 Sept 2026artificialanalysis.aiT2History
219gemma-4-E2BOpen weightsGoogle · Gemma 4 · best of 4 rows38.0%IndependentreasoningonPartially comparable-2.17 ptobs. 12 Sept 2026artificialanalysis.aiT2History
220qwen3-5-omni-flashClosedAlibaba Group · Qwen3.5 · best of 2 rows38.0%IndependentreasoningoffPartially comparable-2.24 ptobs. 12 Sept 2026artificialanalysis.aiT2History
221hyperclova-x-seed-think-32bOpen weightsNaver · Seed · best of 2 rows37.9%IndependentreasoningonPartially comparable-2.31 ptobs. 12 Sept 2026artificialanalysis.aiT2History
222GLM 4.5 AirOpen weightsZ.ai (Zhipu AI) · GLM4.5 · best of 2 rows37.5%IndependentreasoningonPartially comparable-2.65 ptobs. 12 Sept 2026artificialanalysis.aiT2History
223llama-3-instruct-70bOpen weightsMeta AI · Llama 3 · best of 2 rows37.1%IndependentreasoningoffPartially comparable-3.13 ptobs. 12 Sept 2026artificialanalysis.aiT2History
223solar-pro-2ClosedUpstage · Solar · best of 4 rows37.1%IndependentreasoningonPartially comparable-3.13 ptobs. 12 Sept 2026artificialanalysis.aiT2History
225llama-nemotron-super-49b-v1-5Open weightsNVIDIA · Llama · best of 4 rows37.0%IndependentreasoningonPartially comparable-3.19 ptobs. 12 Sept 2026artificialanalysis.aiT2History
226Qwen2.5 72B InstructOpen weightsQwen · Qwen2.5 · best of 2 rows36.9%IndependentreasoningoffPartially comparable-3.33 ptobs. 12 Sept 2026artificialanalysis.aiT2History
227Gemma 3 12BOpen weightsGoogle · Gemma 3 · best of 2 rows36.7%IndependentreasoningoffPartially comparable-3.47 ptobs. 12 Sept 2026artificialanalysis.aiT2History
228jt-miniClosedChina Mobile · best of 2 rows36.7%IndependentreasoningoffPartially comparable-3.53 ptobs. 12 Sept 2026artificialanalysis.aiT2History
229Qwen3-VL-4B-InstructOpen weightsQwen · Qwen3 · best of 4 rows36.6%IndependentreasoningonPartially comparable-3.60 ptobs. 12 Sept 2026artificialanalysis.aiT2History
230Command AOpen weightsCohere · Command · best of 2 rows36.5%IndependentreasoningoffPartially comparable-3.74 ptobs. 12 Sept 2026artificialanalysis.aiT2History
231Qwen3 32BOpen weightsQwen · Qwen336.3%Independentgroup defaultsPartially comparable-3.87 ptobs. 11 Sept 2026artificialanalysis.aiT2History
231exaone-4-0-32bOpen weightsLG AI Research · EXAONE 4.0 · best of 4 rows36.3%IndependentreasoningonPartially comparable-3.87 ptobs. 12 Sept 2026artificialanalysis.aiT2History
231qwen3-32b-instructOpen weightsAlibaba Group · Qwen3 · best of 3 rows36.3%IndependentreasoningonPartially comparable-3.87 ptobs. 12 Sept 2026artificialanalysis.aiT2History
234Mistral Large 3Open weightsMistral AI · Mistral · best of 2 rows36.2%IndependentreasoningoffPartially comparable-4.01 ptobs. 12 Sept 2026artificialanalysis.aiT2History
234nova-premierClosedAmazon Web Services · Nova · best of 2 rows36.2%IndependentreasoningoffPartially comparable-4.01 ptobs. 12 Sept 2026artificialanalysis.aiT2History
236Claude 3 HaikuClosedAnthropic · Claude · best of 2 rows36.1%IndependentreasoningoffPartially comparable-4.08 ptobs. 12 Sept 2026artificialanalysis.aiT2History
237GPT-4o (2024-08-06)ClosedOpenAI · GPT 4 · best of 2 rows36.0%IndependentreasoningoffPartially comparable-4.21 ptobs. 12 Sept 2026artificialanalysis.aiT2History
238nanbeige4-1-3bOpen weightsNanbeige · best of 2 rows35.4%IndependentreasoningonPartially comparable-4.76 ptobs. 12 Sept 2026artificialanalysis.aiT2History
239Qwen3 Coder NextOpen weightsQwen · Qwen3 · best of 2 rows35.2%IndependentreasoningoffPartially comparable-4.96 ptobs. 12 Sept 2026artificialanalysis.aiT2History
240jamba-1-7-largeOpen weightsAI21 Labs · Jamba 1.7 · best of 2 rows35.2%IndependentreasoningoffPartially comparable-5.03 ptobs. 12 Sept 2026artificialanalysis.aiT2History
241hermes-4-llama-3-1-405bOpen weightsNous Research · Llama 3.1 · best of 3 rows34.8%IndependentreasoningoffPartially comparable-5.44 ptobs. 12 Sept 2026artificialanalysis.aiT2History
241ling-1tOpen weightsinclusionAI · best of 2 rows34.8%IndependentreasoningoffPartially comparable-5.44 ptobs. 12 Sept 2026artificialanalysis.aiT2History
243devstral-smallOpen weightsMistral AI · Devstral · best of 2 rows34.6%IndependentreasoningoffPartially comparable-5.64 ptobs. 12 Sept 2026artificialanalysis.aiT2History
244Pixtral LargeOpen weightsMistral AI · Pixtral · best of 2 rows34.5%IndependentreasoningoffPartially comparable-5.71 ptobs. 12 Sept 2026artificialanalysis.aiT2History
245Llama-3.1-70BOpen weightsMeta AI · Llama 3.1 · best of 2 rows34.4%IndependentreasoningoffPartially comparable-5.78 ptobs. 12 Sept 2026artificialanalysis.aiT2History
246ling-flash-2-0Open weightsinclusionAI · best of 2 rows34.4%IndependentreasoningoffPartially comparable-5.85 ptobs. 12 Sept 2026artificialanalysis.aiT2History
246sarvam-105bOpen weightsSarvam · best of 2 rows34.4%Independentreasoning_efforthighPartially comparable-5.85 ptobs. 12 Sept 2026artificialanalysis.aiT2History
248gpt-4oClosedOpenAI · GPT 4 · best of 2 rows34.3%IndependentreasoningoffPartially comparable-5.91 ptobs. 12 Sept 2026artificialanalysis.aiT2History
249GLM 4.5VOpen weightsZ.ai (Zhipu AI) · GLM4.5 · best of 4 rows34.2%IndependentreasoningonPartially comparable-5.98 ptobs. 12 Sept 2026artificialanalysis.aiT2History
250nova-liteClosedAmazon Web Services · Nova · best of 2 rows34.1%IndependentreasoningoffPartially comparable-6.05 ptobs. 12 Sept 2026artificialanalysis.aiT2History
251intellect-3Open weightsPrime Intellect · best of 2 rows34.0%IndependentreasoningonPartially comparable-6.19 ptobs. 12 Sept 2026artificialanalysis.aiT2History
252granite-4.1-3bOpen weightsIBM · Granite 4.1 · best of 2 rows33.7%IndependentreasoningoffPartially comparable-6.53 ptobs. 12 Sept 2026artificialanalysis.aiT2History
253Mistral Small 3.2Open weightsMistral AI · Mistral · best of 2 rows33.5%IndependentreasoningoffPartially comparable-6.73 ptobs. 12 Sept 2026artificialanalysis.aiT2History
253Qwen3 8BOpen weightsQwen · Qwen333.5%Independentgroup defaultsPartially comparable-6.73 ptobs. 11 Sept 2026artificialanalysis.aiT2History
253qwen3-8b-instructOpen weightsAlibaba Group · Qwen3 · best of 3 rows33.5%IndependentreasoningonPartially comparable-6.73 ptobs. 12 Sept 2026artificialanalysis.aiT2History
256gpt-4ClosedOpenAI · GPT 4 · best of 2 rows33.2%IndependentreasoningoffPartially comparable-7.00 ptobs. 12 Sept 2026artificialanalysis.aiT2History
257LFM2.5-VL-1.6BOpen weightsLiquid AI · LFM2.5 · best of 2 rows33.1%IndependentreasoningoffPartially comparable-7.07 ptobs. 12 Sept 2026artificialanalysis.aiT2History
258Olmo-3-7B-InstructOpen weightsAllen Institute for AI · OLMo 3 · best of 2 rows32.8%IndependentreasoningoffPartially comparable-7.41 ptobs. 12 Sept 2026artificialanalysis.aiT2History
259Hermes 4 405BOpen weightsNous Research · Hermes 432.7%Independentgroup defaultsPartially comparable-7.48 ptobs. 11 Sept 2026artificialanalysis.aiT2History
260Qwen3 Coder 30B A3B InstructOpen weightsQwen · Qwen3 · best of 2 rows32.6%IndependentreasoningoffPartially comparable-7.55 ptobs. 12 Sept 2026artificialanalysis.aiT2History
261Qwen3-4BOpen weightsQwen · Qwen332.5%Independentgroup defaultsPartially comparable-7.68 ptobs. 11 Sept 2026artificialanalysis.aiT2History
261qwen3-4b-instructOpen weightsAlibaba Group · Qwen332.5%IndependentreasoningonPartially comparable-7.68 ptobs. 12 Sept 2026artificialanalysis.aiT2History
263Ministral 3 14BOpen weightsMistral AI · Ministral 3 · best of 2 rows32.0%IndependentreasoningoffPartially comparable-8.16 ptobs. 12 Sept 2026artificialanalysis.aiT2History
263gpt-4.1-nanoClosedOpenAI · GPT 4.1 · best of 2 rows32.0%IndependentreasoningoffPartially comparable-8.16 ptobs. 12 Sept 2026artificialanalysis.aiT2History
265nvidia-nemotron-nano-12b-v2-vlOpen weightsNVIDIA · Nemotron · best of 4 rows31.9%IndependentreasoningonPartially comparable-8.30 ptobs. 12 Sept 2026artificialanalysis.aiT2History
266Gemma 3 27BOpen weightsGoogle · Gemma 3 · best of 2 rows31.8%IndependentreasoningoffPartially comparable-8.36 ptobs. 12 Sept 2026artificialanalysis.aiT2History
266sarvam-m-reasoningOpen weightsSarvam · best of 2 rows31.8%IndependentreasoningonPartially comparable-8.36 ptobs. 12 Sept 2026artificialanalysis.aiT2History
268Devstral Small 1.0Open weightsMistral AI · Devstral · best of 2 rows31.6%IndependentreasoningoffPartially comparable-8.57 ptobs. 12 Sept 2026artificialanalysis.aiT2History
268Mistral Large 2.0Open weightsMistral AI · Mistral · best of 2 rows31.6%IndependentreasoningoffPartially comparable-8.57 ptobs. 12 Sept 2026artificialanalysis.aiT2History
270Qwen3.5-2BOpen weightsQwen · Qwen3.5 · best of 4 rows31.5%IndependentreasoningonPartially comparable-8.70 ptobs. 12 Sept 2026artificialanalysis.aiT2History
270granite-4-0-h-smallOpen weightsIBM · Granite 4.0 · best of 2 rows31.5%IndependentreasoningoffPartially comparable-8.70 ptobs. 12 Sept 2026artificialanalysis.aiT2History
272jamba-1-7-miniOpen weightsAI21 Labs · Jamba 1.7 · best of 2 rows31.4%IndependentreasoningoffPartially comparable-8.84 ptobs. 12 Sept 2026artificialanalysis.aiT2History
273Hermes-4-70BRestricted weightsNous Research · Hermes 431.3%Independentgroup defaultsPartially comparable-8.91 ptobs. 11 Sept 2026artificialanalysis.aiT2History
273hermes-4-llama-3-1-70bOpen weightsNous Research · Llama 3.1 · best of 3 rows31.3%IndependentreasoningonPartially comparable-8.91 ptobs. 12 Sept 2026artificialanalysis.aiT2History
275mistral-large-2Open weightsMistral AI · Mistral · best of 2 rows31.2%IndependentreasoningoffPartially comparable-8.98 ptobs. 12 Sept 2026artificialanalysis.aiT2History
276Devstral Small 2Open weightsMistral AI · Devstral · best of 2 rows31.2%IndependentreasoningoffPartially comparable-9.04 ptobs. 12 Sept 2026artificialanalysis.aiT2History
277gpt-4o-miniClosedOpenAI · GPT 4 · best of 2 rows30.9%IndependentreasoningoffPartially comparable-9.25 ptobs. 12 Sept 2026artificialanalysis.aiT2History
278llama-3-1-nemotron-instruct-70bOpen weightsNVIDIA · Llama 3.1 · best of 2 rows30.8%IndependentreasoningoffPartially comparable-9.45 ptobs. 12 Sept 2026artificialanalysis.aiT2History
279Reka Flash 3Open weightsrekaai · best of 2 rows30.4%IndependentreasoningonPartially comparable-9.79 ptobs. 12 Sept 2026artificialanalysis.aiT2History
279llama-3-2-instruct-11b-visionOpen weightsMeta AI · Llama 3.2 · best of 2 rows30.4%IndependentreasoningoffPartially comparable-9.79 ptobs. 12 Sept 2026artificialanalysis.aiT2History
281GLM 4.6VOpen weightsZ.ai (Zhipu AI) · GLM4.6 · best of 4 rows30.1%IndependentreasoningonPartially comparable-10.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
282Mistral Small 3.1Open weightsMistral AI · Mistral · best of 2 rows29.9%IndependentreasoningoffPartially comparable-10.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
282devstral-mediumClosedMistral AI · Devstral · best of 2 rows29.9%IndependentreasoningoffPartially comparable-10.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
284nova-microClosedAmazon Web Services · Nova · best of 2 rows29.4%IndependentreasoningoffPartially comparable-10.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
285Ministral 3 8BOpen weightsMistral AI · Ministral 3 · best of 2 rows29.1%IndependentreasoningoffPartially comparable-11.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
286Llama 3.1 8BRestricted weightsMeta AI · Llama 3.1 · best of 2 rows28.6%IndependentreasoningoffPartially comparable-11.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
287Gemma 3 4BRestricted weightsGoogle · Gemma 3 · best of 2 rows28.3%IndependentreasoningoffPartially comparable-11.9 ptobs. 12 Sept 2026artificialanalysis.aiT2History
288Kimi-Linear-48B-A3B-InstructOpen weightsMoonshot AI · Kimi · best of 2 rows28.1%IndependentreasoningoffPartially comparable-12.1 ptobs. 12 Sept 2026artificialanalysis.aiT2History
289gemma-3n-e4bOpen weightsGoogle · Gemma 3 · best of 2 rows27.9%IndependentreasoningoffPartially comparable-12.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
290NVIDIA-Nemotron-Nano-9B-v2Open weightsNVIDIA · Nemotron · best of 4 rows27.6%IndependentreasoningonPartially comparable-12.6 ptobs. 12 Sept 2026artificialanalysis.aiT2History
291R1 Distill Llama 70BOpen weightsDeepSeek · Llama · best of 2 rows27.6%IndependentreasoningonPartially comparable-12.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
292Molmo2-8BOpen weightsAllen Institute for AI · best of 2 rows26.9%IndependentreasoningoffPartially comparable-13.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
293qwen3-1.7b-instructOpen weightsAlibaba Group · Qwen3.1 · best of 4 rows26.9%IndependentreasoningonPartially comparable-13.3 ptobs. 12 Sept 2026artificialanalysis.aiT2History
294Ministral 3 3BOpen weightsMistral AI · Ministral 3 · best of 2 rows26.8%IndependentreasoningoffPartially comparable-13.4 ptobs. 12 Sept 2026artificialanalysis.aiT2History
295minicpm-v4-6-1-3bOpen weightsOpenBMB · best of 2 rows26.7%IndependentreasoningoffPartially comparable-13.5 ptobs. 12 Sept 2026artificialanalysis.aiT2History
296sarvam-30bOpen weightsSarvam · best of 2 rows26.5%Independentreasoning_efforthighPartially comparable-13.7 ptobs. 12 Sept 2026artificialanalysis.aiT2History
297Mistral Small 3Open weightsMistral AI · Mistral · best of 2 rows26.4%IndependentreasoningoffPartially comparable-13.8 ptobs. 12 Sept 2026artificialanalysis.aiT2History
298lfm2-8b-a1bOpen weightsLiquid AI · LFM2 · best of 2 rows26.3%IndependentreasoningoffPartially comparable-13.9 ptobs. 12 Sept 2026artificialanalysis.aiT2History
299Llama-3.2-3BRestricted weightsMeta AI · Llama 3.2 · best of 2 rows26.2%IndependentreasoningoffPartially comparable-14.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History
299granite-4-0-h-nano-1bOpen weightsIBM · Granite 4.0 · best of 2 rows26.2%IndependentreasoningoffPartially comparable-14.0 ptobs. 12 Sept 2026artificialanalysis.aiT2History

One row per canonical model — its best current row inside this comparability group (effort variants are folded into the model). Bars are relative to the page's best score. “vs leader” reads comparability: partially comparable = same task, conditions differ (reasoning effort, temperature, judge). Rules →