Updated 8 h ago · first seen 11 Sept 2026
bench_01M293SPEGA1XMKH0SMTFTCK2R
- Metric
- resolved · %
- Direction
- Higher is better
- Results
- 180 · 104 filtered
- Leader
- Claude Opus 4.5 79.2%
Score history · claude-3-5-sonnet-20241022 2 rows
- claude-3-5-sonnet-20241022
- 51.6%date=2025-01-22 · board=Verified · system=AutoCodeRover-v2.1 · model_tag=claude-3-5-sonnet-2024102222 Jan 2025
- 46.2%date=2024-11-08 · board=Verified · system=AutoCodeRover-v2.0 · model_tag=claude-3-5-sonnet-202410228 Nov 2024
Leaderboard 104 current results · config contains “false”
Select models with +, then open Compare.
| # | Model | Score | Config | Evaluated | Source | Actions |
|---|---|---|---|---|---|---|
| #1Claude Opus 4.5Anthropic | 79.2% | date=2025-12-05 · board=Verified · system=Sonar Foundation Agent · model_tag=claude-opus-4-5 | 5 Dec 2025 | swebench.comT2 | History | |
| #2Claude Opus 4.5Anthropic | 79.2% | date=2025-12-15 · board=Verified · system=live-SWE-agent · model_tag=claude-opus-4-5-20251101 | 15 Dec 2025 | swebench.comT2 | History | |
| #3doubao-seed-codeByteDance | 78.8% | date=2025-09-28 · board=Verified · system=TRAE · model_tag=Doubao-Seed-Code | 28 Sept 2025 | swebench.comT2 | History | |
| #4Gemini 3 Pro PreviewGoogle | 77.4% | date=2025-11-20 · board=Verified · system=live-SWE-agent · model_tag=gemini-3-pro-preview | 20 Nov 2025 | swebench.comT2 | History | |
| #5MultipleAnthropic | 76.8% | date=2025-09-02 · board=Verified · system=Atlassian Rovo Dev · model_tag=claude-sonnet-4-20250514 | 2 Sept 2025 | swebench.comT2 | History | |
| #6Claude 4 SonnetAnthropic | 76.8% | date=2025-08-04 · board=Verified · system=EPAM AI/Run Developer Agent · model_tag=claude-sonnet-4-20250514 | 4 Aug 2025 | swebench.comT2 | History | |
| #7MultipleAnthropic | 76.4% | date=2025-08-19 · board=Verified · system=ACoder · model_tag=claude-4-sonnet | 19 Aug 2025 | swebench.comT2 | History | |
| #8MultipleAnthropic | 75.6% | date=2025-09-01 · board=Verified · system=Warp · model_tag=gpt-5 | 1 Sept 2025 | swebench.comT2 | History | |
| #9MultipleAnthropic | 75.2% | date=2025-06-12 · board=Verified · system=TRAE · model_tag=claude-4-sonnet-20250522 | 12 Jun 2025 | swebench.comT2 | History | |
| #10Claude 4.5 SonnetAnthropic | 74.8% | date=2025-11-03 · board=Verified · system=Sonar Foundation Agent · model_tag=claude-sonnet-4-5 | 3 Nov 2025 | swebench.comT2 | History | |
| #11Claude Sonnet 4Anthropic | 74.8% | date=2025-07-31 · board=Verified · system=Harness AI · model_tag=claude-sonnet-4-20250514 | 31 Jul 2025 | swebench.comT2 | History | |
| #12MultipleAnthropic | 74.6% | date=2025-09-15 · board=Verified · system=JoyCode · model_tag=claude-4-sonnet | 15 Sept 2025 | swebench.comT2 | History | |
| #13Claude 4 SonnetAnthropic | 74.6% | date=2025-07-20 · board=Verified · system=Lingxi-v1.5 · model_tag=claude-4-sonnet-20250514 | 20 Jul 2025 | swebench.comT2 | History | |
| #14gpt-5OpenAI | 74.4% | date=2025-10-15 · board=Verified · system=Prometheus-v1.2.1 · model_tag=gpt-5-2025-08-07 | 15 Oct 2025 | swebench.comT2 | History | |
| #15MultipleAnthropic | 74.4% | date=2025-06-03 · board=Verified · system=Refact.ai Agent · model_tag=claude-4-sonnet | 3 Jun 2025 | swebench.comT2 | History | |
| #16MultipleAnthropic | 73.8% | date=2025-11-03 · board=Verified · system=Salesforce AI Research SAGE · model_tag=claude-sonnet-4.5 | 3 Nov 2025 | swebench.comT2 | History | |
| #17claude-4-opusAnthropic | 73.2% | date=2025-05-22 · board=Verified · system=Tools · model_tag=claude-4-opus-20250514 | 22 May 2025 | swebench.comT2 | History | |
| #18MultipleAnthropic | 73% | date=2025-10-21 · board=Verified · system=Salesforce AI Research SAGE · model_tag=claude-sonnet-4.5 | 21 Oct 2025 | swebench.comT2 | History | |
| #19Claude 4 SonnetAnthropic | 72.4% | date=2025-05-22 · board=Verified · system=Tools · model_tag=claude-4-sonnet-20250514 | 22 May 2025 | swebench.comT2 | History | |
| #20Claude 4 SonnetAnthropic | 71.2% | date=2025-07-10 · board=Verified · system=Bloop · model_tag=claude-4-sonnet-20250514 | 10 Jul 2025 | swebench.comT2 | History | |
| #21Kimi K2 0711Moonshot AI | 71.2% | date=2025-10-14 · board=Verified · system=Lingxi v1.5 · model_tag=kimi-k2-0905-preview | 14 Oct 2025 | swebench.comT2 | History | |
| #22gpt-5OpenAI | 71.2% | date=2025-09-29 · board=Verified · system=Prometheus-v1.2 · model_tag=gpt-5-2025-08-07 | 29 Sept 2025 | swebench.comT2 | History | |
| #23MultipleAnthropic | 71.2% | date=2025-07-15 · board=Verified · system=Qodo Command · model_tag=claude-sonnet-4-20250514 | 15 Jul 2025 | swebench.comT2 | History | |
| #24MultipleAnthropic | 71% | date=2025-06-23 · board=Verified · system=Warp · model_tag=claude-sonnet-4-20250514 | 23 Jun 2025 | swebench.comT2 | History | |
| #25Undisclosed | 70.6% | date=2025-05-19 · board=Verified · system=TRAE · submission=20250519_trae | 19 May 2025 | swebench.comT2 | History | |
| #26Undisclosed | 70.4% | date=2025-06-10 · board=Verified · system=Augment Agent v1 · submission=20250610_augment_agent_v1 | 10 Jun 2025 | swebench.comT2 | History | |
| #27MultipleAnthropic | 70.4% | date=2025-05-15 · board=Verified · system=Refact.ai Agent · model_tag=claude-3-7-sonnet-20250219 | 15 May 2025 | swebench.comT2 | History | |
| #28MultipleAnthropic | 70.2% | date=2025-05-19 · board=Verified · system=devlo · model_tag=claude-3-7-sonnet-20250219 | 19 May 2025 | swebench.comT2 | History | |
| #29MultipleAnthropic | 70% | date=2025-04-30 · board=Verified · system=Zencoder · model_tag=claude-3-7-sonnet-20250219 | 30 Apr 2025 | swebench.comT2 | History | |
| #30qwen3-coder-480b-a35b-instructAlibaba Group | 69.6% | date=2025-08-05 · board=Verified · system=OpenHands · model_tag=https://huggingface.co/Qwen/Qwen3-Coder-480B-A35B-Instruct | 5 Aug 2025 | swebench.comT2 | History | |
| #31MultipleAnthropic | 68.2% | date=2025-05-16 · board=Verified · system=Nemotron-CORTEXA · model_tag=NV-EmbedCode | 16 May 2025 | swebench.comT2 | History | |
| #32GLM 4.6Z.ai (Zhipu AI) | 68.2% | date=2025-09-30 · board=Verified · system=Undisclosed · model_tag=https://huggingface.co/zai-org/GLM-4.6 | 30 Sept 2025 | swebench.comT2 | History | |
| #33Claude 3.7 SonnetAnthropic | 66.4% | date=2025-05-14 · board=Verified · system=Aime-coder v1 · model_tag=claude-3-7-sonnet-20250219 | 14 May 2025 | swebench.comT2 | History | |
| #34Undisclosed | 65.4% | date=2025-04-05 · board=Verified · system=Amazon Q Developer Agent · submission=20250405_amazon-q-developer-agent-20250405-dev | 5 Apr 2025 | swebench.comT2 | History | |
| #35Undisclosed | 65.4% | date=2025-03-16 · board=Verified · system=Augment Agent v0 · submission=20250316_augment_agent_v0 | 16 Mar 2025 | swebench.comT2 | History | |
| #36o1 PreviewOpenAI | 64.6% | date=2025-01-17 · board=Verified · system=W&B Programmer O1 crosscheck5 · model_tag=o1-preview | 17 Jan 2025 | swebench.comT2 | History | |
| #37o4-miniOpenAI | 64.6% | date=2025-05-03 · board=Verified · system=PatchPilot-v1.1 · model_tag=o4-mini | 3 May 2025 | swebench.comT2 | History | |
| #38GLM 4.5Z.ai (Zhipu AI) | 64.2% | date=2025-07-28 · board=Verified · system=Undisclosed · model_tag=https://huggingface.co/zai-org/GLM-4.5 | 28 Jul 2025 | swebench.comT2 | History | |
| #39claude-35-sonnetAnthropic | 63.4% | date=2025-02-06 · board=Verified · system=AgentScope · model_tag=claude-3-5-sonnet-20241022 | 6 Feb 2025 | swebench.comT2 | History | |
| #40Claude 3.7 SonnetAnthropic | 63.2% | date=2025-02-24 · board=Verified · system=Tools · submission=20250224_tools_claude-3-7-sonnet | 24 Feb 2025 | swebench.comT2 | History | |
| #41claude-35-sonnetAnthropic | 62.8% | date=2025-02-28 · board=Verified · system=EPAM AI/Run Developer Agent · model_tag=claude-3-5-sonnet-20241022 | 28 Feb 2025 | swebench.comT2 | History | |
| #42Undisclosed | 62.8% | date=2025-01-10 · board=Verified · system=Blackbox AI Agent · submission=20250110_blackboxai_agent_v1.1 | 10 Jan 2025 | swebench.comT2 | History | |
| #43swe-searchAnthropic | 62.2% | date=2024-12-21 · board=Verified · system=CodeStory Midwit Agent · model_tag=claude-3-5-sonnet-20241022 | 21 Dec 2024 | swebench.comT2 | History | |
| #44Qwen3 Coder 30B A3B InstructQwen | 60.4% | date=2025-09-01 · board=Verified · system=EntroPO + R2E · model_tag=Qwen3-Coder-30B-A3B-Instruct | 1 Sept 2025 | swebench.comT2 | History | |
| #45claude-35-sonnetAnthropic | 60.2% | date=2025-01-10 · board=Verified · system=Learn-by-interact · model_tag=claude-3-5-sonnet-20241022 | 10 Jan 2025 | swebench.comT2 | History | |
| #46TTS(Bo16)OpenAI | 58.8% | date=2025-06-29 · board=Verified · system=DeepSWE-Preview · model_tag=https://huggingface.co/agentica-org/DeepSWE-Preview | 29 Jun 2025 | swebench.comT2 | History | |
| #47Undisclosed | 58.2% | date=2024-12-13 · board=Verified · system=devlo · submission=20241213_devlo | 13 Dec 2024 | swebench.comT2 | History | |
| #48MultipleAnthropic | 58.2% | date=2025-04-10 · board=Verified · system=Nemotron-CORTEXA · model_tag=NV-EmbedCode | 10 Apr 2025 | swebench.comT2 | History | |
| #49MultipleAnthropic | 57.2% | date=2024-12-23 · board=Verified · system=Emergent E1 · model_tag=claude-3-5-sonnet-20241022 | 23 Dec 2024 | swebench.comT2 | History | |
| #50Claude Sonnet 4Anthropic | 57% | date=2025-09-24 · board=Verified · system=Artemis Agent v2 · model_tag=claude-sonnet-4-20250514 | 24 Sept 2025 | swebench.comT2 | History | |
| #51Undisclosed | 57% | date=2024-12-08 · board=Verified · system=Gru · submission=20241208_gru | 8 Dec 2024 | swebench.comT2 | History | |
| #52Undisclosed | 56.6% | date=2025-04-05 · board=Verified · system=SWE-Rizzo · submission=20250405_swe-rizzo_claude37 | 5 Apr 2025 | swebench.comT2 | History | |
| #53claude-35-sonnetAnthropic | 55.4% | date=2024-12-12 · board=Verified · system=EPAM AI/Run Developer Agent · model_tag=claude-3-5-sonnet-20241022 | 12 Dec 2024 | swebench.comT2 | History | |
| #54Undisclosed | 55% | date=2024-12-02 · board=Verified · system=Amazon Q Developer Agent · submission=20241202_amazon-q-developer-agent-20241202-dev | 2 Dec 2024 | swebench.comT2 | History | |
| #55claude-35-sonnetAnthropic | 54.2% | date=2024-11-08 · board=Verified · system=devlo · model_tag=claude-3-5-sonnet-20241022 | 8 Nov 2024 | swebench.comT2 | History | |
| #56Frogboss 32B 2510 | 53.6% | date=2025-11-10 · board=Verified · system=FrogBoss-32B-2510 · model_tag=FrogBoss-32B-2510 | 10 Nov 2025 | swebench.comT2 | History | |
| #57Kimi K2 0711Moonshot AI | 53.4% | date=2025-08-04 · board=Verified · system=CodeSweep - SWE-agent · model_tag=kimi-k2-instruct | 4 Aug 2025 | swebench.comT2 | History | |
| #58Undisclosed | 53.2% | date=2025-01-20 · board=Verified · system=Bracket.sh · submission=20250120_Bracket | 20 Jan 2025 | swebench.comT2 | History | |
| #59Gemini 2.0 Flash (v20241212-experimental)Google DeepMind | 52.2% | date=2024-12-12 · board=Verified · system=Google Jules · submission=20241212_google_jules_gemini_2.0_flash_experimental | 12 Dec 2024 | swebench.comT2 | History | |
| #60Qwen3 Coder 30B A3B InstructQwen | 52.2% | date=2025-09-01 · board=Verified · system=EntroPO + R2E · model_tag=Qwen3-Coder-30B-A3B-Instruct | 1 Sept 2025 | swebench.comT2 | History | |
| #61claude-35-sonnetAnthropic | 51.8% | date=2024-11-25 · board=Verified · system=Engine Labs · model_tag=claude-3-5-sonnet-20241022 | 25 Nov 2024 | swebench.comT2 | History | |
| #62claude-3-5-sonnet-20241022Anthropic | 51.6% | date=2025-01-22 · board=Verified · system=AutoCodeRover-v2.1 · model_tag=claude-3-5-sonnet-20241022 | 22 Jan 2025 | swebench.comT2 | History | |
| #63Qwen3 Coder 30B A3B InstructQwen | 51.6% | date=2025-08-05 · board=Verified · system=OpenHands · model_tag=https://huggingface.co/Qwen/Qwen3-Coder-30B-A3B-Instruct | 5 Aug 2025 | swebench.comT2 | History | |
| #64claude-35-sonnetAnthropic | 50.8% | date=2024-12-02 · board=Verified · system=Agentless-1.5 · model_tag=claude-3-5-sonnet-20241022 | 2 Dec 2024 | swebench.comT2 | History | |
| #65Undisclosed | 50% | date=2024-10-28 · board=Verified · system=Solver · submission=20241028_solver | 28 Oct 2024 | swebench.comT2 | History | |
| #66Undisclosed | 50% | date=2024-11-25 · board=Verified · system=Bytedance MarsCode Agent · submission=20241125_marscode-agent-dev | 25 Nov 2024 | swebench.comT2 | History | |
| #67MultipleAnthropic | 49.2% | date=2024-11-05 · board=Verified · system=nFactorial · model_tag=claude-3-5-sonnet-20241022 | 5 Nov 2024 | swebench.comT2 | History | |
| #68claude-35-sonnetAnthropic | 49% | date=2024-10-22 · board=Verified · system=Tools · model_tag=claude-3-5-sonnet-20241022 | 22 Oct 2024 | swebench.comT2 | History | |
| #69MultipleAnthropic | 46.6% | date=2024-10-23 · board=Verified · system=Emergent E1 · model_tag=claude-3-5-sonnet-20241022 | 23 Oct 2024 | swebench.comT2 | History | |
| #70claude-3-5-sonnet-20241022Anthropic | 46.2% | date=2024-11-08 · board=Verified · system=AutoCodeRover-v2.0 · model_tag=claude-3-5-sonnet-20241022 | 8 Nov 2024 | swebench.comT2 | History | |
| #71Co-PatcheR | 46% | date=2025-05-28 · board=Verified · system=PatchPilot · submission=20250528_patchpilot_Co-PatcheR | 28 May 2025 | swebench.comT2 | History | |
| #72Undisclosed | 45.4% | date=2024-09-24 · board=Verified · system=Solver · submission=20240924_solver | 24 Sept 2024 | swebench.comT2 | History | |
| #73Undisclosed | 45.2% | date=2024-08-24 · board=Verified · system=Gru · submission=20240824_gru | 24 Aug 2024 | swebench.comT2 | History | |
| #74Frogmini 14B 2510 | 45% | date=2025-11-10 · board=Verified · system=FrogMini-14B-2510 · model_tag=FrogMini-14B-2510 | 10 Nov 2025 | swebench.comT2 | History | |
| #75gemini-2.0-flash-expGoogle | 44.2% | date=2025-01-18 · board=Verified · system=CodeShellAgent · model_tag=gemini-2.0-flash-exp | 18 Jan 2025 | swebench.comT2 | History | |
| #76Undisclosed | 43.6% | date=2024-09-20 · board=Verified · system=Solver · submission=20240920_solver | 20 Sept 2024 | swebench.comT2 | History | |
| #77Amazon.nova Premier v1:0 | 42.4% | date=2025-05-27 · board=Verified · system=Amazon Nova Premier 1.0 · model_tag=amazon.nova-premier-v1:0 | 27 May 2025 | swebench.comT2 | History | |
| #78DeepSWE-Preview | 42.2% | date=2025-06-29 · board=Verified · system=R2E-Gym · model_tag=https://huggingface.co/agentica-org/DeepSWE-Preview | 29 Jun 2025 | swebench.comT2 | History | |
| #79DeepSeek V3 0324DeepSeek | 42% | date=2025-08-06 · board=Verified · system=SWE-Exp · model_tag=DeepSeek-V3-0324 | 6 Aug 2025 | swebench.comT2 | History | |
| #80Undisclosed | 41.6% | date=2025-01-12 · board=Verified · system=ugaiforge · submission=20250112_ugaiforge | 12 Jan 2025 | swebench.comT2 | History | |
| #81MultipleAnthropic | 41.6% | date=2024-10-30 · board=Verified · system=nFactorial · model_tag=claude-3-5-sonnet-20241022 | 30 Oct 2024 | swebench.comT2 | History | |
| #82Llama3-SWE-RL-70BMeta AI | 41.2% | date=2025-02-26 · board=Verified · system=Agentless Mini · model_tag=Llama3-SWE-RL-70B | 26 Feb 2025 | swebench.comT2 | History | |
| #83Undisclosed | 40.6% | date=2024-08-20 · board=Verified · system=Honeycomb · submission=20240820_honeycomb | 20 Aug 2024 | swebench.comT2 | History | |
| #84MultipleAnthropic | 40.6% | date=2024-11-13 · board=Verified · system=Nebius AI · model_tag=Llama 3.1 | 13 Nov 2024 | swebench.comT2 | History | |
| #85Claude Haiku 3.5Anthropic | 40.6% | date=2024-10-22 · board=Verified · system=Tools · model_tag=claude-3-haiku-20240307 | 22 Oct 2024 | swebench.comT2 | History | |
| #86claude-35-sonnetAnthropic | 40.6% | date=2024-10-16 · board=Verified · system=Composio SWEkit · model_tag=claude-3-5-sonnet-20241022 | 16 Oct 2024 | swebench.comT2 | History | |
| #87claude-35-sonnetAnthropic | 39.6% | date=2024-10-29 · board=Verified · system=EPAM AI/Run Developer Agent · model_tag=claude-3-5-sonnet-20241022 | 29 Oct 2024 | swebench.comT2 | History | |
| #88Undisclosed | 38.8% | date=2024-07-21 · board=Verified · system=Amazon Q Developer Agent · submission=20240721_amazon-q-developer-agent-20240719-dev | 21 Jul 2024 | swebench.comT2 | History | |
| #89gpt-4oOpenAI | 38.8% | date=2024-10-28 · board=Verified · system=Agentless-1.5 · model_tag=gpt-4o-2024-05-13 | 28 Oct 2024 | swebench.comT2 | History | |
| #90gpt-4oOpenAI | 38.4% | date=2024-06-28 · board=Verified · system=AutoCodeRover · model_tag=gpt-4o-2024-05-13 | 28 Jun 2024 | swebench.comT2 | History | |
| #91Undisclosed | 37% | date=2024-06-17 · board=Verified · system=Factory Code Droid · submission=20240617_factory_code_droid | 17 Jun 2024 | swebench.comT2 | History | |
| #92gpt-4oOpenAI | 32.6% | date=2024-06-12 · board=Verified · system=MASAI · model_tag=gpt-4o-2024-08-06 | 12 Jun 2024 | swebench.comT2 | History | |
| #93Undisclosed | 32% | date=2024-11-20 · board=Verified · system=Artemis Agent v1 · submission=20241120_artemis_agent | 20 Nov 2024 | swebench.comT2 | History | |
| #94Undisclosed | 31.6% | date=2024-10-07 · board=Verified · system=nFactorial · submission=20241007_nfactorial | 7 Oct 2024 | swebench.comT2 | History | |
| #95Qwen2.5 (7B + 72B)Qwen | 30.2% | date=2024-11-28 · board=Verified · system=SWE-Fixer · model_tag=Qwen 2.5 | 28 Nov 2024 | swebench.comT2 | History | |
| #96Lingma SWE-GPT 72b (v0925)OpenAI | 28.8% | date=2024-10-02 · board=Verified · system=Lingma Agent · submission=20241002_lingma-agent_lingma-swe-gpt-72b | 2 Oct 2024 | swebench.comT2 | History | |
| #97gpt-4oOpenAI | 27% | date=2024-10-16 · board=Verified · system=EPAM AI/Run Developer Agent · model_tag=gpt-4o-2024-08-06 | 16 Oct 2024 | swebench.comT2 | History | |
| #98Undisclosed | 25.8% | date=2024-10-01 · board=Verified · system=nFactorial · submission=20241001_nfactorial | 1 Oct 2024 | swebench.comT2 | History | |
| #99Undisclosed | 25.6% | date=2024-05-09 · board=Verified · system=Amazon Q Developer Agent · submission=20240509_amazon-q-developer-agent-20240430-dev | 9 May 2024 | swebench.comT2 | History | |
| #100Lingma SWE-GPT 72b (v0918)OpenAI | 25% | date=2024-09-18 · board=Verified · system=Lingma Agent · submission=20240918_lingma-agent_lingma-swe-gpt-72b | 18 Sept 2024 | swebench.comT2 | History |
Scores are reported as published, with their evaluation configuration (harness, prompting, judge). The bar is relative to the best score on this page. Results with different configs are not directly comparable — see methodology.
The config filter matches a value inside each result's configuration (server-side, `config=` on the API). Chips are the values shared by several rows on the first page; per-model identifiers are not offered.
Definition
- Category
- coding
Source:AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2observed 8 h agomedium
- Task
- resolve real GitHub issues (500 human-validated instances)
Source:AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2observed 8 h agomedium
- Metric
- resolved · %
Source:AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2observed 8 h agomedium
- Direction
- Higher is better
- Paper
- https://arxiv.org/abs/2310.06770
Source:AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2observed 8 h agomedium
- Known limitations
- Scaffold/agent dependent; results are not comparable across harnesses.
Source:AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2observed 8 h agomedium
- Website
Source:AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2observed 8 h agomedium
Each value shows its source, tier and observation time. Missing rows mean no source stated them. How results are recorded →
- Evaluated models
- Claude Opus 4.6, doubao-seed-code, gemini-3-flash, MiniMax M2.5, Claude Opus 4.5, Gemini 3 Pro Preview, GLM 5, GPT-5.2-Codex +74(82 total)
As of
Rewind the record: see this entity's attributes exactly as AI Atlas knew them on a given day.
Claim history
Categorycategory1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| coding | → current | current | AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2 | medium | curated |
Known limitationsknown_limitations1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| Scaffold/agent dependent; results are not comparable across harnesses. | → current | current | AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2 | medium | curated |
Metricmetric1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| resolved | → current | current | AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2 | medium | curated |
Paperpaper1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| https://arxiv.org/abs/2310.06770 | → current | current | AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2 | medium | curated |
Tasktask1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| resolve real GitHub issues (500 human-validated instances) | → current | current | AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2 | medium | curated |
Unitunit1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| % | → current | current | AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2 | medium | curated |
Websitewebsite1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| https://www.swebench.com | → current | current | AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2 | medium | curated |
Claims are temporal and append-only: a new observation closes the previous claim (valid_to) instead of overwriting it. Conflicting claims from different sources are kept side by side and flagged — never averaged. Methodology →
| Source | Document | Type | Tier | Last observed | Snapshots |
|---|---|---|---|---|---|
| SWE-bench leaderboards | swebench.com/ | leaderboard | T2· Quality secondary | 6 h ago | 1 |
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.