Updated 24 min ago · first seen 11 Sept 2026
bench_01M293SPECN73H13J38KYN4DNN
- Metric
- pass@1 · %
- Direction
- —
- Results
- 0
- Leader
- —
Leaderboard 0 current results
Select models with +, then open Compare.
No benchmark results recorded
Definition
- Category
- coding
Source:AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2observed 24 min agomedium
- Task
- Python function synthesis from docstrings
Source:AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2observed 24 min agomedium
- Metric
- pass@1 · %
Source:AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2observed 24 min agomedium
- Paper
- https://arxiv.org/abs/2107.03374
Source:AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2observed 24 min agomedium
- Known limitations
- Saturated; small (164 problems).
Source:AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2observed 24 min agomedium
- Website
Source:AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2observed 24 min agomedium
Each value shows its source, tier and observation time. Missing rows mean no source stated them. How results are recorded →
No relations recorded.
As of
Rewind the record: see this entity's attributes exactly as AI Atlas knew them on a given day.
Claim history
Categorycategory1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| coding | → current | current | AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2 | medium | curated |
Known limitationsknown_limitations1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| Saturated; small (164 problems). | → current | current | AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2 | medium | curated |
Metricmetric1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| pass@1 | → current | current | AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2 | medium | curated |
Paperpaper1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| https://arxiv.org/abs/2107.03374 | → current | current | AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2 | medium | curated |
Tasktask1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| Python function synthesis from docstrings | → current | current | AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2 | medium | curated |
Unitunit1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| % | → current | current | AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2 | medium | curated |
Websitewebsite1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| https://github.com/openai/human-eval | → current | current | AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2 | medium | curated |
Claims are temporal and append-only: a new observation closes the previous claim (valid_to) instead of overwriting it. Conflicting claims from different sources are kept side by side and flagged — never averaged. Methodology →
No documents recorded