Updated 1 h ago · first seen 11 Sept 2026
bench_01M293SPE4R3B1PPEGHN82HBCE
- Metric
- accuracy · %
- Direction
- —
- Results
- 0
- Leader
- —
Leaderboard 0 current results
Select models with +, then open Compare.
No benchmark results recorded
Definition
- Category
- knowledge
Source:AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2observed 1 h agomedium
- Task
- harder, 10-option MMLU variant
Source:AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2observed 1 h agomedium
- Metric
- accuracy · %
Source:AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2observed 1 h agomedium
- Paper
- https://arxiv.org/abs/2406.01574
Source:AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2observed 1 h agomedium
- Website
Source:AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2observed 1 h agomedium
Each value shows its source, tier and observation time. Missing rows mean no source stated them. How results are recorded →
No relations recorded.
As of
Rewind the record: see this entity's attributes exactly as AI Atlas knew them on a given day.
Claim history
Categorycategory1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| knowledge | → current | current | AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2 | medium | curated |
Metricmetric1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| accuracy | → current | current | AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2 | medium | curated |
Paperpaper1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| https://arxiv.org/abs/2406.01574 | → current | current | AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2 | medium | curated |
Tasktask1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| harder, 10-option MMLU variant | → current | current | AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2 | medium | curated |
Unitunit1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| % | → current | current | AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2 | medium | curated |
Websitewebsite1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| https://github.com/TIGER-AI-Lab/MMLU-Pro | → current | current | AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2 | medium | curated |
Claims are temporal and append-only: a new observation closes the previous claim (valid_to) instead of overwriting it. Conflicting claims from different sources are kept side by side and flagged — never averaged. Methodology →
No documents recorded