Updated 5 h ago · first seen 11 Sept 2026
bench_01M293SPERF0D8QG7TSCGE8GGF
- Metric
- average score · %
- Direction
- —
- Results
- 456 · 0 filtered
Score history · Gemini 3.8 Flash 8 rows
Not enough history to chart — 8 observations, all dated 25 Jun 2026. Rows under different configurations count separately; the list below shows each one.
- 81.41%release=2026-06-25 · subtasks=["paraphrase","simplify","story_generation","summarize"] · livebench_model_id=gemini-3.8-flash-high25 Jun 2026
- 87.79%release=2026-06-25 · subtasks=["connections","plot_unscrambling","typos"] · livebench_model_id=gemini-3.8-flash-high25 Jun 2026
- 54.01%release=2026-06-25 · subtasks=["consecutive_events","tablejoin","tablereformat"] · livebench_model_id=gemini-3.8-flash-high25 Jun 2026
- 91.56%release=2026-06-25 · subtasks=["AMPS_Hard","integrals_with_game","math_comp","olympiad"] · livebench_model_id=gemini-3.8-flash-high25 Jun 2026
- 54.24%release=2026-06-25 · subtasks=["javascript","typescript","python"] · livebench_model_id=gemini-3.8-flash-high25 Jun 2026
- 72.49%release=2026-06-25 · subtasks=["code_generation","code_completion"] · livebench_model_id=gemini-3.8-flash-high25 Jun 2026
- 89.29%release=2026-06-25 · subtasks=["theory_of_mind","zebra_puzzle","spatial","logic_with_navigation"] · livebench_model_id=gemini-3.8-flash-high25 Jun 2026
- 75.83%release=2026-06-25 · aggregation=mean of category averages; category = mean of its subtasks · livebench_model_id=gemini-3.8-flash-high25 Jun 2026
Leaderboard 0 current results · config contains “["code_generation","code_completion"]”
Select models with +, then open Compare.
No result matches this configuration filter
The config filter matches a value inside each result's configuration (server-side, `config=` on the API). Chips are the values shared by several rows on the first page; per-model identifiers are not offered.
Definition
- Category
- general
Source:AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2observed 5 h agomedium
- Task
- contamination-limited, monthly refreshed questions across 6 categories
Source:AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2observed 5 h agomedium
- Metric
- average score · %
Source:AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2observed 5 h agomedium
- Paper
- https://arxiv.org/abs/2406.19314
Source:AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2observed 5 h agomedium
- Website
Source:AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2observed 5 h agomedium
Each value shows its source, tier and observation time. Missing rows mean no source stated them. How results are recorded →
As of
Rewind the record: see this entity's attributes exactly as AI Atlas knew them on a given day.
Claim history
Categorycategory1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| general | → current | current | AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2 | medium | curated |
Metricmetric1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| average score | → current | current | AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2 | medium | curated |
Paperpaper1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| https://arxiv.org/abs/2406.19314 | → current | current | AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2 | medium | curated |
Tasktask1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| contamination-limited, monthly refreshed questions across 6 categories | → current | current | AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2 | medium | curated |
Unitunit1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| % | → current | current | AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2 | medium | curated |
Websitewebsite1
| Value | Valid from → to | Status | Source | Confidence | Extractor |
|---|---|---|---|---|---|
| https://livebench.ai | → current | current | AI Atlas curated registry (YAML, versioned in git, every entry carries its source URL)T2 | medium | curated |
Claims are temporal and append-only: a new observation closes the previous claim (valid_to) instead of overwriting it. Conflicting claims from different sources are kept side by side and flagged — never averaged. Methodology →
| Source | Document | Type | Tier | Last observed | Snapshots |
|---|---|---|---|---|---|
| LiveBench | livebench.ai/table_2026_06_25.csv | leaderboard | T2· Quality secondary | 3 h ago | 1 |
Tier 1 = official/primary, 2 = quality secondary, 3 = community, 4 = unverified. Every snapshot is archived; see all sources and the methodology.