Skip to content
AI Atlas
FrameworkActive

evals

OpenAIgithub.com/openai/evals

Evals is a framework for evaluating LLMs and LLM systems, and an open-source registry of benchmarks.

Updated 46 min ago · first seen 11 Sept 2026

framework_01M294NHSHM7YQDAFJSC10DRCG

Version
2.0.0
T2 · 51 min ago
Released
12 Jan 2024
T2 · 51 min ago
Stars
19,436
T2 · 46 min ago

As of

Rewind the record: see this entity's attributes exactly as AI Atlas knew them on a given day.

Claim history · Kind

1 claims · 1 propertiesShow all properties

Kindkind1

Claim history for Kind
ValueValid from → toStatusSourceConfidenceExtractor
toolcurrentcurrentGitHub public repositories (HTML, releases.atom, raw files)T2mediumdeterministic

Claims are temporal and append-only: a new observation closes the previous claim (valid_to) instead of overwriting it. Conflicting claims from different sources are kept side by side and flagged — never averaged. Methodology →