h-matched tracker

Time from benchmark release to human-level AI performance

A benchmark is h-matched on the date an AI system first reaches the human baseline its authors published. This page records that interval for every benchmark with a measured human baseline, and shows what it is doing over time. Benchmarks whose "human baseline" turns out to be an estimated ceiling, a passing threshold, or a model's own score are excluded.

Benchmarks
62
H-matched
35
Open
23
Unreported
4
Median interval
4.33 yr
Updated Sep 4, 2026Most recent h-match: ARC-AGI-3, GPT-6 Astra (OpenAI), after 5m 9d (Sep 3, 2026)

1.Interval by release year

One point per h-matched benchmark, placed at its release year against the years it then took to reach the human baseline.

Filled circles are clean h-matches, open circles contested ones, and the larger circle is the most recent. The dashed line is an ordinary least-squares fit through the plotted points only. It is shown for continuity with earlier versions of this page and should not be read as the trend: it cannot see the benchmarks that are still open, which is exactly the bias section 2 corrects.

2.Share not yet h-matched, by release cohort

Kaplan-Meier estimate of the share of each cohort still below its human baseline at a given age.

Benchmarks that are still open are censored at today's date, and unreported ones at the date of their last published score, so each curve stops where the evidence stops rather than implying the recent cohorts are finished. This is the estimate to read: averaging only the benchmarks that have already been h-matched understates the interval, because a benchmark released recently that will take years to fall cannot appear in that average yet.

3.Summary statistics

The survival estimate is the figure to quote. The others are included because they are the ones usually cited.

Median interval, survival estimateKaplan-Meier over all 62 benchmarks; open censored today, unreported at last published score4.33 years
Median interval, h-matched onlyBiased low: ignores benchmarks not yet h-matched1.82 years
Mean interval, h-matched onlyBiased low for the same reason2.65 years
Mean interval, released in the last 3 yearsThe most heavily censored subset, and the most often quoted0.97 years
Shortest interval-0.06 years
Longest interval8.83 years
H-matched within 1 / 2 / 3 yearsShare of h-matched benchmarks28.6% / 57.1% / 68.6%
Longest still openSince release; unreported benchmarks excluded7.94 years
H-matched / open / unreported35 / 23 / 4

4.H-matched benchmarks

Benchmarks an AI system has taken to the published human baseline, with the system that did it and the conditions. Any column sorts; the information icon opens the note and its sources.

35 of 35
By
ARC-AGI-3Mar 25, 2026Sep 3, 2026GPT-6 Astra (OpenAI)contested99.9% vs 100%Provider Adapter harness, $19K; 62.7% on the Standard harness5 mo
SocialIQASep 9, 2019Sep 1, 2026Claude Fable 5.186.44% vs 84.4%zero-shot, full 1,954-item dev set, measured for this tracker6 yr 11 mo
SimpleBenchOct 31, 2024Sep 1, 2026Claude Fable 5.186.6% vs 83.7%AVG@51 yr 9 mo
LAB-Bench (FigQA)Jul 15, 2024Apr 16, 2026Claude Opus 4.779.3% vs 77%no tools (85.4% with Python tools)1 yr 8 mo
BoolQMay 24, 2019Feb 6, 2026Claude Opus 4.6contested91.9% vs 90%zero-shot, full 3,270-item validation set, measured for this tracker6 yr 8 mo
OSWorldApr 11, 2024Feb 6, 2026Claude Opus 4.672.7% vs 72.36%1 yr 9 mo
FrontierMath (Tier 1–3)Nov 7, 2024Dec 11, 2025GPT-5.2 Thinking (OpenAI)contested40.3% vs 35%with tool use1 yr 1 mo
CharXiv-RJun 26, 2024Apr 16, 2025o3 (OpenAI)78.6% vs 71.3%9 mo
EgoSchemaAug 17, 2023Jan 26, 2025Qwen2-VL-72B-Instruct77.9% vs 76%1 yr 5 mo
ARC-AGI-1 (Verified)Nov 5, 2019Dec 20, 2024o3 (OpenAI)contested87.5% vs 85%semi-private eval, high-compute configuration5 yr 1 mo
LongBench v2Jan 3, 2025Dec 12, 2024o1-preview (OpenAI)57.7% vs 53.7%-22 dbefore release
MATHNov 8, 2021Sep 12, 2024o1 (OpenAI)94.8% vs 90%2 yr 10 mo
GPQANov 29, 2023Sep 12, 2024o1 (OpenAI)78.3% vs 69.7%9 mo
TextVQAMay 13, 2019Sep 1, 2024Qwen2-VL-72B-Instruct85.5% vs 85%5 yr 3 mo
ScienceQASep 20, 2022Aug 1, 2024Phi-3.5-vision-instruct (Microsoft)91.3% vs 88.4%1 yr 10 mo
BIG-Bench-HardOct 17, 2022Jun 21, 2024Claude 3.5 Sonnet93.1% vs 94.4%3-shot CoT1 yr 8 mo
MathVistaJan 21, 2024May 13, 2024GPT-4o63.8% vs 60.3%3 mo
HellaSwagJun 19, 2019Mar 4, 2024Claude 3 Opus95.4% vs 95.6%10-shot4 yr 8 mo
PubMedQANov 3, 2019Mar 4, 2024Claude 3 Sonnet79.7% vs 78%4 yr 3 mo
GSM8KNov 18, 2021Mar 14, 2023GPT-487.1% vs 60%1 yr 3 mo
HumanEvalJul 14, 2021Mar 14, 2023unattributed1 yr 7 mo
MMLUSep 7, 2020Dec 15, 2022unattributed89.8%2 yr 3 mo
TriviaQAMay 13, 2017Aug 8, 2022Atlas (Meta AI)84.7% vs 79.7%5 yr 2 mo
CommonsenseQANov 2, 2018Jul 23, 2022KEAR (Microsoft)89.4% vs 88.9%3 yr 8 mo
VQAMay 3, 2019Jun 15, 2022unattributed3 yr 1 mo
Adversarial NLIOct 31, 2019Jun 15, 2021unattributed1 yr 7 mo
SuperGLUEFeb 13, 2020Dec 15, 2020unattributed89.8 pts10 mo
WinoGradJan 1, 2011Nov 1, 2019RoBERTa fine-tuned on WinoGrandecontested90.1% vs 92%8 yr 9 mo
GLUENov 1, 2018Jul 1, 2019XLNet (Yang et al.)88.4 pts vs 87.1 pts7 mo
CoQAAug 21, 2018Mar 29, 2019Microsoft Research Asia ensemble89.4 F1 vs 88.8 F1overall test-set F17 mo
SQuAD 2.0Jun 11, 2018Mar 15, 2019unattributed89.47 F1 vs 89.45 F19 mo
SQuAD 1.1Jun 16, 2016Sep 15, 2018unattributed91.22 F12 yr 2 mo
ImageNet ChallengeJan 1, 2009Mar 15, 2016ResNet ensemble (He et al.)96.4% vs 95%top-5 accuracy7 yr 2 mo
Arcade Learning EnvironmentJun 21, 2013Sep 22, 2015Double DQN114.7% vs 100%median human-normalized score, 49 games, 5-minute episodes2 yr 2 mo
GTSRBJan 19, 2011Aug 1, 2011IDSIA committee of CNNs99.46% vs 98.84%final IJCNN competition round, multi-column DNN committee6 mo

5.Not yet h-matched

Two different situations, kept apart: benchmarks the field is still failing, and benchmarks nobody has measured a current model on for years.

Open

Below the human baseline, and labs still publish scores on them.

23
BenchmarkReleasedOpen forHuman baseline
PaperBenchApr 2, 20251y 5m41.4%expert, n=8
ZeroBenchFeb 13, 20251y 6m29.5%small sample, n=15
VSI-BenchDec 18, 20241y 8m79%small sample
BioLP-benchAug 31, 20242y 4d38.4%expert
BELEBELEJul 25, 20242y 1m97.6%unspecified
BLINKJul 3, 20242y 2m95.7%unspecified
MMMUJun 13, 20242y 2m88.6%expert
ReMIJun 13, 20242y 2m95.8%unspecified
MathVerseMar 21, 20242y 5m64.9%small sample, n=10
TempCompassMar 1, 20242y 6m97.3%small sample, n=3
MMVPJan 11, 20242y 7m95.7%small sample, n=4
GAIANov 21, 20232y 9m92%crowd
BIRD-SQLNov 15, 20232y 9m92.96%expert
METATOOLOct 5, 20232y 10m96%unspecified
WebArenaJul 25, 20233y 1m78.24%small sample, n=5
Perception TestMay 23, 20233y 3m91.4%unspecified
TruthfulQAMay 8, 20224y 3m94%unspecified
InfographicVQAAug 22, 20215y 13d95.7%unspecified
NExT-QAMay 18, 20215y 3m88.38%crowd
Habitat ObjectNavJun 23, 20206y 2m88.9%unspecified
ALFREDDec 3, 20196y 9m91%small sample, n=5
WinoGrandeNov 21, 20196y 9m94%crowd
HotpotQASep 25, 20187y 11m82.55 F1crowd

Unreported

No published frontier-model score for about two years. Status unknown: a reporting gap, not evidence that they are hard.

4
BenchmarkReleasedLast reportedHuman baseline
PIQANov 26, 2019Dec 27, 202494.9%crowd
SpatialSenseAug 29, 2019unknown94.6%unspecified
DROPApr 16, 2019Dec 27, 202496.4 F1expert
RACEApr 17, 2017Dec 27, 202494.5%expert

6.Excluded benchmarks

Benchmarks considered and left out, with the reason. Checking whether a published human baseline is really a human baseline is most of the work behind this page, so the results are recorded rather than discarded.

A benchmark earns a place on this tracker by publishing a human score measured on the same metric and split that models are scored on. These did not. In 2 cases the number circulating as the human baseline is a model's own score, and one of those was live on this page until it was checked.

BenchmarkWhat the number isDetail
HallusionBenchquoted as 65.28%a model's scoreThe paper's Table 2 gives two rows per model, labelled Human and GPT4-Assisted. Those are the two ways GPT-4V's free-text answers were graded, not two contestants: 65.28 is GPT-4V's own all-accuracy under GPT-4-assisted grading. The paper publishes no human baseline. This entry was live on this tracker until September 2026.
OCRBench v2quoted as variousa model's scoreThe figures circulating as human baselines are InternVL3-14B scores, mislabelled by a downstream aggregator and repeated from there.
ChartQAannotator agreementReports inter-annotator agreement, which measures whether two people give the same answer, not whether either is right. It is widely quoted as a human baseline anyway.
LAMBADAquoted as ~86%not in the sourceThe paper contains no measured human accuracy. Its only human-related number describes how many candidate passages were discarded during dataset construction. The circulating 86% looks like a rounding of GPT-3's own 86.4%.
MedQA (USMLE)quoted as ~60%a pass markThat is the USMLE pass/fail policy cutoff, not measured human accuracy. The paper explicitly declines to give a human baseline, citing high variance among examinees.
BrowseCompquoted as 29.2%a pass markA bounded-effort solve rate: what trainers managed within a few hours, not what people can do. It was passed almost immediately, which says more about the bar than the systems.
DocVQAquoted as 94.36 ANLSdifferent measurementThe paper reports two human numbers: 0.981 ANLS and 94.36% accuracy. The 94.36 figure is the accuracy one, routinely quoted against model ANLS scores. Compared correctly, the best systems are around 0.971 ANLS and the benchmark is not yet h-matched.
COCO Captionsquoted as 0.854 CIDEr-Ddifferent measurementThe baseline is a leave-one-out n-gram score for held-out reference captions, not measured human quality. Machines passed it in 2015 while the same project’s human judges still preferred human captions, which is the clearest example on this list of a metric being beaten without the claim behind it holding.
TextCapsquoted as 125.5 CIDErdifferent measurementSame leave-one-out CIDEr construction as COCO Captions, and the same objection.
BigCodeBenchquoted as 97%authors marking own workEleven of the paper's own annotators solved 33 of the tasks they had just written, as a data-quality check. It is not an independent measurement of human ability, and it covers a fraction of the benchmark.
ARC-AGI-2an estimated ceilingSeveral incompatible human figures circulate, from a panel measure where a task counts as solved if any of about ten testers gets it, down to a recomputed individual average near 53%. Until the intended baseline is settled the interval cannot be computed.
Windows Agent Arenaquoted as 74.5%an estimated ceilingReported from a single casual user over roughly ninety minutes. Too thin to treat as a population baseline, though the benchmark is otherwise a good fit.
CodeElo / CodeContestsa percentileThe human comparison is a rank in the live Codeforces rating distribution. There is no fixed score to cross, and the bar moves as the population does, so a date of first match is not well defined.
MLE-benchout of scopeThe human reference is Kaggle medal thresholds across 75 separate competitions. Real and meaningful, but it is 75 baselines on 75 metrics rather than one number to cross.
AGIEvalout of scopePublishes per-exam figures for average and top test-takers rather than one baseline, so which number counts as human level is a choice the tracker would be making, not the paper.
HLE, SWE-bench, MMLU-Pro, RULER, TAU-bench, miniF2F and othersno human baselineChecked and no human baseline is published. Also in this group: ARC, SNLI, MultiNLI, XNLI, HumanEval, MBPP, APPS, DS-1000, AI2D, DVQA, WikiTableQuestions, MMStar, SEED-Bench, MME, MVBench, Video-MME, ActivityNet-QA, NarrativeQA, CNN/DailyMail and XSum. Numbers circulate for several of them; none trace to a paper.

7.Data and corrections

Export

The whole table as one Markdown document: the summary statistics, the survival table by cohort, both benchmark lists, and every per-benchmark note with its references.

Corrections

Missing a benchmark, know of a newer score, or think a human baseline is mislabelled? Open an issue with a source and it will be added. Baselines that turn out to be estimated ceilings, passing thresholds, or a model's own score get removed, so those reports are welcome too.

Open an issueEmailPortfolio