1.Interval by release year
One point per h-matched benchmark, placed at its release year against the years it then took to reach the human baseline.
Filled circles are clean h-matches, open circles contested ones, and the larger circle is the most recent. The dashed line is an ordinary least-squares fit through the plotted points only. It is shown for continuity with earlier versions of this page and should not be read as the trend: it cannot see the benchmarks that are still open, which is exactly the bias section 2 corrects.
2.Share not yet h-matched, by release cohort
Kaplan-Meier estimate of the share of each cohort still below its human baseline at a given age.
Benchmarks that are still open are censored at today's date, and unreported ones at the date of their last published score, so each curve stops where the evidence stops rather than implying the recent cohorts are finished. This is the estimate to read: averaging only the benchmarks that have already been h-matched understates the interval, because a benchmark released recently that will take years to fall cannot appear in that average yet.
3.Summary statistics
The survival estimate is the figure to quote. The others are included because they are the ones usually cited.
| Median interval, survival estimateKaplan-Meier over all 62 benchmarks; open censored today, unreported at last published score | 4.33 years |
|---|---|
| Median interval, h-matched onlyBiased low: ignores benchmarks not yet h-matched | 1.82 years |
| Mean interval, h-matched onlyBiased low for the same reason | 2.65 years |
| Mean interval, released in the last 3 yearsThe most heavily censored subset, and the most often quoted | 0.97 years |
| Shortest interval | -0.06 years |
| Longest interval | 8.83 years |
| H-matched within 1 / 2 / 3 yearsShare of h-matched benchmarks | 28.6% / 57.1% / 68.6% |
| Longest still openSince release; unreported benchmarks excluded | 7.94 years |
| H-matched / open / unreported | 35 / 23 / 4 |
4.H-matched benchmarks
Benchmarks an AI system has taken to the published human baseline, with the system that did it and the conditions. Any column sorts; the information icon opens the note and its sources.
5.Not yet h-matched
Two different situations, kept apart: benchmarks the field is still failing, and benchmarks nobody has measured a current model on for years.
Open
Below the human baseline, and labs still publish scores on them.
Unreported
No published frontier-model score for about two years. Status unknown: a reporting gap, not evidence that they are hard.
6.Excluded benchmarks
Benchmarks considered and left out, with the reason. Checking whether a published human baseline is really a human baseline is most of the work behind this page, so the results are recorded rather than discarded.
A benchmark earns a place on this tracker by publishing a human score measured on the same metric and split that models are scored on. These did not. In 2 cases the number circulating as the human baseline is a model's own score, and one of those was live on this page until it was checked.
7.Data and corrections
Export
The whole table as one Markdown document: the summary statistics, the survival table by cohort, both benchmark lists, and every per-benchmark note with its references.
Corrections
Missing a benchmark, know of a newer score, or think a human baseline is mislabelled? Open an issue with a source and it will be added. Baselines that turn out to be estimated ceilings, passing thresholds, or a model's own score get removed, so those reports are welcome too.