Reading the benchmark plateau: saturation vs. real progress
When a 99% score stops meaning anything, and what this site does about it.
The TowardSingularity Team · Oct 2026 · 4 min read
A benchmark is useful right up until everyone aces it. Past that point, a higher score says nothing about capability, only that the test has run out of headroom. Treating a saturated benchmark as live progress is one of the most common ways timeline charts lie.
Saturation looks like progress
When a model goes from 97% to 99% on a test that humans score 95% on, the chart still ticks up and to the right. But the marginal points are coming from the easy tail, not from new ability. The benchmark has stopped discriminating between models.
“A maxed-out leaderboard measures the test, not the model.”
It also runs the other way, which is less often noticed. A score climbing towards 100% has to slow down as it gets there, because there is less and less room left. Plot it on a plain percentage axis and steady progress looks like a plateau. So a saturating test can fool you twice: once by flattering a model that is merely polishing the easy questions, and again by making real improvement look like a stall.
Two fixes, depending on the reading
The textbook fix is renormalisation: down-weight a benchmark as it gets close to saturation, so a 99% on a dead test cannot keep pushing a number up. This site does not do that, and for the reading it publishes it does not need to. The benchmark reading is the top rating on LMArena, which is relative and head-to-head rather than a percentage on a fixed test.
A rating like that has no ceiling to ace. When a model stops winning matchups its rating stops climbing, and nobody has to declare the test dead for that to happen. The top rating today is 1508.
Ratings have their own weakness, though: they measure what voters prefer, and preference can drift towards style. So the pipeline cross-checks the reading monthly against Stanford HELM, a fixed suite scored the old-fashioned way. On the latest check, 25 models appeared on both boards, and the rank correlation between them was 0.53. That is real agreement, and far from lockstep. Two boards that agree on the broad ordering and argue about the details are about what you would expect. A sudden collapse in that number would mean one of them has started measuring something else, and the site would flag it rather than absorb it.
Where saturation bites here
Saturation does bite in a domain forecast fitted on an accuracy over a fixed test set. Scientific reasoning is the clearest case.
GPQA Diamond is 198 four-choice questions written by PhD students in biology, chemistry and physics. It was built so that domain experts reach about 65% and skilled non-experts with web access about 34%. The best score in the data is now 94.8%, which is about ten questions wrong. The bar the site uses is 99%, which allows two.
So a date fitted to it says as much about a test being used up as about capability. The fit runs on a logit scale, which stretches out the last few points before 100%. A move from 94% to 96% counts for as much there as a move from 50% to 60% does lower down, so a score flattening against its ceiling is not read as a slowdown. And once a domain’s latest score clears its bar, the site flags it and labels the date as history rather than as a forecast. The fix for the underlying problem is a harder benchmark, not a cleverer fit.
When the test is small
Mathematics adds a second problem on top of saturation: size. FrontierMath Tier 4 is 41 research-level problems set by working mathematicians, each taking a specialist days, held privately so they cannot leak into training data. When it was published, o3-mini scored 0 of 41.
The latest reading is 87.8%, which is 36 of 41. The bar is 90%, or 37. One problem is worth almost two and a half percentage points, and the same model run twice on a test this size will typically land a couple of problems apart. The gap between the reading and the bar is smaller than the noise in the reading.
The site does not pretend otherwise. The maths page carries a flag saying the domain is at its bar, and that one more correct answer could come from a lucky run as easily as from a better model. A score on a small test is a count, and a count of 36 out of 41 cannot distinguish the next model from this one.
The ceiling ahead for task length
Task length, the measurement this site leads with, has no ceiling in principle. Twice as long is always possible. METR’s task suite does have one, because a suite contains only the tasks someone has built and timed. METR cautions that horizons above about 16 hours are unreliable on its current tasks, and the latest software reading is 17 hours.
So the software domain is heading into the same problem from a different direction. It will not saturate at 100%, but it will run out of tasks long enough to tell the next model from the last one. The bar the forecast uses, about 79 hours, sits well past what the suite can confirm today. A longer task suite is what would fix it, and until one exists the site says the bar is beyond its measurement’s reach.
The pattern across all three is the same. A benchmark is an instrument with a range, and any reading near the edge of that range needs to say so out loud. That is a property of the instrument, and it says nothing either way about the model being measured.
Related notes
Methodology
What these forecasts do not measure
The dated forecasts are not about AGI. Each one is about a single measurement in a single domain. What those numbers cover, and why raising the bar does not widen them.
Oct 2026 · 5 min read
Methodology
Why the forecast gives no single date
A point estimate hides everything that matters. The case for publishing a distribution, and how this one is built.
Oct 2026 · 4 min read
Data
Training compute is still doubling about every five months
The clearest leading indicator hasn't bent, and why it still does not move a single date on this site.
Oct 2026 · 4 min read
See the data
behind
the notes.
Ten domains, each on its own measure, with the gaps published beside the measurements.