How this date is made
- A straight line through 18 measurements of GPQA Diamond accuracy, on a logit scale. It rises 0.54 on that scale each year, and fits the points with r² 0.95.
- The bar is 99.0%, 99% on GPQA Diamond, which leaves no headroom on an exam where domain experts score about 65%. It is a choice, not a measurement: a bar that cuts the remaining error about tenfold costs about 1.9 years.
- Solving for when the line reaches the bar, 5,000 times with the line, the bar and a single model's scatter drawn from their uncertainty, gives 2027, with an 80% interval of 2026–2028.
Latest measurement 94.8% (Gemini 3.7 Flash). Source: Epoch AI · Benchmarking Hub.
How far to trust it
- The band is too narrow. Refit at 13 past cutoffs, its 80% intervals held 57% of the measurements that followed.
- No clear change of pace. A curved fit does not beat the straight line here.
- It is an extrapolation. The bar is 0.66 steps on the fitted scale past the latest reading, about 1.2 years at the fitted pace.
- 7% of the bar's uncertainty range sits at or below today's reading, so it is already reached; the date covers the rest.
The numbers and equations
- y(t) = ȳ + b · (t − t̄)
- y* ~ N(μH*, σH*), above today's reading
- t* = t̄ + ( y* − ȳ ) / b
A draw that would have crossed before the latest measurement, which did not, is dropped (6.1% here). Years are calendar years. No acceleration split has been justified for this series: GPQA's SOTA points sit on one line from GPT-4 onward, with no residual pattern of the kind that forced software's 2023 cut. Every SOTA measurement is fit, and this note says so.
Training compute and efficiency (β_C 0.715, β_E 0.438 a year) are fit and published, but γ₁(β_C + β_E) equals b exactly, so they cancel out of the date.
| b | 0.539 | slope: change in y per calendar year |
| se(b) | 0.043 | its standard error, widened 1.40× for residuals that run in streaks |
| σ | 0.111 | how far a single model sits off the line |
| μH* | 1.90 | the bar on the fitted scale (99.0%) |
| σH* | 0.45 | uncertainty on the bar |
| γ1 | 0.467 | the same slope in compute units, for context |
Questions
GPQA is answered by the same frontier language models Epoch's training-run series describes, so the systems measured and the systems setting the compute frontier are the same class. The reading is close to its ceiling: domain experts score about 65% and the frontier is at 94.8%, so this domain's date is as much about when the exam is exhausted as about capability, and the page says so.
No acceleration split has been justified for this series: GPQA's SOTA points sit on one line from GPT-4 onward, with no residual pattern of the kind that forced software's 2023 cut. Every SOTA measurement is fit, and this note says so.
No. Each domain is measured on its own scale - a task length in hours, an accuracy on a fixed test set - and every one of them scores a model on a bounded, pre-specified set of tasks. That is a measure of DEPTH on that set, not of breadth of competence. Depth is not breadth: a model can clear the bar on GPQA Diamond and still fail at things a child does.