87.8% on FrontierMath T4.
Research-level problems set by working mathematicians, and contest problems below them.
Clears the bar in
2026
80% interval 2026–2026
20252028
90% of FrontierMath Tier 4: nine in ten research-level problems, each a specialist's day of work is the bar this date is solved against.
⚠ FrontierMath Tier 4 accuracy needs 37 of 41 items and the latest reading has 36: 1 more. One run of the same model moves by about ±2 items on a test this size, so this domain is at its bar within measurement noise, and the year below is a forecast of that noise rather than of new capability.
When FrontierMath T4 reaches 90.0%
The year a frontier model first scores 90.0% on FrontierMath Tier 4 accuracy. One benchmark, not Advanced mathematics as a whole.
FrontierMath Tier 4 accuracy needs 37 of 41 items and the latest reading has 36: 1 more. One run of the same model moves by about ±2 items on a test this size, so this domain is at its bar within measurement noise, and the year below is a forecast of that noise rather than of new capability.
The published curve.
The band has held up in backtests so far.How far to trust it
Change the assumptions
The curve redraws in your browser, labelled as yours. Nothing you set here changes the published date.
Today's reading
If you think Advanced mathematics is further along or behind than the latest measurement.
The bar
What score on FrontierMath Tier 4 accuracy counts as done. Widen the spread if you are less sure where it sits.
Pace of the trend
as fittedWhat if progress runs faster or slower than the fitted line?
Questions
FrontierMath is answered by the same frontier language models Epoch's largest-training-run series describes, so the systems measured and the systems setting the compute frontier are the same class. Contest sets such as Mock AIME are solved outright by frontier models, so they have no signal left to fit, which is why this domain is measured on the research-level tier.
The whole Tier 4 series is post-2025, because it did not exist earlier, so there is no earlier era to split off. Every SOTA measurement is fit.
No. Each domain is measured on its own scale - a task length in hours, an accuracy on a fixed test set - and every one of them scores a model on a bounded, pre-specified set of tasks. That is a measure of DEPTH on that set, not of breadth of competence. Depth is not breadth: a model can clear the bar on FrontierMath T4 and still fail at things a child does.