Home/Domains/Software & ML research

17 h, doubling every 125 days.

Writing, debugging and shipping code; running machine-learning experiments.

Clears the bar in

2027

80% interval 2026–2028

20252030

79 h · roughly two working weeks, unattended is the bar this date is solved against.

When a 79-hour task is reached

The year a frontier model first finishes a 79-hour Software & ML research task unattended, half the time. Task length, not breadth of competence.

2025 · 0.0%
2026 · 27.7%
2027 · 60.0%
2028 · 11.6%
2029 · 0.6%
2030 · 0.0%
2031 · 0.0%
2032 · 0.0%
2033 · 0.0%
2034 · 0.0%
202520292034
Median
2027
80% interval
2026–2028
By 2030
100%

The published curve.

Treat the band as too narrow: in backtests it held 29% of later measurements, not 80%. How far to trust it

Change the assumptions

The curve redraws in your browser, labelled as yours. Nothing you set here changes the published date.

Today's reading

If you think Software & ML research is further along or behind than the latest measurement.

Task horizon17 h

The bar

How long a task counts as done. Widen the spread if you are less sure where it sits.

Task length79 h · ±0.45

Pace of the trend

as fitted

What if progress runs faster or slower than the fitted line?

half as fasthalf again as fast

Questions

METR times the same frontier language models that Epoch's largest-training-run series describes, so extrapolating along that line is a claim about the same systems. It is still not evidence that compute causes capability: the frontier level is a function of the date, so this fit is a horizon-vs-time trend expressed in compute units.

METR publish a doubling time of 187.8 days all-time against 128.7 days from 2023 on, and 2023 is where they split it. Fit on everything and the last six points sit above the line by a growing margin, ending +0.51 dex out.

No. Each domain is measured on its own scale - a task length in hours, an accuracy on a fixed test set - and every one of them scores a model on a bounded, pre-specified set of tasks. That is a measure of DEPTH on that set, not of breadth of competence. Depth is not breadth: a model can hold a 79-hour Software & ML research job unattended and still fail at things a child does.