Home/Domains/Software & ML research

17 h, doubling every 125 days.

Writing, debugging and shipping code; running machine-learning experiments.

Clears the bar in

2027

80% interval 2026–2028

20252030

79 h · roughly two working weeks, unattended is the bar this date is solved against.

How this date is made

  1. A straight line through 15 measurements of 50% task horizon since 2023, on a log scale. It doubles about every 4.1 months, and fits the points with r² 0.94.
  2. The bar is 79 h, roughly two working weeks, unattended. It is a choice, not a measurement: a task ten times longer costs about 1.1 years.
  3. Solving for when the line reaches the bar, 5,000 times with the line, the bar and a single model's scatter drawn from their uncertainty, gives 2027, with an 80% interval of 2026–2028.

Latest measurement 17 h (claude_mythos_preview_early_inspect). Source: METR · METR-Horizon-v1.1.

How far to trust it

  • The band is too narrow. Refit at 10 past cutoffs, its 80% intervals held 29% of the measurements that followed.
  • The pace has been speeding up. A curved fit that allows for it gives 2026 (80%: 2026–2027).
  • It is an extrapolation. The bar is 4.6× past the latest reading, about 9 months at the fitted pace. METR caution that horizons above about 16 hours are unreliable on their current task suite, and the latest reading is already 17 h. A 79 h bar is about five times past what that suite can measure, so confirming it will need a longer suite, and METR's revision from version 1.0 to 1.1 moved recent readings by up to 20%.
  • 7% of the bar's uncertainty range sits at or below today's reading, so it is already reached; the date covers the rest.
The numbers and equations
  1. y(t) = ȳ + b · (t − t̄)
  2. y* ~ N(μH*, σH*), above today's reading
  3. t* = t̄ + ( y* − ȳ ) / b

A draw that would have crossed before the latest measurement, which did not, is dropped (0.9% here). Years are calendar years. METR publish a doubling time of 187.8 days all-time against 128.7 days from 2023 on, and 2023 is where they split it. Fit on everything and the last six points sit above the line by a growing margin, ending +0.51 dex out.

Training compute and efficiency (β_C 0.715, β_E 0.438 a year) are fit and published, but γ₁(β_C + β_E) equals b exactly, so they cancel out of the date.

b0.882slope: change in y per calendar year
se(b)0.095its standard error, widened 1.54× for residuals that run in streaks
σ0.205how far a single model sits off the line
μH*1.90the bar on the fitted scale (79 h)
σH*0.45uncertainty on the bar
γ10.765the same slope in compute units, for context

Questions

METR times the same frontier language models that Epoch's largest-training-run series describes, so extrapolating along that line is a claim about the same systems. It is still not evidence that compute causes capability: the frontier level is a function of the date, so this fit is a horizon-vs-time trend expressed in compute units.

METR publish a doubling time of 187.8 days all-time against 128.7 days from 2023 on, and 2023 is where they split it. Fit on everything and the last six points sit above the line by a growing margin, ending +0.51 dex out.

No. Each domain is measured on its own scale - a task length in hours, an accuracy on a fixed test set - and every one of them scores a model on a bounded, pre-specified set of tasks. That is a measure of DEPTH on that set, not of breadth of competence. Depth is not breadth: a model can hold a 79-hour Software & ML research job unattended and still fail at things a child does.