Home/Domains/Software & ML research

17 h, doubling every 125 days.

Writing, debugging and shipping code; running machine-learning experiments.

Clears the bar in

2027

80% interval 2026–2028

20252030

79 h · roughly two working weeks, unattended is the bar this date is solved against.

The measurement

Each dot is one model's 50% task horizon, plotted on its release date. Nothing on this chart is a forecast.

Task length finished unattended, half the time · log scale, each gridline ≈ 10×METR · METR-Horizon-v1.1 ↗
1 sec10 sec1 min10 min1.0 h10 h201920202021202220232024202520262027gpt2 · 3 sec · Feb 2019davinci-002 · 9 sec · May 2020gpt-3-5-turbo-instruct · 36 sec · Mar 2022gpt-4 · 4 min · Mar 2023gpt-4-1106 · 4 min · Nov 2023gpt-4o · 7 min · May 2024claude-3-5-sonnet-20240620 · 11 min · Jun 2024o1-preview · 20 min · Sep 2024claude-3-5-sonnet-20241022 · 21 min · Oct 2024o1 · 39 min · Dec 2024claude-3-7-sonnet · 1.0 h · Feb 2025o3 · 2.0 h · Apr 2025gpt-5-2025-08-07 · 3.4 h · Aug 2025gemini-3-pro · 3.7 h · Nov 2025claude-opus-4-5 · 4.9 h · Nov 2025gpt-5-2 · 5.9 h · Dec 2025claude-opus-4-6 · 12 h · Feb 2026claude-mythos-preview-early · 17 h · Apr 2026Latest17 hclaude-mythos-preview-early · Apr 2026
Measured, before the fit windowMeasured (METR · METR-Horizon-v1.1)Fitted trend · the rows in the fit window
Doubling
~4.1 mo
125 days · 2023 onward
Fit
r² 0.94
15 of 18 measurements fitted
Measured span
3 sec 17 h
Feb 2019 → Apr 2026
How to read this chart

Our fit over 2023 onward doubles every 125 days. METR publish 129 days over the same window, and 188 all-time against our 183. We reproduce their number rather than assert our own.

Measurements

18 models

Every point on the chart above, named. The date is the model's release, which is what the trend is regressed on.

Software & ML research measurements, most recent first
Model Released Position on the measured rangeTask horizon
claude-mythos-preview-earlyApr 202617 h
claude-opus-4-6Feb 202612 h
gpt-5-2Dec 20255.9 h
claude-opus-4-5Nov 20254.9 h
gemini-3-proNov 20253.7 h
gpt-5-2025-08-07Aug 20253.4 h
o3Apr 20252.0 h
claude-3-7-sonnetFeb 20251.0 h

Questions

METR times the same frontier language models that Epoch's largest-training-run series describes, so extrapolating along that line is a claim about the same systems. It is still not evidence that compute causes capability: the frontier level is a function of the date, so this fit is a horizon-vs-time trend expressed in compute units.

METR publish a doubling time of 187.8 days all-time against 128.7 days from 2023 on, and 2023 is where they split it. Fit on everything and the last six points sit above the line by a growing margin, ending +0.51 dex out.

No. Each domain is measured on its own scale - a task length in hours, an accuracy on a fixed test set - and every one of them scores a model on a bounded, pre-specified set of tasks. That is a measure of DEPTH on that set, not of breadth of competence. Depth is not breadth: a model can hold a 79-hour Software & ML research job unattended and still fail at things a child does.