Writing, debugging and shipping code; running machine-learning experiments.
Every domain, and where it stands.
Ten areas of cognitive work. Three have a dated series to fit, and get a forecast. The other seven do not, and each one says why.
Domain coverage
3 of 10 domains carry a date
Graduate-level questions in physics, chemistry and biology.
Research-level problems set by working mathematicians, and contest problems below them.
Within one test run's noise of the bar.
The other 7
A benchmark exists, but no dated series
Someone measures these, but not in a form a trend can be fitted to yet.
- Agentic computer use
Driving a real desktop or browser to finish an errand end to end.
Would be measured on OSWorld success rate, from OSWorld leaderboard. Needs a dated OSWorld / WebArena series. This is the gap most likely to set an AGI floor.
Reported elsewhere: Horizons 40–100× shorter than software, improving at a similar rate.
- Robotic manipulation
Physical tasks in the world: the coffee test, the flat-pack test.
Would be measured on RLBench success rate, from RLBench. Needs a dated RLBench horizon series with human baselines.
Reported elsewhere: Improving far slower than the software cluster.
- Self-driving
Kilometres of unsupervised driving between human interventions.
Would be measured on miles between disengagements, from California DMV disengagement reports. Needs a defensible series; the public tracker is community telemetry, not a controlled benchmark.
Reported elsewhere: About 0.6 doublings per year, roughly five times slower than software.
Nothing anyone maintains
No independent, dated measurement exists, and this site will not invent one.
- Medicine
Diagnosis, treatment planning, and carrying a case over time.
Clinical work carries over weeks and the outcome is often unobservable at the horizon of the task, so "a doctor takes N hours" is not a well-posed baseline for most of it. Existing medical benchmarks score answer accuracy on vignettes, which is a different quantity from how long a case a system can carry unattended.
- Law
Research, drafting, and running a matter to a conclusion.
Billable-hour records exist in enormous volume but are confidential and are not task-decomposed, so the one dataset that would give human baselines cheaply is the one nobody can publish. Public legal benchmarks test retrieval and classification, not how long a matter a system can run.
- Finance
Analysis, modelling, and decisions carried over a horizon.
The work whose value is easiest to measure is also the work whose results are least likely to be shared, and success is confounded with market luck over exactly the horizons that would need measuring. Timed professional baselines exist inside firms and are not published.
The wrong instrument
Data exists, but it measures the wrong thing for this question.
- Video understanding
Following and reasoning about long-form footage.
Video length is used as a proxy for task length, and METR flag that the correlation with difficulty is weak. A two-hour film is not a two-hour task. The instrument is wrong here, not merely missing.