Domain coverage

3 of 10 domains carry a date

Today
87.8%
Clears
2026
FrontierMath T4Bar 90.0%

Within one test run's noise of the bar.

No dated series yet for the other 7: Agentic computer use, Robotic manipulation, Self-driving, Video understanding, Medicine, Law, Finance.

Why a full board still is not a date.

Filling every row would raise the floor, not settle the question. Two reasons a complete board would still stop short of a date.

The slowest domain decides

If generality is the claim, the right aggregation is a minimum, not a mean.

A system superhuman at code and helpless at everything else is not general, and an average would hide exactly that. So an AGI date would be set by the slowest domain on the board, and the slowest domains are the ones with the least data. Adding more well-measured coding benchmarks would not move it at all.

A mean rides the one tall bar. The floor is set by the domains nobody has measured.

Breadth is not a domain

Coverage is necessary here, not sufficient.

Most definitions of AGI require transferring to problems nobody enumerated, and a checklist of domains cannot test for that: you can clear any fixed list by building one specialist per domain. Even a full board would not settle it.

Every box ticked, and the list still does not close on the right.

How fast is frontier AI moving?

Four separate measurements, each in its own unit, each tracked publicly by someone else, last read on October 6, 2026. They are not combined into a score: there is no defensible scale on which a training run and a price are the same kind of quantity, and inventing one would smuggle in a weighting nobody could argue with.

From cited data to a dated trend.

Read the
cited sources

Training compute from Epoch AI, task horizons from METR, test scores from Epoch AI’s benchmarking hub, and cost, rating and market readings from OpenRouter, LMArena and Manifold. Every input traces back to its source.

Report each
in its own unit

No shared scale and no composite. The four readings are published as they come, a FLOP count, a price, a rating and a probability, because there is no defensible axis on which those are the same kind of quantity.

Publish & score
the predictions

A measured trend per domain, on that domain’s own measure, with a dated forecast where the data supports one, and public predictions graded true or false as they come due.

The tools moving the numbers.

One standout per layer of the stack: assistant, coding, research, agents, media and safety. Picked by the editors, and none of it feeds a calculation on this site.

OpenAI · Assistant

ChatGPT

The mass-market on-ramp to frontier models: a general assistant for writing, reasoning and code.

Visit site

Anysphere · Coding

Cursor

AI-native code editor that edits across an entire repository.

Visit site

Perplexity AI · Research

Perplexity

Answer engine that cites its sources as it searches the live web.

Visit site

Cognition · Agents

Devin

Autonomous software agent that plans and ships multi-step tasks.

Visit site
01 / 01