Home/Domains/Scientific reasoning

94.8% on GPQA Diamond.

Graduate-level questions in physics, chemistry and biology.

Clears the bar in

2027

80% interval 2026–2028

20252031

99% on GPQA Diamond, which leaves no headroom on an exam where domain experts score about 65% is the bar this date is solved against.

The measurement

Each dot is one model's GPQA Diamond accuracy, plotted on its release date. Nothing on this chart is a forecast.

Accuracy on GPQA DiamondEpoch AI · Benchmarking Hub ↗
26.8%46.4%67.2%82.9%92.1%96.6%20232024202520262027GPT-4 (Mar 2023) · 35.7% · Mar 2023GPT-4 Turbo (Nov 2023) · 42.4% · Nov 2023Claude 3 Opus · 47.2% · Feb 2024GPT-4o · 48.9% · May 2024Claude 3.5 Sonnet · 54.0% · Jun 2024o1-mini · 62.4% · Sep 2024o1 · 76.8% · Dec 2024o3-mini · 77.0% · Jan 2025Claude 3.7 Sonnet · 79.7% · Feb 2025Gemini 2.5 Pro (Mar 2025) · 83.8% · Mar 2025Gemini 2.5 Pro (Jun 2025) · 84.8% · Jun 2025Gemini 2.5 Pro (Jun 2025) · 85.3% · Jun 2025Grok 4 · 87.0% · Jul 2025GPT-5.1 · 87.6% · Nov 2025Gemini 3 Pro · 92.6% · Nov 2025Gemini 3.1 Pro · 94.4% · Feb 2026GPT-5.4 Pro · 94.6% · Mar 2026Gemini 3.7 Flash · 94.8% · Aug 2026Latest94.8%Gemini 3.7 Flash · Aug 2026
Measured, before the fit windowMeasured (Epoch AI · Benchmarking Hub)Fitted trend · the rows in the fit window
Doubling
~6.7 mo
204 days · all measurements
Fit
r² 0.95
18 of 18 measurements fitted
Measured span
35.7% 94.8%
Mar 2023 → Aug 2026
How to read this chart

The axis is a log-odds scale, the same one the trend is fit on. Equal distance means an equal cut in the remaining error, so 50% to 90% is about the same step as 90% to 99%, which is why the labels crowd near the top. A percentage axis would flatten near 100% whether or not capability flattened. Very close to the ceiling the steps shorten, because a 198-question test cannot resolve past its own last item.

Our fit over all measurements doubles every 204 days. Epoch AI · Benchmarking Hub publish the scores and the release dates. The trend line through them is ours, and it is the only thing on this chart that is.

Measurements

18 models

Every point on the chart above, named. The date is the model's release, which is what the trend is regressed on.

Scientific reasoning measurements, most recent first
Model Released Position on the measured rangeGPQA Diamond
Gemini 3.7 FlashAug 202694.8%
GPT-5.4 ProMar 202694.6%
Gemini 3.1 ProFeb 202694.4%
Gemini 3 ProNov 202592.6%
GPT-5.1Nov 202587.6%
Grok 4Jul 202587.0%
Gemini 2.5 Pro (Jun 2025)Jun 202585.3%
Gemini 2.5 Pro (Jun 2025)Jun 202584.8%

Questions

GPQA is answered by the same frontier language models Epoch's training-run series describes, so the systems measured and the systems setting the compute frontier are the same class. The reading is close to its ceiling: domain experts score about 65% and the frontier is at 94.8%, so this domain's date is as much about when the exam is exhausted as about capability, and the page says so.

No acceleration split has been justified for this series: GPQA's SOTA points sit on one line from GPT-4 onward, with no residual pattern of the kind that forced software's 2023 cut. Every SOTA measurement is fit, and this note says so.

No. Each domain is measured on its own scale - a task length in hours, an accuracy on a fixed test set - and every one of them scores a model on a bounded, pre-specified set of tasks. That is a measure of DEPTH on that set, not of breadth of competence. Depth is not breadth: a model can clear the bar on GPQA Diamond and still fail at things a child does.