Getting the same result, for less.
AI gets better partly because companies spend more on training, and partly because the methods themselves improve. This page separates the two, using 208 published test results, and shows every step of the arithmetic. The number is published as context: no date on this site depends on it, as the methodology explains.
- βE
- 0.438
- powers of ten a year
- in plain terms
- 2.74×
- cheaper every year
- cost halves every
- 8.2
- months
- worked out from
- 208
- real results, 2015 to 2023
Two things make AI better. We want to measure one of them.
Suppose a model this year beats one from three years ago. Some of that is simply money: it was trained on far more computing power. The rest is skill, meaning people found better ways to build and train it. Only the second one is efficiency.
To measure it, you have to hold the money still and see what is left. That is the whole idea, and the picture beside this is what it looks like.
Each dot is one model, tested once. Left to right is how much compute trained it. Up and down is its score, and lower is better.
- Dots fall as you go right, because more compute buys a better score.
- The two lines are the same relationship in 2018 and in 2023. The later one sits lower everywhere.
- That drop, marked by the upright bar between the two lines, happened without extra compute. It is what we are trying to put a number on.
Every result on the WT103 test
score down, compute acrossWait a year and the same result costs 2.74 times less
0.438 is a power of ten, not a multiplier, so turn it back into one: 100.438 = 2.74. A job that needs a million pounds of computing today needs about 364 thousand next year, and about 133 thousand the year after, for the same result.
Put another way, the price of any fixed level of ability halves roughly every 8.2 months. Halving means dividing by 2, and 2 is 100.3010, so the question is how long 0.438 powers of ten a year takes to add up to 0.3010 of them:
0.3010 ÷ 0.438 = 0.6867 years
0.6867 × 12 = 8.2 months
That is faster than computer chips have ever got cheaper on their own, and it is why capability spreads outward so quickly: what only the largest labs could afford three years ago is ordinary now.
Two real models, same test
- 2015.2genCNN + dyn eval
- score 106.3 · trained on 1016.86 of compute
- 2019.9bRSM + cache
- score 103.5 · trained on 1014.33 of compute
16.86 − 14.33 = 2.53 powers of ten
102.53 = 339× less compute
The later one scored a shade better on that much less compute, 4.7 years on. It was picked by rule, not by eye: same test, at least three years apart, the later score no worse, and of every pair meeting that, the widest compute gap. One pair proves nothing by itself. The number above is the same thing averaged over all 208.
Where this goes next: the forecast adds this rate to the rate at which training runs are getting bigger, and calls the total "effective compute". A domain's own measurements are then fitted against that, and the year the fitted line reaches the bar you set is the date the site publishes.
Show the full working
The equation, all 208 results, how they are prepared, the arithmetic step by step, and how sure the number is.
Show the full working
The equation, all 208 results, how they are prepared, the arithmetic step by step, and how sure the number is.
One line that says what a score is made of
A model's score comes from three things: how hard the test is, how much compute it was trained on, and what the field had learned by then. Written down, that is:
Scores and compute are both written as powers of ten, because the numbers involved run from thousands to trillions and nothing else fits on one page. A compute of 20 means 1020, a one followed by twenty zeros.
- P
- the score a model got on a set test. Lower is better here, like a golf score. Its technical name is perplexity.
- C
- how much calculation went into training that model.
- b
- what ten times the compute is worth. One number, shared by every row.
- atest
- how hard each test is. Three tests are used and they are not equally hard.
- gyear
- whatever that year did better that the compute does not explain. This is the part we are after.
We know the scores and we know the compute, because both are published. The two unknowns are b and each year's own gain gyear, and the computer finds the values for them that come closest to matching all 208 results at once. Everything below is what came back.
Every result the answer was built from
208 results covering 202 models, on 3 standard tests, from 2015 to 2023. Nothing is weighted or held back. The table beside this is the entire input.
These are the rows that survived the filtering, which the next section walks through row by row.
What a row has to contain
- system
- which model was tested
- year
- when it came out
- flop
- the compute its training used
- benchmark
- which test the score is from
- perplexity
- the score itself. The column keeps its technical name, which is what this kind of score is called.
Nothing here is stored or typed in by hand. The whole page is worked out again on every update, so new results move the answer on their own.
| model | year | compute | test | score |
|---|---|---|---|---|
| LLaMA-33B (LoRA finetuned) | 2023.39 | 23.48 | PTB | 7.68 |
| LLaMA-13B (LoRA finetuned) | 2023.39 | 22.93 | PTB | 8.64 |
| LLaMA-7B (LoRA finetuned) | 2023.39 | 22.66 | PTB | 9.69 |
| LLaMA-13B (LoRA finetuned) | 2023.39 | 22.93 | WT2 | 5.54 |
| LLaMA-65B (LoRA finetuned) | 2023.39 | 23.78 | WT2 | 4.27 |
| LLaMA-7B (LoRA finetuned) | 2023.39 | 22.66 | WT2 | 6.19 |
| MPT-7B | 2023.34 | 22.62 | WT2 | 9.96 |
| Pythia-12b | 2023.25 | 22.33 | WT2 | 10.54 |
| Pythia-6.9b | 2023.25 | 22.09 | WT2 | 11.41 |
| Pythia-160m | 2023.25 | 20.46 | WT2 | 33.43 |
| Pythia-1b | 2023.25 | 21.26 | WT2 | 16.45 |
| Pythia-1.4b | 2023.25 | 21.40 | WT2 | 14.72 |
| Pythia-410m | 2023.25 | 20.87 | WT2 | 20.11 |
| Pythia-2.8b | 2023.25 | 21.70 | WT2 | 12.69 |
| Sparse Wide GPT-3 Small | 2023.22 | 19.95 | WT103 | 20.4 |
| LLaMA-33B | 2023.16 | 23.48 | WT2 | 6.9 |
| LLaMA-13B | 2023.16 | 22.93 | WT2 | 13.99 |
| LLaMA-7B | 2023.16 | 22.66 | WT2 | 9.49 |
| LLaMA-65B | 2023.16 | 23.78 | WT2 | 4.96 |
| GPT-2+Active-SGD | 2023.06 | 17.49 | WT2 | 20.59 |
| Hybrid H3-355M | 2022.99 | 20.05 | WT103 | 16.9 |
| Hybrid H3-125M | 2022.99 | 19.59 | WT103 | 23.7 |
| Hybrid H3-2.7B | 2022.99 | 20.93 | WT103 | 10.6 |
| Hybrid H3-1.3B | 2022.99 | 20.61 | WT103 | 12.5 |
| Transformer + GFM | 2022.91 | 18.91 | WT103 | 20.05 |
| Mogrifier RLSTM (PTB) | 2022.84 | 16.73 | PTB | 42.9 |
| Mogrifier RLSTM (WT2) | 2022.84 | 17.04 | WT2 | 38 |
| Decaying Fast Weights Transformer | 2022.77 | 19.11 | WT103 | 20.5 |
| NMST+GPT-2 | 2022.75 | 20.08 | WT103 | 20.69 |
| BLOOM-1.7B | 2022.51 | 21.56 | WT2 | 20.17 |
| BLOOM-1B | 2022.51 | 21.35 | WT2 | 23.7 |
| BLOOM-560M | 2022.51 | 21.07 | WT2 | 30.05 |
| BLOOM-3B | 2022.51 | 21.80 | WT2 | 17.57 |
| BLOOM-7.1B | 2022.51 | 22.17 | WT2 | 14.72 |
| OPT-125M (finetuned on PTB) | 2022.47 | 20.35 | PTB | 16.5 |
| OPT-2.7B (finetuned on PTB) | 2022.47 | 21.69 | PTB | 10.8 |
| OPT-1.3B (finetuned on PTB) | 2022.47 | 21.37 | PTB | 12.02 |
| OPT-66B | 2022.47 | 23.08 | WT2 | 9.34 |
| OPT-2.7B (finetuned on WT2) | 2022.47 | 21.69 | WT2 | 10.27 |
| OPT-6.7B | 2022.47 | 22.08 | WT2 | 10.86 |
| OPT-1.3B (finetuned) | 2022.47 | 21.37 | WT2 | 12.22 |
| OPT-2.7B | 2022.47 | 21.69 | WT2 | 12.47 |
| OPT-175B | 2022.47 | 23.62 | WT2 | 8.35 |
| OPT-13B | 2022.47 | 22.37 | WT2 | 10.13 |
| OPT-125M (finetuned) | 2022.47 | 20.35 | WT2 | 19.85 |
| OPT-1.3B | 2022.47 | 21.37 | WT2 | 16.41 |
| OPT-30B | 2022.47 | 22.73 | WT2 | 10.67 |
| OPT-350M | 2022.47 | 20.35 | WT2 | 25.42 |
| DITTO | 2022.43 | 19.04 | WT103 | 24.33 |
| B2T connection (16L) | 2022.41 | 19.45 | WT103 | 19.2 |
| GPT-NeoX-20B | 2022.28 | 22.75 | WT2 | 9.2 |
| LaMemo | 2022.28 | 18.87 | WT103 | 23.77 |
| Monarch-GPT-2-Medium | 2022.25 | 20.64 | WT103 | 20.3 |
| Monarch-GPT-2-Small | 2022.25 | 20.28 | WT103 | 20.7 |
| Chinchilla | 2022.24 | 23.76 | WT103 | 7.16 |
| NoPos | 2022.24 | 20.21 | WT103 | 20.97 |
| Segatron-XL large, M=384 + HCP | 2022.22 | 19.42 | WT103 | 17 |
| Transformer Large + HCP | 2022.22 | 18.78 | WT103 | 25.3 |
| Segatron -XL base, M=150 + HCP | 2022.22 | 18.24 | WT103 | 22.1 |
| MemSizer | 2022.22 | 18.86 | WT103 | 20.8 |
| GPT3-6.7B + muP | 2022.18 | 22.11 | WT103 | 8.56 |
| HSO | 2021.96 | 20.54 | WT103 | 20.3 |
| Gopher (7.1B) | 2021.93 | 23.80 | WT103 | 10.81 |
| Gopher (280B) | 2021.93 | 22.11 | WT103 | 8.12 |
| GPT-2-Medium+Pixelfly | 2021.91 | 19.10 | WT103 | 21 |
| GPT-2-Small+Pixelfly | 2021.91 | 18.62 | WT103 | 22.5 |
| GPT2+CoreLM+Fine-Tuning | 2021.84 | 16.50 | WT103 | 29.51 |
| GPT2+CoreLM+Fine-Tuning | 2021.84 | 16.50 | WT2 | 31.8 |
| S4 | 2021.83 | 19.89 | WT103 | 20.95 |
| GPT-2 (fine-tuned with HYDRA) | 2021.79 | 16.28 | WT2 | 15.17 |
| base LM+GNN+kNN | 2021.79 | 18.86 | WT103 | 16.8 |
| PermuteFormer | 2021.68 | 18.49 | WT103 | 32.49 |
| $\infty$-former (SM) | 2021.67 | 20.08 | WT103 | 16.61 |
| ALiBi (L=3072, Lvalid = 3072) | 2021.65 | 20.26 | WT103 | 18.3 |
| GPT-2 (1.5B, Curriculum Learning 45K) | 2021.61 | 20.78 | WT103 | 13.72 |
| DEQ-Transformer (Post-LN) + Jacobian Regularisation | 2021.49 | 19.46 | WT103 | 24.9 |
| Adaptive Input Transformer + RD | 2021.49 | 19.91 | WT103 | 18.07 |
| GPT-J-6B | 2021.44 | 22.16 | WT2 | 10.88 |
| Delta RNN (+ full context) | 2021.44 | 18.04 | WT103 | 32.8 |
| Transformer-C | 2021.27 | 18.26 | WT103 | 25.1 |
| GPT-Neo-2.7B (finetuned on PTB) | 2021.22 | 21.81 | PTB | 14.7 |
| GPT-Neo-2.7B (finetuned) | 2021.22 | 21.81 | WT2 | 10.78 |
| GPT-Neo-125M(finetuned) | 2021.22 | 20.48 | WT2 | 21.96 |
| GPT-Neo-125M | 2021.22 | 20.48 | WT2 | 32.29 |
| GPT-Neo-2.7B | 2021.22 | 21.81 | WT2 | 11.39 |
| GPT-Neo-1.3B (finetuned) | 2021.22 | 21.81 | WT2 | 12.09 |
| GLM-10B-bidirectional | 2021.21 | 22.58 | WT103 | 11.33 |
| GLM-10B-unidirectional | 2021.21 | 22.58 | WT103 | 12.22 |
| RFA-GATE-Gaussian-Stateful Big | 2021.17 | 18.85 | WT103 | 23.5 |
| SRU++ Large | 2021.15 | 19.04 | WT103 | 17.1 |
| SRU++ Large only 2 attention layers (k=5) | 2021.15 | 18.90 | WT103 | 17.3 |
| SRU++ Base | 2021.15 | 18.76 | WT103 | 18.3 |
| Linear Transformer (large) | 2021.14 | 18.59 | WT103 | 31.5 |
| Linear Transformer (small) | 2021.14 | 18.47 | WT103 | 35.5 |
| Selfish-RNN (ON-LSTM) | 2021.06 | 16.80 | PTB | 55.82 |
| Selfish-RNN (SNT-ASGD) Stacked LSTMs | 2021.06 | 16.15 | PTB | 71.42 |
| Selfish-RNN (SNT-ASGD)RHNs | 2021.06 | 16.33 | PTB | 64.03 |
| Selfish-RNN (AWD-LSTM-MoS) | 2021.06 | 17.29 | WT2 | 63.05 |
| Shortformer | 2021.00 | 18.48 | WT103 | 18.15 |
| ERNIE-Doc (151M) | 2021.00 | 19.25 | WT103 | 21 |
| ERNIE-Doc (247M) | 2021.00 | 19.46 | WT103 | 16.8 |
| Subformer (83M) | 2021.00 | 18.56 | WT103 | 20.88 |
| Subformer (122M) | 2021.00 | 18.72 | WT103 | 19.9 |
| Subformer (96M) | 2021.00 | 18.62 | WT103 | 20.39 |
| CT-MoS (PTB) | 2020.98 | 17.13 | PTB | 54.69 |
| CT-MoS + DynamicEval (PTB) | 2020.98 | 17.13 | PTB | 47.42 |
| CT-MoS + DynamicEval (WT2) | 2020.98 | 17.75 | WT2 | 40.96 |
| CT-MoS (WT2) | 2020.98 | 17.75 | WT2 | 62.21 |
| AWD-FWM (PTB) | 2020.88 | 17.13 | PTB | 54.48 |
| AWD-FWM (WT2) | 2020.88 | 17.87 | WT2 | 61.65 |
| Transformer+Recurrent Windows of Context | 2020.62 | 20.07 | WT103 | 26.73 |
| DeLight | 2020.59 | 19.38 | WT103 | 24.14 |
| 3-Layer-Tensor-Transformer+AdaHessian | 2020.42 | 15.30 | PTB | 51.5 |
| 6-Layer-Tensor-Transformer+AdaHessian | 2020.42 | 18.20 | WT103 | 19.9 |
| GPT3-6.7B (rerun of original) | 2020.41 | 22.08 | WT103 | 9.13 |
| rTop-k(distributed setting) | 2020.39 | 16.16 | PTB | 82.49 |
| ONLSTM-SYD | 2020.36 | 17.14 | PTB | 55.7 |
| Segatron XL base, M=384 | 2020.33 | 18.24 | WT103 | 22.5 |
| Segatron XL large, M=384 | 2020.33 | 19.42 | WT103 | 17.1 |
| DiffStk-MRNN | 2020.26 | 14.45 | PTB | 115 |
| Tensor-Transformer(1core)+PN (PTB) | 2020.21 | 15.30 | PTB | 47.6 |
| Tensor-Transformer(1core)+PN (WT103) | 2020.21 | 18.20 | WT103 | 17.9 |
| TransformerXL + spectrum control | 2020.19 | 17.66 | WT103 | 23.2 |
| LSTM-3-layer+Gadam | 2020.17 | 16.43 | PTB | 58.77 |
| Feedback Transformer | 2020.14 | 19.32 | WT103 | 18.3 |
| Turing-NLG | 2020.12 | 22.20 | WT103 | 10.21 |
| TaLK Convolution | 2020.10 | 19.44 | WT103 | 23.3 |
| bRSM + cache | 2019.92 | 14.33 | PTB | 103.5 |
| AWD-LSTM + DeFINE | 2019.90 | 15.35 | PTB | 54.2 |
| Transformer-XL DeFINE (107M) | 2019.90 | 18.72 | WT103 | 25.72 |
| Adaptive LSTM + DeFINE | 2019.90 | 18.79 | WT103 | 35.94 |
| Transformer-XL DeFINE (141M) | 2019.90 | 18.79 | WT103 | 24.17 |
| Compressive Transformers for Long-Range Sequence Modelling | 2019.87 | 20.20 | WT103 | 17.1 |
| Sandwich Transformer | 2019.86 | 20.20 | WT103 | 17.84 |
| LSTM(medium)+Sememe+cell | 2019.80 | 15.70 | WT2 | 89.16 |
| Megatron-LM (8.3B) | 2019.71 | 21.96 | WT103 | 10.81 |
| Megatron-LM (355M) | 2019.71 | 20.64 | WT103 | 19.31 |
| DEQ-TrellisNet | 2019.67 | 17.91 | PTB | 57.1 |
| DEQ-Transformer (Medium, Adaptive Embedding) | 2019.67 | 17.91 | WT103 | 23.2 |
| R-Transformer | 2019.53 | 15.92 | PTB | 84.38 |
| All-attention network + adaptive span | 2019.50 | 19.66 | WT103 | 20.6 |
| Tensorized Transformer (small) | 2019.48 | 15.30 | PTB | 57.9 |
| Tensorized Transformer (large PTB) | 2019.48 | 15.60 | PTB | 52.7 |
| Tensorized Transformer (257M) | 2019.48 | 18.68 | WT103 | 21.2 |
| Tensorized Transformer (core-2) | 2019.48 | 18.20 | WT103 | 18.9 |
| Tensorized Transformer (151M) | 2019.48 | 18.45 | WT103 | 18.8 |
| Adversarial + AWD-LSTM-MoS + partial shuffled | 2019.44 | 16.74 | PTB | 46.01 |
| AdvSoft + 4 layer QRNN + dynamic evaluation | 2019.44 | 17.56 | WT103 | 28 |
| 4 layer QRNN + dynamic evaluation | 2019.44 | 17.56 | WT103 | 31.6 |
| AWD-LSTM + MoS + Partial Shuffled | 2019.44 | 17.52 | WT2 | 38.07 |
| Transformer-XL Large + Phrase Induction | 2019.42 | 17.20 | WT103 | 17.4 |
| AWD-LSTM-DRILL + dynamic evaluation† (PTB) | 2019.36 | 17.13 | PTB | 49.4 |
| AWD-LSTM-DRILL + dynamic evaluation† (WT2) | 2019.36 | 17.63 | WT2 | 42 |
| GPT-2 (1542M) | 2019.12 | 21.18 | PTB | 35.76 |
| GPT-2 (1542M) | 2019.12 | 21.18 | WT103 | 17.48 |
| GPT-2 (762M) | 2019.12 | 20.88 | WT103 | 22.05 |
| GPT-2 (345M) | 2019.12 | 20.54 | WT103 | 26.37 |
| GPT-2 (1542M) | 2019.12 | 21.18 | WT2 | 18.34 |
| Transformer-XL-ptb | 2019.02 | 19.04 | PTB | 54.52 |
| Transformer-XL Large | 2019.02 | 19.04 | WT103 | 18.3 |
| Multi-cell LSTM | 2018.87 | 15.30 | PTB | 77.12 |
| Fine-tuned-AWD-LSTM-DOC(fin) | 2018.86 | 15.28 | PTB | 52.12 |
| TrellisNet-MoS (1.4x larger) | 2018.79 | 18.44 | PTB | 54.19 |
| TrellisNet | 2018.79 | 18.44 | WT103 | 29.19 |
| TrellisNet-MoS (1.4x larger) | 2018.79 | 18.44 | WT103 | 29.19 |
| Transformer (Adaptive Input Embeddings) | 2018.74 | 18.86 | WT103 | 18.7 |
| LSTM+NeuralCache | 2018.73 | 15.01 | WT2 | 66.2 |
| AWD-LSTM-DOC (fin) (23M) | 2018.66 | 16.94 | PTB | 52.38 |
| AWD-LSTM-DOC (fin) (37M) | 2018.66 | 17.14 | WT2 | 58.03 |
| AWD-LSTM-MoS+PDR + dynamic evaluation (PTB) | 2018.62 | 17.15 | PTB | 47.3 |
| aLSTM(depth-2)+RecurrentPolicy (PTB) | 2018.39 | 16.38 | PTB | 55.3 |
| aLSTM(depth-2)+RecurrentPolicy (WT2) | 2018.39 | 16.88 | WT2 | 64.5 |
| AWD-LSTM-MoS+Noisin+dynamic evaluation | 2018.33 | 16.69 | PTB | 47.6 |
| Dropout-LSTM+Noise(Bernoulli) (PTB) | 2018.33 | 16.76 | PTB | 66.1 |
| LSTM+Noise(Beta) | 2018.33 | 17.10 | WT2 | 82.9 |
| Dropout-LSTM+Noise(Laplace) | 2018.33 | 16.51 | WT2 | 82.1 |
| Dropout-LSTM+Noise(Bernoulli) (WT2) | 2018.33 | 17.10 | WT2 | 76.8 |
| LSTM (Hebbian, Cache, MbPA) | 2018.23 | 19.38 | WT103 | 29.2 |
| 4 layer QRNN (h=2500) | 2018.22 | 17.38 | WT103 | 33 |
| QRNN | 2018.08 | 17.56 | WT103 | 33 |
| RNNLM + Dynamic KL Regularization | 2018.00 | 15.49 | PTB | 77.8 |
| RNNLM + Dynamic KL Regularization (WT2) | 2018.00 | 16.34 | WT2 | 86.8 |
| AWD-LSTM-MoS + dynamic evaluation (PTB, 2017) | 2017.86 | 17.09 | PTB | 47.69 |
| AWD-LSTM-MoS + dynamic evaluation (WT2, 2017) | 2017.86 | 17.64 | WT2 | 40.68 |
| Fraternal dropout + AWD-LSTM 3-layer (PTB) | 2017.83 | 16.84 | PTB | 56.8 |
| Fraternal dropout + AWD-LSTM 3-layer (WT2) | 2017.83 | 17.34 | WT2 | 64.1 |
| AWD-LSTM+WT+Cache+IOG (PTB) | 2017.73 | 14.92 | PTB | 53 |
| AWD-LSTM+WT+Cache+IOG (WT2) | 2017.73 | 15.52 | WT2 | 51.7 |
| GL-LWGC-AWD-MoS-LSTM + dynamic evaluation (PTB) | 2017.66 | 17.16 | PTB | 46.34 |
| GL-LWGC-AWD-MoS-LSTM + dynamic evaluation (WT2) | 2017.66 | 17.68 | WT2 | 40.46 |
| EI-REHN-1200D | 2017.62 | 15.83 | PTB | 66.2 |
| EI-REHN-1000D | 2017.62 | 16.03 | PTB | 68.7 |
| AWD-LSTM - 3-layer LSTM (tied) + continuous cache pointer (PTB) | 2017.60 | 16.83 | PTB | 52.8 |
| AWD-LSTM - 3-layer LSTM (tied) + continuous cache pointer (WT2) | 2017.60 | 17.49 | WT2 | 52 |
| 4 layer Densely Connected LSTM | 2017.54 | 15.89 | PTB | 76.8 |
| Densely Connected LSTM + Var. Dropout | 2017.54 | 16.11 | PTB | 78.3 |
| GCRN-M1, dropout | 2016.97 | 15.48 | PTB | 98.67 |
| VD-LSTM+REAL Large | 2016.84 | 16.33 | PTB | 68.5 |
| Pointer Sentinel-LSTM (medium) | 2016.74 | 15.87 | PTB | 70.9 |
| Zoneout + Variational LSTM (PTB) | 2016.74 | 15.87 | PTB | 80.6 |
| Zoneout + Variational LSTM (WT2) | 2016.74 | 16.23 | WT2 | 100.9 |
| Pointer Sentinel-LSTM | 2016.74 | 16.20 | WT2 | 80.8 |
| Variational RHN + WT | 2016.53 | 15.41 | PTB | 65.4 |
| VD-RHN | 2016.53 | 15.55 | PTB | 68.5 |
| Variational (untied weights, MC) LSTM (Large) | 2015.96 | 15.75 | PTB | 73.4 |
| LSTM-Char-Large | 2015.65 | 15.42 | PTB | 78.9 |
| Search-Proven Best LSTM | 2015.51 | 15.52 | PTB | 79.83 |
| genCNN + dyn eval | 2015.21 | 16.86 | PTB | 106.3 |
Turning one row of that table into terms of the equation
The file gives four things about each result. The equation asks for the same four in a different shape. Nothing is added, estimated or filled in: every line below is one column rewritten. Here is the newest WT103 row doing it.
- what the file has
- score 20.4
- what the equation needs
- log10 P
- worked out
- log10 20.4 = 1.310
- what the file has
- training compute, 1019.95 FLOP
- what the equation needs
- log10 C
- worked out
- the power itself: 19.95
- what the file has
- dated 2023.22
- what the equation needs
- gyear
- worked out
- the whole part: 2023, so this row uses g2023
- what the file has
- test WT103
- what the equation needs
- atest
- worked out
- one of 3 baselines, so this row uses aWT103
Doing that to every row leaves one number per column and nothing else. Not every row gets through, and this is the whole of what is dropped:
- fetched from the source212
- every published result
- with both a compute and a score212
- a row missing either cannot be placed on the chart
- in years holding 3 results or more208
- a year holding one model tells you about that model, not about the year
That last rule drops 4 rows across 2012, 2013, 2014, and it is the one choice on this page that moves the answer. Section 08 says what the number would be without it.
What is left is 208 rows, each carrying a score, a compute, a year and a test. That is one equation per row, and 208 of them to solve at once.
Four steps. Each one: the data, the equation, the result.
Step 1 · finding b
What ten times the compute is worth
The data
Every row in a group was run on the same test in the same year, so atest and gyear are one shared number for all of them. Neither has to be known yet.
The equation
x is a row's compute measured from its own group's average, and y is its score measured the same way. Σ means "add up the column", over every row of every group. Both columns are worked out below.
The result
−0.0826
what ten times the compute does to the score, as a power of ten: 10−0.0826 = 0.827, so the score drops by about 17%.
One b, shared by every row and every year. Step 2 takes it back off the data.
Why the rows are grouped first
Rows on the same test in the same year carry the same atest and the same gyear, so inside a group those two are a single constant. Call it k. Every row in the group then reads log10 P = k + b * log10 C, and taking the group's average off both sides removes k entirely. What is left is b, in the last two columns.
So every group has its own b, printed in its heading. They do not agree, and they do not count equally: a group whose models all trained on much the same compute has an x² total near zero and a slope that is mostly noise, so the totals at the foot weight it to nothing. That weight is in the heading too. Years read top to bottom, with the 3 tests side by side inside each one.
| model | its equation | x = log C − group avg | y = log P − group avg | x * y | x² |
|---|---|---|---|---|---|
PTB · 20154 rows shared k = aPTB + g2015 avg log C 15.89 avg log P 1.923 | b = 0.09740.3% | ||||
| LSTM-Char-Large | 1.897 = k + b * 15.42 | −0.467 | −0.026 | 0.0121 | 0.2186 |
| Search-Proven Best LSTM | 1.902 = k + b * 15.52 | −0.367 | −0.021 | 0.0076 | 0.1351 |
| Variational (untied weights, MC) LSTM (Large) | 1.866 = k + b * 15.75 | −0.137 | −0.057 | 0.0079 | 0.0189 |
| genCNN + dyn eval | 2.027 = k + b * 16.86 | 0.973 | 0.104 | 0.1008 | 0.9458 |
| added up over 4 rows | 0.1283 | 1.3183 | |||
b = 0.1283 ÷ 1.3183 = 0.0974 weight = 1.3183 ÷ 401.0633 = 0.0033 | |||||
PTB · 20166 rows shared k = aPTB + g2016 avg log C 15.75 avg log P 1.873 | b = −0.04390.1% | ||||
| Variational RHN + WT | 1.816 = k + b * 15.41 | −0.342 | −0.057 | 0.0196 | 0.1167 |
| GCRN-M1, dropout | 1.994 = k + b * 15.48 | −0.272 | 0.121 | −0.0329 | 0.0738 |
| VD-RHN | 1.836 = k + b * 15.55 | −0.202 | −0.037 | 0.0075 | 0.0407 |
| Pointer Sentinel-LSTM (medium) | 1.851 = k + b * 15.87 | 0.118 | −0.022 | −0.0026 | 0.0140 |
| Zoneout + Variational LSTM (PTB) | 1.906 = k + b * 15.87 | 0.118 | 0.033 | 0.0039 | 0.0140 |
| VD-LSTM+REAL Large | 1.836 = k + b * 16.33 | 0.578 | −0.037 | −0.0216 | 0.3345 |
| added up over 6 rows | −0.0261 | 0.5937 | |||
b = −0.0261 ÷ 0.5937 = −0.0439 weight = 0.5937 ÷ 401.0633 = 0.0015 | |||||
WT2 · 20162 rows shared k = aWT2 + g2016 avg log C 16.21 avg log P 1.956 | b = 3.21600.0% | ||||
| Pointer Sentinel-LSTM | 1.907 = k + b * 16.20 | −0.015 | −0.048 | 0.0007 | 0.0002 |
| Zoneout + Variational LSTM (WT2) | 2.004 = k + b * 16.23 | 0.015 | 0.048 | 0.0007 | 0.0002 |
| added up over 2 rows | 0.0014 | 0.0005 | |||
b = 0.0014 ÷ 0.0005 = 3.2160 weight = 0.0005 ÷ 401.0633 = 0.0000 | |||||
PTB · 20179 rows shared k = aPTB + g2017 avg log C 16.30 avg log P 1.776 | b = −0.05651.1% | ||||
| AWD-LSTM+WT+Cache+IOG (PTB) | 1.724 = k + b * 14.92 | −1.380 | −0.052 | 0.0712 | 1.9044 |
| EI-REHN-1200D | 1.821 = k + b * 15.83 | −0.470 | 0.045 | −0.0212 | 0.2209 |
| 4 layer Densely Connected LSTM | 1.885 = k + b * 15.89 | −0.410 | 0.110 | −0.0449 | 0.1681 |
| EI-REHN-1000D | 1.837 = k + b * 16.03 | −0.270 | 0.061 | −0.0165 | 0.0729 |
| Densely Connected LSTM + Var. Dropout | 1.894 = k + b * 16.11 | −0.190 | 0.118 | −0.0224 | 0.0361 |
| AWD-LSTM - 3-layer LSTM (tied) + continuous cache pointer (PTB) | 1.723 = k + b * 16.83 | 0.530 | −0.053 | −0.0282 | 0.2809 |
| Fraternal dropout + AWD-LSTM 3-layer (PTB) | 1.754 = k + b * 16.84 | 0.540 | −0.021 | −0.0116 | 0.2916 |
| AWD-LSTM-MoS + dynamic evaluation (PTB, 2017) | 1.678 = k + b * 17.09 | 0.790 | −0.097 | −0.0770 | 0.6241 |
| GL-LWGC-AWD-MoS-LSTM + dynamic evaluation (PTB) | 1.666 = k + b * 17.16 | 0.860 | −0.110 | −0.0945 | 0.7396 |
| added up over 9 rows | −0.2451 | 4.3386 | |||
b = −0.2451 ÷ 4.3386 = −0.0565 weight = 4.3386 ÷ 401.0633 = 0.0108 | |||||
WT2 · 20175 rows shared k = aWT2 + g2017 avg log C 17.13 avg log P 1.691 | b = −0.02720.8% | ||||
| AWD-LSTM+WT+Cache+IOG (WT2) | 1.713 = k + b * 15.52 | −1.614 | 0.023 | −0.0370 | 2.6050 |
| Fraternal dropout + AWD-LSTM 3-layer (WT2) | 1.807 = k + b * 17.34 | 0.206 | 0.116 | 0.0240 | 0.0424 |
| AWD-LSTM - 3-layer LSTM (tied) + continuous cache pointer (WT2) | 1.716 = k + b * 17.49 | 0.356 | 0.025 | 0.0091 | 0.1267 |
| AWD-LSTM-MoS + dynamic evaluation (WT2, 2017) | 1.609 = k + b * 17.64 | 0.506 | −0.081 | −0.0411 | 0.2560 |
| GL-LWGC-AWD-MoS-LSTM + dynamic evaluation (WT2) | 1.607 = k + b * 17.68 | 0.546 | −0.084 | −0.0456 | 0.2981 |
| added up over 5 rows | −0.0907 | 3.3283 | |||
b = −0.0907 ÷ 3.3283 = −0.0272 weight = 3.3283 ÷ 401.0633 = 0.0083 | |||||
PTB · 20189 rows shared k = aPTB + g2018 avg log C 16.49 avg log P 1.763 | b = −0.04192.1% | ||||
| Fine-tuned-AWD-LSTM-DOC(fin) | 1.717 = k + b * 15.28 | −1.212 | −0.046 | 0.0553 | 1.4695 |
| Multi-cell LSTM | 1.887 = k + b * 15.30 | −1.192 | 0.125 | −0.1485 | 1.4214 |
| RNNLM + Dynamic KL Regularization | 1.891 = k + b * 15.49 | −1.002 | 0.128 | −0.1286 | 1.0044 |
| aLSTM(depth-2)+RecurrentPolicy (PTB) | 1.743 = k + b * 16.38 | −0.112 | −0.020 | 0.0022 | 0.0126 |
| AWD-LSTM-MoS+Noisin+dynamic evaluation | 1.678 = k + b * 16.69 | 0.198 | −0.085 | −0.0168 | 0.0391 |
| Dropout-LSTM+Noise(Bernoulli) (PTB) | 1.820 = k + b * 16.76 | 0.268 | 0.058 | 0.0154 | 0.0717 |
| AWD-LSTM-DOC (fin) (23M) | 1.719 = k + b * 16.94 | 0.448 | −0.043 | −0.0195 | 0.2005 |
| AWD-LSTM-MoS+PDR + dynamic evaluation (PTB) | 1.675 = k + b * 17.15 | 0.658 | −0.088 | −0.0577 | 0.4327 |
| TrellisNet-MoS (1.4x larger) | 1.734 = k + b * 18.44 | 1.948 | −0.029 | −0.0559 | 3.7938 |
| added up over 9 rows | −0.3541 | 8.4458 | |||
b = −0.3541 ÷ 8.4458 = −0.0419 weight = 8.4458 ÷ 401.0633 = 0.0211 | |||||
WT103 · 20186 rows shared k = aWT103 + g2018 avg log C 18.34 avg log P 1.451 | b = −0.06640.7% | ||||
| 4 layer QRNN (h=2500) | 1.519 = k + b * 17.38 | −0.963 | 0.068 | −0.0652 | 0.9280 |
| QRNN | 1.519 = k + b * 17.56 | −0.783 | 0.068 | −0.0531 | 0.6136 |
| TrellisNet | 1.465 = k + b * 18.44 | 0.097 | 0.014 | 0.0014 | 0.0093 |
| TrellisNet-MoS (1.4x larger) | 1.465 = k + b * 18.44 | 0.097 | 0.014 | 0.0014 | 0.0093 |
| Transformer (Adaptive Input Embeddings) | 1.272 = k + b * 18.86 | 0.517 | −0.179 | −0.0925 | 0.2669 |
| LSTM (Hebbian, Cache, MbPA) | 1.465 = k + b * 19.38 | 1.037 | 0.015 | 0.0151 | 1.0747 |
| added up over 6 rows | −0.1928 | 2.9019 | |||
b = −0.1928 ÷ 2.9019 = −0.0664 weight = 2.9019 ÷ 401.0633 = 0.0072 | |||||
WT2 · 20187 rows shared k = aWT2 + g2018 avg log C 16.58 avg log P 1.864 | b = 0.00380.9% | ||||
| LSTM+NeuralCache | 1.821 = k + b * 15.01 | −1.573 | −0.044 | 0.0685 | 2.4739 |
| RNNLM + Dynamic KL Regularization (WT2) | 1.939 = k + b * 16.34 | −0.243 | 0.074 | −0.0180 | 0.0590 |
| Dropout-LSTM+Noise(Laplace) | 1.914 = k + b * 16.51 | −0.073 | 0.050 | −0.0036 | 0.0053 |
| aLSTM(depth-2)+RecurrentPolicy (WT2) | 1.810 = k + b * 16.88 | 0.297 | −0.055 | −0.0163 | 0.0883 |
| LSTM+Noise(Beta) | 1.919 = k + b * 17.10 | 0.517 | 0.054 | 0.0280 | 0.2674 |
| Dropout-LSTM+Noise(Bernoulli) (WT2) | 1.885 = k + b * 17.10 | 0.517 | 0.021 | 0.0108 | 0.2674 |
| AWD-LSTM-DOC (fin) (37M) | 1.764 = k + b * 17.14 | 0.557 | −0.101 | −0.0561 | 0.3104 |
| added up over 7 rows | 0.0133 | 3.4717 | |||
b = 0.0133 ÷ 3.4717 = 0.0038 weight = 3.4717 ÷ 401.0633 = 0.0087 | |||||
PTB · 201910 rows shared k = aPTB + g2019 avg log C 16.85 avg log P 1.756 | b = −0.04389.5% | ||||
| bRSM + cache | 2.015 = k + b * 14.33 | −2.520 | 0.259 | −0.6518 | 6.3504 |
| Tensorized Transformer (small) | 1.763 = k + b * 15.30 | −1.550 | 0.006 | −0.0099 | 2.4025 |
| AWD-LSTM + DeFINE | 1.734 = k + b * 15.35 | −1.500 | −0.022 | 0.0334 | 2.2500 |
| Tensorized Transformer (large PTB) | 1.722 = k + b * 15.60 | −1.250 | −0.034 | 0.0431 | 1.5625 |
| R-Transformer | 1.926 = k + b * 15.92 | −0.930 | 0.170 | −0.1581 | 0.8649 |
| Adversarial + AWD-LSTM-MoS + partial shuffled | 1.663 = k + b * 16.74 | −0.110 | −0.093 | 0.0103 | 0.0121 |
| AWD-LSTM-DRILL + dynamic evaluation† (PTB) | 1.694 = k + b * 17.13 | 0.280 | −0.063 | −0.0175 | 0.0784 |
| DEQ-TrellisNet | 1.757 = k + b * 17.91 | 1.060 | 0.000 | 0.0004 | 1.1236 |
| Transformer-XL-ptb | 1.737 = k + b * 19.04 | 2.190 | −0.020 | −0.0432 | 4.7961 |
| GPT-2 (1542M) | 1.553 = k + b * 21.18 | 4.330 | −0.203 | −0.8785 | 18.7489 |
| added up over 10 rows | −1.6718 | 38.1894 | |||
b = −1.6718 ÷ 38.1894 = −0.0438 weight = 38.1894 ÷ 401.0633 = 0.0952 | |||||
WT103 · 201919 rows shared k = aWT103 + g2019 avg log C 19.27 avg log P 1.325 | b = −0.04348.4% | ||||
| Transformer-XL Large + Phrase Induction | 1.241 = k + b * 17.20 | −2.072 | −0.084 | 0.1740 | 4.2914 |
| AdvSoft + 4 layer QRNN + dynamic evaluation | 1.447 = k + b * 17.56 | −1.712 | 0.123 | −0.2099 | 2.9295 |
| 4 layer QRNN + dynamic evaluation | 1.500 = k + b * 17.56 | −1.712 | 0.175 | −0.2998 | 2.9295 |
| DEQ-Transformer (Medium, Adaptive Embedding) | 1.365 = k + b * 17.91 | −1.362 | 0.041 | −0.0557 | 1.8539 |
| Tensorized Transformer (core-2) | 1.276 = k + b * 18.20 | −1.072 | −0.048 | 0.0515 | 1.1483 |
| Tensorized Transformer (151M) | 1.274 = k + b * 18.45 | −0.822 | −0.050 | 0.0414 | 0.6750 |
| Tensorized Transformer (257M) | 1.326 = k + b * 18.68 | −0.592 | 0.002 | −0.0011 | 0.3500 |
| Transformer-XL DeFINE (107M) | 1.410 = k + b * 18.72 | −0.552 | 0.086 | −0.0473 | 0.3042 |
| Adaptive LSTM + DeFINE | 1.556 = k + b * 18.79 | −0.482 | 0.231 | −0.1113 | 0.2319 |
| Transformer-XL DeFINE (141M) | 1.383 = k + b * 18.79 | −0.482 | 0.059 | −0.0283 | 0.2319 |
| Transformer-XL Large | 1.262 = k + b * 19.04 | −0.232 | −0.062 | 0.0144 | 0.0536 |
| All-attention network + adaptive span | 1.314 = k + b * 19.66 | 0.388 | −0.011 | −0.0041 | 0.1509 |
| Sandwich Transformer | 1.251 = k + b * 20.20 | 0.928 | −0.073 | −0.0679 | 0.8620 |
| Compressive Transformers for Long-Range Sequence Modelling | 1.233 = k + b * 20.20 | 0.928 | −0.092 | −0.0850 | 0.8620 |
| GPT-2 (345M) | 1.421 = k + b * 20.54 | 1.268 | 0.097 | 0.1225 | 1.6089 |
| Megatron-LM (355M) | 1.286 = k + b * 20.64 | 1.368 | −0.039 | −0.0530 | 1.8726 |
| GPT-2 (762M) | 1.343 = k + b * 20.88 | 1.608 | 0.019 | 0.0303 | 2.5870 |
| GPT-2 (1542M) | 1.243 = k + b * 21.18 | 1.908 | −0.082 | −0.1565 | 3.6421 |
| Megatron-LM (8.3B) | 1.034 = k + b * 21.96 | 2.688 | −0.291 | −0.7816 | 7.2276 |
| added up over 19 rows | −1.4673 | 33.8123 | |||
b = −1.4673 ÷ 33.8123 = −0.0434 weight = 33.8123 ÷ 401.0633 = 0.0843 | |||||
WT2 · 20194 rows shared k = aWT2 + g2019 avg log C 18.01 avg log P 1.604 | b = −0.11893.9% | ||||
| LSTM(medium)+Sememe+cell | 1.950 = k + b * 15.70 | −2.308 | 0.346 | −0.7980 | 5.3246 |
| AWD-LSTM + MoS + Partial Shuffled | 1.581 = k + b * 17.52 | −0.488 | −0.024 | 0.0116 | 0.2377 |
| AWD-LSTM-DRILL + dynamic evaluation† (WT2) | 1.623 = k + b * 17.63 | −0.378 | 0.019 | −0.0071 | 0.1425 |
| GPT-2 (1542M) | 1.263 = k + b * 21.18 | 3.172 | −0.341 | −1.0817 | 10.0648 |
| added up over 4 rows | −1.8752 | 15.7695 | |||
b = −1.8752 ÷ 15.7695 = −0.1189 weight = 15.7695 ÷ 401.0633 = 0.0393 | |||||
PTB · 20209 rows shared k = aPTB + g2020 avg log C 16.24 avg log P 1.781 | b = −0.06772.0% | ||||
| DiffStk-MRNN | 2.061 = k + b * 14.45 | −1.791 | 0.279 | −0.5004 | 3.2081 |
| Tensor-Transformer(1core)+PN (PTB) | 1.678 = k + b * 15.30 | −0.941 | −0.104 | 0.0976 | 0.8857 |
| 3-Layer-Tensor-Transformer+AdaHessian | 1.712 = k + b * 15.30 | −0.941 | −0.069 | 0.0654 | 0.8857 |
| rTop-k(distributed setting) | 1.916 = k + b * 16.16 | −0.081 | 0.135 | −0.0110 | 0.0066 |
| LSTM-3-layer+Gadam | 1.769 = k + b * 16.43 | 0.189 | −0.012 | −0.0023 | 0.0357 |
| AWD-FWM (PTB) | 1.736 = k + b * 17.13 | 0.889 | −0.045 | −0.0400 | 0.7901 |
| CT-MoS (PTB) | 1.738 = k + b * 17.13 | 0.889 | −0.043 | −0.0386 | 0.7901 |
| CT-MoS + DynamicEval (PTB) | 1.676 = k + b * 17.13 | 0.889 | −0.105 | −0.0936 | 0.7901 |
| ONLSTM-SYD | 1.746 = k + b * 17.14 | 0.899 | −0.035 | −0.0319 | 0.8080 |
| added up over 9 rows | −0.5548 | 8.2001 | |||
b = −0.5548 ÷ 8.2001 = −0.0677 weight = 8.2001 ÷ 401.0633 = 0.0204 | |||||
WT103 · 202014 rows shared k = aWT103 + g2020 avg log C 19.39 avg log P 1.266 | b = −0.07335.9% | ||||
| TransformerXL + spectrum control | 1.365 = k + b * 17.66 | −1.726 | 0.100 | −0.1724 | 2.9781 |
| Tensor-Transformer(1core)+PN (WT103) | 1.253 = k + b * 18.20 | −1.186 | −0.013 | 0.0151 | 1.4059 |
| 6-Layer-Tensor-Transformer+AdaHessian | 1.299 = k + b * 18.20 | −1.186 | 0.033 | −0.0395 | 1.4059 |
| Segatron XL base, M=384 | 1.352 = k + b * 18.24 | −1.146 | 0.087 | −0.0992 | 1.3127 |
| Shortformer | 1.259 = k + b * 18.48 | −0.906 | −0.007 | 0.0061 | 0.8203 |
| ERNIE-Doc (151M) | 1.322 = k + b * 19.25 | −0.136 | 0.057 | −0.0077 | 0.0184 |
| Feedback Transformer | 1.262 = k + b * 19.32 | −0.066 | −0.003 | 0.0002 | 0.0043 |
| DeLight | 1.383 = k + b * 19.38 | −0.006 | 0.117 | −0.0007 | 0.0000 |
| Segatron XL large, M=384 | 1.233 = k + b * 19.42 | 0.034 | −0.033 | −0.0011 | 0.0012 |
| TaLK Convolution | 1.367 = k + b * 19.44 | 0.054 | 0.102 | 0.0055 | 0.0029 |
| ERNIE-Doc (247M) | 1.225 = k + b * 19.46 | 0.074 | −0.040 | −0.0030 | 0.0055 |
| Transformer+Recurrent Windows of Context | 1.427 = k + b * 20.07 | 0.684 | 0.161 | 0.1105 | 0.4682 |
| GPT3-6.7B (rerun of original) | 0.960 = k + b * 22.08 | 2.694 | −0.305 | −0.8220 | 7.2592 |
| Turing-NLG | 1.009 = k + b * 22.20 | 2.814 | −0.257 | −0.7220 | 7.9202 |
| added up over 14 rows | −1.7303 | 23.6029 | |||
b = −1.7303 ÷ 23.6029 = −0.0733 weight = 23.6029 ÷ 401.0633 = 0.0589 | |||||
WT2 · 20203 rows shared k = aWT2 + g2020 avg log C 17.79 avg log P 1.732 | b = 0.72350.0% | ||||
| CT-MoS + DynamicEval (WT2) | 1.612 = k + b * 17.75 | −0.040 | −0.120 | 0.0048 | 0.0016 |
| CT-MoS (WT2) | 1.794 = k + b * 17.75 | −0.040 | 0.062 | −0.0025 | 0.0016 |
| AWD-FWM (WT2) | 1.790 = k + b * 17.87 | 0.080 | 0.058 | 0.0046 | 0.0064 |
| added up over 3 rows | 0.0069 | 0.0096 | |||
b = 0.0069 ÷ 0.0096 = 0.7235 weight = 0.0096 ÷ 401.0633 = 0.0000 | |||||
PTB · 20214 rows shared k = aPTB + g2021 avg log C 17.77 avg log P 1.644 | b = −0.11845.5% | ||||
| Selfish-RNN (SNT-ASGD) Stacked LSTMs | 1.854 = k + b * 16.15 | −1.623 | 0.210 | −0.3411 | 2.6325 |
| Selfish-RNN (SNT-ASGD)RHNs | 1.806 = k + b * 16.33 | −1.443 | 0.163 | −0.2348 | 2.0808 |
| Selfish-RNN (ON-LSTM) | 1.747 = k + b * 16.80 | −0.973 | 0.103 | −0.1004 | 0.9458 |
| GPT-Neo-2.7B (finetuned on PTB) | 1.167 = k + b * 21.81 | 4.037 | −0.476 | −1.9229 | 16.3014 |
| added up over 4 rows | −2.5992 | 21.9605 | |||
b = −2.5992 ÷ 21.9605 = −0.1184 weight = 21.9605 ÷ 401.0633 = 0.0548 | |||||
WT103 · 202127 rows shared k = aWT103 + g2021 avg log C 19.57 avg log P 1.291 | b = −0.078917.2% | ||||
| GPT2+CoreLM+Fine-Tuning | 1.470 = k + b * 16.50 | −3.069 | 0.179 | −0.5484 | 9.4204 |
| Delta RNN (+ full context) | 1.516 = k + b * 18.04 | −1.529 | 0.225 | −0.3435 | 2.3386 |
| Transformer-C | 1.400 = k + b * 18.26 | −1.309 | 0.108 | −0.1419 | 1.7142 |
| Linear Transformer (small) | 1.550 = k + b * 18.47 | −1.099 | 0.259 | −0.2846 | 1.2084 |
| PermuteFormer | 1.512 = k + b * 18.49 | −1.079 | 0.220 | −0.2379 | 1.1648 |
| Subformer (83M) | 1.320 = k + b * 18.56 | −1.009 | 0.028 | −0.0287 | 1.0186 |
| Linear Transformer (large) | 1.498 = k + b * 18.59 | −0.979 | 0.207 | −0.2027 | 0.9589 |
| Subformer (96M) | 1.309 = k + b * 18.62 | −0.949 | 0.018 | −0.0172 | 0.9011 |
| GPT-2-Small+Pixelfly | 1.352 = k + b * 18.62 | −0.949 | 0.061 | −0.0578 | 0.9011 |
| Subformer (122M) | 1.299 = k + b * 18.72 | −0.849 | 0.008 | −0.0064 | 0.7212 |
| SRU++ Base | 1.262 = k + b * 18.76 | −0.809 | −0.029 | 0.0233 | 0.6549 |
| RFA-GATE-Gaussian-Stateful Big | 1.371 = k + b * 18.85 | −0.719 | 0.080 | −0.0574 | 0.5173 |
| base LM+GNN+kNN | 1.225 = k + b * 18.86 | −0.709 | −0.066 | 0.0468 | 0.5030 |
| SRU++ Large only 2 attention layers (k=5) | 1.238 = k + b * 18.90 | −0.669 | −0.053 | 0.0356 | 0.4479 |
| SRU++ Large | 1.233 = k + b * 19.04 | −0.529 | −0.058 | 0.0309 | 0.2801 |
| GPT-2-Medium+Pixelfly | 1.322 = k + b * 19.10 | −0.469 | 0.031 | −0.0145 | 0.2202 |
| DEQ-Transformer (Post-LN) + Jacobian Regularisation | 1.396 = k + b * 19.46 | −0.109 | 0.105 | −0.0115 | 0.0119 |
| S4 | 1.321 = k + b * 19.89 | 0.321 | 0.030 | 0.0096 | 0.1029 |
| Adaptive Input Transformer + RD | 1.257 = k + b * 19.91 | 0.341 | −0.034 | −0.0117 | 0.1161 |
| $\infty$-former (SM) | 1.220 = k + b * 20.08 | 0.511 | −0.071 | −0.0362 | 0.2609 |
| ALiBi (L=3072, Lvalid = 3072) | 1.262 = k + b * 20.26 | 0.691 | −0.029 | −0.0199 | 0.4771 |
| HSO | 1.307 = k + b * 20.54 | 0.971 | 0.016 | 0.0157 | 0.9423 |
| GPT-2 (1.5B, Curriculum Learning 45K) | 1.137 = k + b * 20.78 | 1.211 | −0.154 | −0.1864 | 1.4659 |
| Gopher (280B) | 0.910 = k + b * 22.11 | 2.541 | −0.382 | −0.9699 | 6.4554 |
| GLM-10B-bidirectional | 1.054 = k + b * 22.58 | 3.011 | −0.237 | −0.7137 | 9.0646 |
| GLM-10B-unidirectional | 1.087 = k + b * 22.58 | 3.011 | −0.204 | −0.6148 | 9.0646 |
| Gopher (7.1B) | 1.034 = k + b * 23.80 | 4.231 | −0.257 | −1.0893 | 17.8992 |
| added up over 27 rows | −5.4326 | 68.8316 | |||
b = −5.4326 ÷ 68.8316 = −0.0789 weight = 68.8316 ÷ 401.0633 = 0.1716 | |||||
WT2 · 20219 rows shared k = aWT2 + g2021 avg log C 19.85 avg log P 1.282 | b = −0.070812.0% | ||||
| GPT-2 (fine-tuned with HYDRA) | 1.181 = k + b * 16.28 | −3.567 | −0.101 | 0.3619 | 12.7211 |
| GPT2+CoreLM+Fine-Tuning | 1.502 = k + b * 16.50 | −3.347 | 0.220 | −0.7362 | 11.2002 |
| Selfish-RNN (AWD-LSTM-MoS) | 1.800 = k + b * 17.29 | −2.557 | 0.517 | −1.3224 | 6.5365 |
| GPT-Neo-125M(finetuned) | 1.342 = k + b * 20.48 | 0.633 | 0.059 | 0.0375 | 0.4011 |
| GPT-Neo-125M | 1.509 = k + b * 20.48 | 0.633 | 0.227 | 0.1435 | 0.4011 |
| GPT-Neo-2.7B (finetuned) | 1.033 = k + b * 21.81 | 1.963 | −0.250 | −0.4905 | 3.8547 |
| GPT-Neo-2.7B | 1.057 = k + b * 21.81 | 1.963 | −0.226 | −0.4436 | 3.8547 |
| GPT-Neo-1.3B (finetuned) | 1.082 = k + b * 21.81 | 1.963 | −0.200 | −0.3927 | 3.8547 |
| GPT-J-6B | 1.037 = k + b * 22.16 | 2.313 | −0.246 | −0.5687 | 5.3515 |
| added up over 9 rows | −3.4111 | 48.1756 | |||
b = −3.4111 ÷ 48.1756 = −0.0708 weight = 48.1756 ÷ 401.0633 = 0.1201 | |||||
PTB · 20224 rows shared k = aPTB + g2022 avg log C 20.04 avg log P 1.241 | b = −0.11963.9% | ||||
| Mogrifier RLSTM (PTB) | 1.632 = k + b * 16.73 | −3.305 | 0.392 | −1.2944 | 10.9230 |
| OPT-125M (finetuned on PTB) | 1.217 = k + b * 20.35 | 0.315 | −0.023 | −0.0074 | 0.0992 |
| OPT-1.3B (finetuned on PTB) | 1.080 = k + b * 21.37 | 1.335 | −0.161 | −0.2148 | 1.7822 |
| OPT-2.7B (finetuned on PTB) | 1.033 = k + b * 21.69 | 1.655 | −0.207 | −0.3432 | 2.7390 |
| added up over 4 rows | −1.8598 | 15.5435 | |||
b = −1.8598 ÷ 15.5435 = −0.1196 weight = 15.5435 ÷ 401.0633 = 0.0388 | |||||
WT103 · 202219 rows shared k = aWT103 + g2022 avg log C 19.94 avg log P 1.249 | b = −0.10457.8% | ||||
| Segatron -XL base, M=150 + HCP | 1.344 = k + b * 18.24 | −1.704 | 0.096 | −0.1628 | 2.9043 |
| Transformer Large + HCP | 1.403 = k + b * 18.78 | −1.164 | 0.154 | −0.1796 | 1.3554 |
| MemSizer | 1.318 = k + b * 18.86 | −1.084 | 0.069 | −0.0750 | 1.1755 |
| LaMemo | 1.376 = k + b * 18.87 | −1.074 | 0.127 | −0.1366 | 1.1539 |
| Transformer + GFM | 1.302 = k + b * 18.91 | −1.034 | 0.053 | −0.0551 | 1.0696 |
| DITTO | 1.386 = k + b * 19.04 | −0.904 | 0.137 | −0.1241 | 0.8176 |
| Decaying Fast Weights Transformer | 1.312 = k + b * 19.11 | −0.834 | 0.063 | −0.0525 | 0.6959 |
| Segatron-XL large, M=384 + HCP | 1.230 = k + b * 19.42 | −0.524 | −0.018 | 0.0097 | 0.2748 |
| B2T connection (16L) | 1.283 = k + b * 19.45 | −0.494 | 0.034 | −0.0170 | 0.2442 |
| Hybrid H3-125M | 1.375 = k + b * 19.59 | −0.354 | 0.126 | −0.0446 | 0.1255 |
| Hybrid H3-355M | 1.228 = k + b * 20.05 | 0.106 | −0.021 | −0.0022 | 0.0112 |
| NMST+GPT-2 | 1.316 = k + b * 20.08 | 0.136 | 0.067 | 0.0091 | 0.0184 |
| NoPos | 1.322 = k + b * 20.21 | 0.266 | 0.073 | 0.0193 | 0.0706 |
| Monarch-GPT-2-Small | 1.316 = k + b * 20.28 | 0.336 | 0.067 | 0.0225 | 0.1128 |
| Hybrid H3-1.3B | 1.097 = k + b * 20.61 | 0.666 | −0.152 | −0.1012 | 0.4433 |
| Monarch-GPT-2-Medium | 1.307 = k + b * 20.64 | 0.696 | 0.059 | 0.0408 | 0.4841 |
| Hybrid H3-2.7B | 1.025 = k + b * 20.93 | 0.986 | −0.224 | −0.2204 | 0.9718 |
| GPT3-6.7B + muP | 0.932 = k + b * 22.11 | 2.166 | −0.316 | −0.6852 | 4.6906 |
| Chinchilla | 0.855 = k + b * 23.76 | 3.816 | −0.394 | −1.5032 | 14.5602 |
| added up over 19 rows | −3.2581 | 31.1799 | |||
b = −3.2581 ÷ 31.1799 = −0.1045 weight = 31.1799 ÷ 401.0633 = 0.0777 | |||||
WT2 · 202218 rows shared k = aWT2 + g2022 avg log C 21.58 avg log P 1.177 | b = −0.11398.6% | ||||
| Mogrifier RLSTM (WT2) | 1.580 = k + b * 17.04 | −4.540 | 0.403 | −1.8282 | 20.6116 |
| OPT-125M (finetuned) | 1.298 = k + b * 20.35 | −1.230 | 0.121 | −0.1484 | 1.5129 |
| OPT-350M | 1.405 = k + b * 20.35 | −1.230 | 0.228 | −0.2805 | 1.5129 |
| BLOOM-560M | 1.478 = k + b * 21.07 | −0.510 | 0.301 | −0.1534 | 0.2601 |
| BLOOM-1B | 1.375 = k + b * 21.35 | −0.230 | 0.198 | −0.0455 | 0.0529 |
| OPT-1.3B (finetuned) | 1.087 = k + b * 21.37 | −0.210 | −0.090 | 0.0189 | 0.0441 |
| OPT-1.3B | 1.215 = k + b * 21.37 | −0.210 | 0.038 | −0.0080 | 0.0441 |
| BLOOM-1.7B | 1.305 = k + b * 21.56 | −0.020 | 0.128 | −0.0026 | 0.0004 |
| OPT-2.7B (finetuned on WT2) | 1.012 = k + b * 21.69 | 0.110 | −0.166 | −0.0182 | 0.0121 |
| OPT-2.7B | 1.096 = k + b * 21.69 | 0.110 | −0.081 | −0.0089 | 0.0121 |
| BLOOM-3B | 1.245 = k + b * 21.80 | 0.220 | 0.068 | 0.0149 | 0.0484 |
| OPT-6.7B | 1.036 = k + b * 22.08 | 0.500 | −0.141 | −0.0706 | 0.2500 |
| BLOOM-7.1B | 1.168 = k + b * 22.17 | 0.590 | −0.009 | −0.0054 | 0.3481 |
| OPT-13B | 1.006 = k + b * 22.37 | 0.790 | −0.171 | −0.1355 | 0.6241 |
| OPT-30B | 1.028 = k + b * 22.73 | 1.150 | −0.149 | −0.1713 | 1.3225 |
| GPT-NeoX-20B | 0.964 = k + b * 22.75 | 1.170 | −0.213 | −0.2496 | 1.3689 |
| OPT-66B | 0.970 = k + b * 23.08 | 1.500 | −0.207 | −0.3101 | 2.2500 |
| OPT-175B | 0.922 = k + b * 23.62 | 2.040 | −0.255 | −0.5210 | 4.1616 |
| added up over 18 rows | −3.9234 | 34.4368 | |||
b = −3.9234 ÷ 34.4368 = −0.1139 weight = 34.4368 ÷ 401.0633 = 0.0859 | |||||
PTB · 20233 rows shared k = aPTB + g2023 avg log C 23.02 avg log P 0.936 | b = −0.11870.1% | ||||
| LLaMA-7B (LoRA finetuned) | 0.986 = k + b * 22.66 | −0.363 | 0.050 | −0.0183 | 0.1320 |
| LLaMA-13B (LoRA finetuned) | 0.937 = k + b * 22.93 | −0.093 | 0.000 | −0.0000 | 0.0087 |
| LLaMA-33B (LoRA finetuned) | 0.885 = k + b * 23.48 | 0.457 | −0.051 | −0.0232 | 0.2085 |
| added up over 3 rows | −0.0415 | 0.3493 | |||
b = −0.0415 ÷ 0.3493 = −0.1187 weight = 0.3493 ÷ 401.0633 = 0.0009 | |||||
WT2 · 202316 rows shared k = aWT2 + g2023 avg log C 22.03 avg log P 1.033 | b = −0.12449.1% | ||||
| GPT-2+Active-SGD | 1.314 = k + b * 17.49 | −4.538 | 0.281 | −1.2729 | 20.5889 |
| Pythia-160m | 1.524 = k + b * 20.46 | −1.567 | 0.491 | −0.7696 | 2.4571 |
| Pythia-410m | 1.303 = k + b * 20.87 | −1.157 | 0.270 | −0.3128 | 1.3398 |
| Pythia-1b | 1.216 = k + b * 21.26 | −0.767 | 0.183 | −0.1405 | 0.5891 |
| Pythia-1.4b | 1.168 = k + b * 21.40 | −0.628 | 0.135 | −0.0846 | 0.3938 |
| Pythia-2.8b | 1.103 = k + b * 21.70 | −0.328 | 0.070 | −0.0230 | 0.1073 |
| Pythia-6.9b | 1.057 = k + b * 22.09 | 0.063 | 0.024 | 0.0015 | 0.0039 |
| Pythia-12b | 1.023 = k + b * 22.33 | 0.302 | −0.010 | −0.0031 | 0.0915 |
| MPT-7B | 0.998 = k + b * 22.62 | 0.593 | −0.035 | −0.0207 | 0.3511 |
| LLaMA-7B | 0.977 = k + b * 22.66 | 0.633 | −0.056 | −0.0353 | 0.4001 |
| LLaMA-7B (LoRA finetuned) | 0.792 = k + b * 22.66 | 0.633 | −0.241 | −0.1527 | 0.4001 |
| LLaMA-13B | 1.146 = k + b * 22.93 | 0.902 | 0.113 | 0.1017 | 0.8145 |
| LLaMA-13B (LoRA finetuned) | 0.744 = k + b * 22.93 | 0.902 | −0.290 | −0.2614 | 0.8145 |
| LLaMA-33B | 0.839 = k + b * 23.48 | 1.453 | −0.194 | −0.2822 | 2.1098 |
| LLaMA-65B | 0.695 = k + b * 23.78 | 1.753 | −0.338 | −0.5917 | 3.0713 |
| LLaMA-65B (LoRA finetuned) | 0.630 = k + b * 23.78 | 1.753 | −0.403 | −0.7057 | 3.0713 |
| added up over 16 rows | −4.5531 | 36.6037 | |||
b = −4.5531 ÷ 36.6037 = −0.1244 weight = 36.6037 ÷ 401.0633 = 0.0913 | |||||
That b column, drawn against the year:
One b out of those 22
Each group gets a say in proportion to its weight, and a weight is one division: that group's x² total, carried down from the table above, divided by every group's x² added together. That total is 401.0633, in the foot of this table. Then multiply every group's b by its weight and add that column up.
weight = group's x² ÷ 401.0633
x² is how far apart on compute a group's models are: square each row's x from the table above and add them up. A group whose models all trained on much the same compute has almost none of it, so almost none of the say.
| group | rows | its b | its x² | ÷ 401.0633 = weight | b * weight |
|---|---|---|---|---|---|
| PTB · 2015 | 4 | 0.0974 | 1.3183 | 0.0033 | 0.00032 |
| PTB · 2016 | 6 | −0.0439 | 0.5937 | 0.0015 | −0.00006 |
| WT2 · 2016 | 2 | 3.2160 | 0.0005 | 0.0000 | 0.00000 |
| PTB · 2017 | 9 | −0.0565 | 4.3386 | 0.0108 | −0.00061 |
| WT2 · 2017 | 5 | −0.0272 | 3.3283 | 0.0083 | −0.00023 |
| PTB · 2018 | 9 | −0.0419 | 8.4458 | 0.0211 | −0.00088 |
| WT103 · 2018 | 6 | −0.0664 | 2.9019 | 0.0072 | −0.00048 |
| WT2 · 2018 | 7 | 0.0038 | 3.4717 | 0.0087 | 0.00003 |
| PTB · 2019 | 10 | −0.0438 | 38.1894 | 0.0952 | −0.00417 |
| WT103 · 2019 | 19 | −0.0434 | 33.8123 | 0.0843 | −0.00366 |
| WT2 · 2019 | 4 | −0.1189 | 15.7695 | 0.0393 | −0.00468 |
| PTB · 2020 | 9 | −0.0677 | 8.2001 | 0.0204 | −0.00138 |
| WT103 · 2020 | 14 | −0.0733 | 23.6029 | 0.0589 | −0.00431 |
| WT2 · 2020 | 3 | 0.7235 | 0.0096 | 0.0000 | 0.00002 |
| PTB · 2021 | 4 | −0.1184 | 21.9605 | 0.0548 | −0.00648 |
| WT103 · 2021 | 27 | −0.0789 | 68.8316 | 0.1716 | −0.01355 |
| WT2 · 2021 | 9 | −0.0708 | 48.1756 | 0.1201 | −0.00851 |
| PTB · 2022 | 4 | −0.1196 | 15.5435 | 0.0388 | −0.00464 |
| WT103 · 2022 | 19 | −0.1045 | 31.1799 | 0.0777 | −0.00812 |
| WT2 · 2022 | 18 | −0.1139 | 34.4368 | 0.0859 | −0.00978 |
| PTB · 2023 | 3 | −0.1187 | 0.3493 | 0.0009 | −0.00010 |
| WT2 · 2023 | 16 | −0.1244 | 36.6037 | 0.0913 | −0.01135 |
| added up | 401.0633 | 1.0000 | −0.0826 | ||
−0.0826
what ten times the compute does to the score, from these 207 rows
There is a shorter road to the same number, and it is the one the pipeline drives. Add every group's x * y together, add every group's x² together, and divide once:
b = −33.1370 ÷ 401.0633 = −0.0826
Weighting each group's b by its x² is the same arithmetic as pooling those two totals, written out one group at a time. That is why the column above adds to exactly this figure rather than merely close to it, and it is the check that the long way and the short way are one calculation.
That is the b the rest of the page uses, and it is the one the model uses too. Every comparison behind it is between two rows on the same test in the same year, which is the only comparison that needs nothing assumed about how a test's baseline and a year's gain combine.
Step 2 · finding atest and gyear
What is left in a score once the compute is off it
The data
A leftover can only be two things, because they are the only two left in the equation: how hard that test is, and what that year had learned.
The equation
atest = that test's leftovers, each less its own year's gyear, averaged
gyear = that year's leftovers, each less its own test's atest, averaged
Two plain averages, and each one needs the other's answer first. So start with every gyear at zero and run them in turn until neither moves: 8 passes here.
The result
−0.3536
the 2023 gain: on the same test, for the same compute, a 2023 score sits that far below a 2015 one, in powers of ten.
9 of these, one a year, and one baseline for each of the 3 tests. 2015 reads zero because every other year is measured against it.
Take the compute term off a row and what is left can only be those two. So take that year's gain off as well, and a test's baseline is simply the average of its own rows.
Neither of those two is known yet, and each one needs the other: a test's baseline needs that year's gain taken off first, and a year's gain needs that test's baseline taken off first. Nothing has to go first, though. Start by assuming there was no progress at all, every gyear at zero, and run the two averages in turn.
Pass 1, first half
With every gain at zero there is nothing to take off, so a test's baseline is the plain average of log10 P − b * log10 C over its own rows.
PTB
3.2355
WT103
3.0314
WT2
3.1254
Pass 1, second half
Take those three off instead, and average within each year. That is a first set of gains, and they are all too small: 2023 comes out at −0.2811 because the baselines it used were built on the assumption that no year had gained anything. So go round again with these gains in hand.
Each pass moves the column less than the one before it, and by pass 6 none of these four decimals changes again. The loop runs 8 passes because it keeps going until a pass moves nothing anywhere by more than a hundred-thousandth, which is finer than what is printed here. The bottom row is what the model actually runs: one least-squares solve that jumps straight to the place this loop walks to. The two agree to within 0.000001, and the three baselines settle on the same values too, which is why the loop is worth showing and the solve is worth using.
year
2023, pass by pass
Every number in that table is one average over real rows, the solve row included. Pick any cell and the rows behind it open below.
Pass 1 · 2023
−0.2811
one average over the 20 results dated 2023, then shifted so 2015 reads zero.
A · the baselines this cell was measured against
This pass started with every gain at zero, so there was nothing to take off: each baseline is the plain average of that test's leftovers.
the gains it started from
| test | rows | added up, each row's leftover , nothing taken off | ÷ rows | = baseline |
|---|---|---|---|---|
| 58 | 180.361 | 58 | 3.1097 | |
| 86 | 249.875 | 86 | 2.9055 | |
| 64 | 191.970 | 64 | 2.9995 |
These are the baselines step B measures against. The three cards higher up print the same ones with this pass's shift already added, so PTB reads 3.2355 there and 3.1097 here. The gap between the two is step C, and taking it off here or off the year below gives the same cell.
Each baseline is a sum over that test's own rows. Open one and they are listed below.
B · every 2023 row, against its own test's baseline
These are the 20 results the cell is an average of. Each one's leftover is its score as a power of ten with the compute term taken off, and the last column takes off the baseline of the test it was run on. Open a row and it shows its own six pieces of arithmetic, starting from the two numbers the file holds for it.
| model | test | log P | b * log C | leftover | its baseline | leftover − baseline |
|---|---|---|---|---|---|---|
| WT2 | 1.3137 | −1.4451 | 2.7587 | 2.9995 | −0.2408 | |
| WT2 | 0.8388 | −1.9400 | 2.7788 | 2.9995 | −0.2207 | |
| WT2 | 1.1458 | −1.8945 | 3.0404 | 2.9995 | 0.0408 | |
| WT2 | 0.9773 | −1.8722 | 2.8495 | 2.9995 | −0.1500 | |
| WT2 | 0.6955 | −1.9648 | 2.6603 | 2.9995 | −0.3393 | |
| WT103 | 1.3096 | −1.6483 | 2.9580 | 2.9055 | 0.0524 | |
| WT2 | 1.0228 | −1.8450 | 2.8678 | 2.9995 | −0.1317 | |
| WT2 | 1.0573 | −1.8251 | 2.8824 | 2.9995 | −0.1171 | |
| WT2 | 1.5241 | −1.6905 | 3.2146 | 2.9995 | 0.2151 | |
| WT2 | 1.2162 | −1.7566 | 2.9727 | 2.9995 | −0.0268 | |
| WT2 | 1.1679 | −1.7681 | 2.9360 | 2.9995 | −0.0635 | |
| WT2 | 1.3034 | −1.7243 | 3.0278 | 2.9995 | 0.0282 | |
| WT2 | 1.1035 | −1.7929 | 2.8964 | 2.9995 | −0.1031 | |
| WT2 | 0.9983 | −1.8689 | 2.8672 | 2.9995 | −0.1323 | |
| PTB | 0.8854 | −1.9400 | 2.8253 | 3.1097 | −0.2843 | |
| PTB | 0.9365 | −1.8945 | 2.8311 | 3.1097 | −0.2786 | |
| PTB | 0.9863 | −1.8722 | 2.8586 | 3.1097 | −0.2511 | |
| WT2 | 0.7435 | −1.8945 | 2.6381 | 2.9995 | −0.3615 | |
| WT2 | 0.6304 | −1.9648 | 2.5952 | 2.9995 | −0.4043 | |
| WT2 | 0.7917 | −1.8722 | 2.6639 | 2.9995 | −0.3356 |
- added up over 20 rows
- −3.104
- ÷ 20 rows
- −0.1552
The total is the pipeline's own, added before anything was rounded, so adding the printed column by hand lands within a thousandth of it.
C · pin 2015 at zero
−0.1552 − 0.1259 = −0.2811
The number taken off is what step B gives for 2015 on this same pass. Taking it off every year is what makes the 2015 column read zero and every other cell a reading against it. It moves the whole row together, so no gap between two years changes.
Every row of a given year carries that year's settled number, which is why the gyear column below repeats down each year while the other two columns change row by row. g2015 reads zero in every pass because each pass is shifted to put it there. The whole column can slide up or down together without changing a single gap between two years, so one year has to be pinned before the numbers mean anything, and pinning 2015 is what makes every gain a reading against it. The division below then arrives at that zero on its own, which is the subject of the last paragraph in this step.
| model | year | log P | b * log C | gyear | leftover |
|---|---|---|---|---|---|
| PTB · 58 rows | |||||
| genCNN + dyn eval | 2015 | 2.0265 | −1.3930 | 0.0000 | 3.4196 |
| Search-Proven Best LSTM | 2015 | 1.9022 | −1.2823 | 0.0000 | 3.1845 |
| LSTM-Char-Large | 2015 | 1.8971 | −1.2740 | 0.0000 | 3.1711 |
| Variational (untied weights, MC) LSTM (Large) | 2015 | 1.8657 | −1.3013 | 0.0000 | 3.1670 |
| Variational RHN + WT | 2016 | 1.8156 | −1.2732 | −0.0252 | 3.1140 |
| VD-RHN | 2016 | 1.8357 | −1.2848 | −0.0252 | 3.1457 |
| Pointer Sentinel-LSTM (medium) | 2016 | 1.8506 | −1.3112 | −0.0252 | 3.1871 |
| Zoneout + Variational LSTM (PTB) | 2016 | 1.9063 | −1.3112 | −0.0252 | 3.2428 |
| VD-LSTM+REAL Large | 2016 | 1.8357 | −1.3492 | −0.0252 | 3.2101 |
| GCRN-M1, dropout | 2016 | 1.9942 | −1.2790 | −0.0252 | 3.2984 |
| 4 layer Densely Connected LSTM | 2017 | 1.8854 | −1.3129 | −0.1107 | 3.3090 |
| Densely Connected LSTM + Var. Dropout | 2017 | 1.8938 | −1.3311 | −0.1107 | 3.3355 |
| AWD-LSTM - 3-layer LSTM (tied) + continuous cache pointer (PTB) | 2017 | 1.7226 | −1.3905 | −0.1107 | 3.2239 |
| EI-REHN-1200D | 2017 | 1.8209 | −1.3079 | −0.1107 | 3.2395 |
| EI-REHN-1000D | 2017 | 1.8370 | −1.3244 | −0.1107 | 3.2721 |
| GL-LWGC-AWD-MoS-LSTM + dynamic evaluation (PTB) | 2017 | 1.6660 | −1.4178 | −0.1107 | 3.1945 |
| AWD-LSTM+WT+Cache+IOG (PTB) | 2017 | 1.7243 | −1.2327 | −0.1107 | 3.0677 |
| Fraternal dropout + AWD-LSTM 3-layer (PTB) | 2017 | 1.7543 | −1.3914 | −0.1107 | 3.2564 |
| AWD-LSTM-MoS + dynamic evaluation (PTB, 2017) | 2017 | 1.6784 | −1.4120 | −0.1107 | 3.2012 |
| RNNLM + Dynamic KL Regularization | 2018 | 1.8910 | −1.2798 | −0.0698 | 3.2406 |
| AWD-LSTM-MoS+Noisin+dynamic evaluation | 2018 | 1.6776 | −1.3790 | −0.0698 | 3.1264 |
| Dropout-LSTM+Noise(Bernoulli) (PTB) | 2018 | 1.8202 | −1.3848 | −0.0698 | 3.2747 |
| aLSTM(depth-2)+RecurrentPolicy (PTB) | 2018 | 1.7427 | −1.3534 | −0.0698 | 3.1659 |
| AWD-LSTM-MoS+PDR + dynamic evaluation (PTB) | 2018 | 1.6749 | −1.4170 | −0.0698 | 3.1616 |
| AWD-LSTM-DOC (fin) (23M) | 2018 | 1.7192 | −1.3996 | −0.0698 | 3.1886 |
| TrellisNet-MoS (1.4x larger) | 2018 | 1.7339 | −1.5236 | −0.0698 | 3.3273 |
| Fine-tuned-AWD-LSTM-DOC(fin) | 2018 | 1.7170 | −1.2625 | −0.0698 | 3.0493 |
| Multi-cell LSTM | 2018 | 1.8872 | −1.2641 | −0.0698 | 3.2211 |
| Transformer-XL-ptb | 2019 | 1.7366 | −1.5731 | −0.1361 | 3.4458 |
| GPT-2 (1542M) | 2019 | 1.5534 | −1.7500 | −0.1361 | 3.4395 |
| AWD-LSTM-DRILL + dynamic evaluation† (PTB) | 2019 | 1.6937 | −1.4153 | −0.1361 | 3.2452 |
| Adversarial + AWD-LSTM-MoS + partial shuffled | 2019 | 1.6629 | −1.3831 | −0.1361 | 3.1821 |
| Tensorized Transformer (small) | 2019 | 1.7627 | −1.2641 | −0.1361 | 3.1629 |
| Tensorized Transformer (large PTB) | 2019 | 1.7218 | −1.2889 | −0.1361 | 3.1469 |
| R-Transformer | 2019 | 1.9262 | −1.3154 | −0.1361 | 3.3777 |
| DEQ-TrellisNet | 2019 | 1.7566 | −1.4798 | −0.1361 | 3.3725 |
| AWD-LSTM + DeFINE | 2019 | 1.7340 | −1.2683 | −0.1361 | 3.1384 |
| bRSM + cache | 2019 | 2.0149 | −1.1840 | −0.1361 | 3.3351 |
| LSTM-3-layer+Gadam | 2020 | 1.7692 | −1.3575 | −0.1558 | 3.2824 |
| Tensor-Transformer(1core)+PN (PTB) | 2020 | 1.6776 | −1.2641 | −0.1558 | 3.0975 |
| DiffStk-MRNN | 2020 | 2.0607 | −1.1939 | −0.1558 | 3.4104 |
| ONLSTM-SYD | 2020 | 1.7459 | −1.4162 | −0.1558 | 3.3178 |
| rTop-k(distributed setting) | 2020 | 1.9164 | −1.3352 | −0.1558 | 3.4074 |
| 3-Layer-Tensor-Transformer+AdaHessian | 2020 | 1.7118 | −1.2641 | −0.1558 | 3.1317 |
| AWD-FWM (PTB) | 2020 | 1.7362 | −1.4153 | −0.1558 | 3.3074 |
| CT-MoS (PTB) | 2020 | 1.7379 | −1.4153 | −0.1558 | 3.3090 |
| CT-MoS + DynamicEval (PTB) | 2020 | 1.6760 | −1.4153 | −0.1558 | 3.2471 |
| Selfish-RNN (ON-LSTM) | 2021 | 1.7468 | −1.3881 | −0.1951 | 3.3300 |
| Selfish-RNN (SNT-ASGD) Stacked LSTMs | 2021 | 1.8538 | −1.3344 | −0.1951 | 3.3833 |
| Selfish-RNN (SNT-ASGD)RHNs | 2021 | 1.8064 | −1.3492 | −0.1951 | 3.3507 |
| GPT-Neo-2.7B (finetuned on PTB) | 2021 | 1.1673 | −1.8020 | −0.1951 | 3.1644 |
| OPT-125M (finetuned on PTB) | 2022 | 1.2175 | −1.6814 | −0.2300 | 3.1288 |
| OPT-2.7B (finetuned on PTB) | 2022 | 1.0334 | −1.7921 | −0.2300 | 3.0555 |
| OPT-1.3B (finetuned on PTB) | 2022 | 1.0799 | −1.7657 | −0.2300 | 3.0755 |
| Mogrifier RLSTM (PTB) | 2022 | 1.6325 | −1.3823 | −0.2300 | 3.2447 |
| LLaMA-33B (LoRA finetuned) | 2023 | 0.8854 | −1.9400 | −0.3536 | 3.1790 |
| LLaMA-13B (LoRA finetuned) | 2023 | 0.9365 | −1.8945 | −0.3536 | 3.1847 |
| LLaMA-7B (LoRA finetuned) | 2023 | 0.9863 | −1.8722 | −0.3536 | 3.2122 |
| added up over 58 rows | 187.661 | ||||
| ÷ 58 rows | aPTB = 3.2355 | ||||
| WT103 · 86 rows | |||||
| QRNN | 2018 | 1.5185 | −1.4509 | −0.0698 | 3.0392 |
| 4 layer QRNN (h=2500) | 2018 | 1.5185 | −1.4360 | −0.0698 | 3.0243 |
| LSTM (Hebbian, Cache, MbPA) | 2018 | 1.4654 | −1.6012 | −0.0698 | 3.1364 |
| Transformer (Adaptive Input Embeddings) | 2018 | 1.2718 | −1.5583 | −0.0698 | 2.8999 |
| TrellisNet | 2018 | 1.4652 | −1.5236 | −0.0698 | 3.0586 |
| TrellisNet-MoS (1.4x larger) | 2018 | 1.4652 | −1.5236 | −0.0698 | 3.0586 |
| Transformer-XL Large | 2019 | 1.2625 | −1.5731 | −0.1361 | 2.9717 |
| GPT-2 (1542M) | 2019 | 1.2425 | −1.7500 | −0.1361 | 3.1286 |
| GPT-2 (762M) | 2019 | 1.3434 | −1.7252 | −0.1361 | 3.2047 |
| GPT-2 (345M) | 2019 | 1.4211 | −1.6971 | −0.1361 | 3.2543 |
| Transformer-XL Large + Phrase Induction | 2019 | 1.2405 | −1.4211 | −0.1361 | 2.7978 |
| AdvSoft + 4 layer QRNN + dynamic evaluation | 2019 | 1.4472 | −1.4509 | −0.1361 | 3.0341 |
| 4 layer QRNN + dynamic evaluation | 2019 | 1.4997 | −1.4509 | −0.1361 | 3.0867 |
| Tensorized Transformer (257M) | 2019 | 1.3263 | −1.5434 | −0.1361 | 3.0059 |
| Tensorized Transformer (core-2) | 2019 | 1.2765 | −1.5037 | −0.1361 | 2.9163 |
| Tensorized Transformer (151M) | 2019 | 1.2742 | −1.5244 | −0.1361 | 2.9347 |
| All-attention network + adaptive span | 2019 | 1.3139 | −1.6244 | −0.1361 | 3.0744 |
| DEQ-Transformer (Medium, Adaptive Embedding) | 2019 | 1.3655 | −1.4798 | −0.1361 | 2.9814 |
| Megatron-LM (8.3B) | 2019 | 1.0338 | −1.8144 | −0.1361 | 2.9844 |
| Megatron-LM (355M) | 2019 | 1.2858 | −1.7053 | −0.1361 | 3.1272 |
| Sandwich Transformer | 2019 | 1.2514 | −1.6690 | −0.1361 | 3.0565 |
| Compressive Transformers for Long-Range Sequence Modelling | 2019 | 1.2330 | −1.6690 | −0.1361 | 3.0381 |
| Transformer-XL DeFINE (107M) | 2019 | 1.4103 | −1.5467 | −0.1361 | 3.0931 |
| Adaptive LSTM + DeFINE | 2019 | 1.5556 | −1.5525 | −0.1361 | 3.2442 |
| Transformer-XL DeFINE (141M) | 2019 | 1.3833 | −1.5525 | −0.1361 | 3.0719 |
| TaLK Convolution | 2020 | 1.3674 | −1.6062 | −0.1558 | 3.1293 |
| Turing-NLG | 2020 | 1.0090 | −1.8342 | −0.1558 | 2.9991 |
| Feedback Transformer | 2020 | 1.2625 | −1.5963 | −0.1558 | 3.0145 |
| TransformerXL + spectrum control | 2020 | 1.3655 | −1.4591 | −0.1558 | 2.9804 |
| Tensor-Transformer(1core)+PN (WT103) | 2020 | 1.2529 | −1.5037 | −0.1558 | 2.9124 |
| Segatron XL base, M=384 | 2020 | 1.3522 | −1.5070 | −0.1558 | 3.0150 |
| Segatron XL large, M=384 | 2020 | 1.2330 | −1.6045 | −0.1558 | 2.9933 |
| GPT3-6.7B (rerun of original) | 2020 | 0.9605 | −1.8243 | −0.1558 | 2.9406 |
| 6-Layer-Tensor-Transformer+AdaHessian | 2020 | 1.2989 | −1.5037 | −0.1558 | 2.9584 |
| DeLight | 2020 | 1.3827 | −1.6012 | −0.1558 | 3.1398 |
| Transformer+Recurrent Windows of Context | 2020 | 1.4270 | −1.6582 | −0.1558 | 3.2410 |
| Shortformer | 2020 | 1.2589 | −1.5269 | −0.1558 | 2.9415 |
| ERNIE-Doc (151M) | 2020 | 1.3222 | −1.5905 | −0.1558 | 3.0685 |
| ERNIE-Doc (247M) | 2020 | 1.2253 | −1.6078 | −0.1558 | 2.9890 |
| Subformer (83M) | 2021 | 1.3197 | −1.5335 | −0.1951 | 3.0483 |
| Subformer (122M) | 2021 | 1.2989 | −1.5467 | −0.1951 | 3.0407 |
| Subformer (96M) | 2021 | 1.3094 | −1.5384 | −0.1951 | 3.0430 |
| Linear Transformer (large) | 2021 | 1.4983 | −1.5360 | −0.1951 | 3.2294 |
| Linear Transformer (small) | 2021 | 1.5502 | −1.5260 | −0.1951 | 3.2714 |
| SRU++ Large | 2021 | 1.2330 | −1.5731 | −0.1951 | 3.0013 |
| SRU++ Large only 2 attention layers (k=5) | 2021 | 1.2380 | −1.5616 | −0.1951 | 2.9947 |
| SRU++ Base | 2021 | 1.2625 | −1.5500 | −0.1951 | 3.0076 |
| RFA-GATE-Gaussian-Stateful Big | 2021 | 1.3711 | −1.5574 | −0.1951 | 3.1236 |
| GLM-10B-bidirectional | 2021 | 1.0542 | −1.8656 | −0.1951 | 3.1150 |
| GLM-10B-unidirectional | 2021 | 1.0871 | −1.8656 | −0.1951 | 3.1478 |
| Transformer-C | 2021 | 1.3997 | −1.5087 | −0.1951 | 3.1035 |
| Delta RNN (+ full context) | 2021 | 1.5159 | −1.4905 | −0.1951 | 3.2015 |
| DEQ-Transformer (Post-LN) + Jacobian Regularisation | 2021 | 1.3962 | −1.6078 | −0.1951 | 3.1992 |
| Adaptive Input Transformer + RD | 2021 | 1.2570 | −1.6450 | −0.1951 | 3.0971 |
| GPT-2 (1.5B, Curriculum Learning 45K) | 2021 | 1.1374 | −1.7169 | −0.1951 | 3.0494 |
| ALiBi (L=3072, Lvalid = 3072) | 2021 | 1.2625 | −1.6739 | −0.1951 | 3.1315 |
| $\infty$-former (SM) | 2021 | 1.2204 | −1.6591 | −0.1951 | 3.0746 |
| PermuteFormer | 2021 | 1.5117 | −1.5277 | −0.1951 | 3.2346 |
| base LM+GNN+kNN | 2021 | 1.2253 | −1.5583 | −0.1951 | 2.9787 |
| S4 | 2021 | 1.3212 | −1.6434 | −0.1951 | 3.1597 |
| GPT2+CoreLM+Fine-Tuning | 2021 | 1.4700 | −1.3633 | −0.1951 | 3.0284 |
| GPT-2-Medium+Pixelfly | 2021 | 1.3222 | −1.5781 | −0.1951 | 3.0954 |
| GPT-2-Small+Pixelfly | 2021 | 1.3522 | −1.5384 | −0.1951 | 3.0857 |
| Gopher (7.1B) | 2021 | 1.0338 | −1.9664 | −0.1951 | 3.1954 |
| Gopher (280B) | 2021 | 0.9096 | −1.8268 | −0.1951 | 2.9315 |
| HSO | 2021 | 1.3075 | −1.6971 | −0.1951 | 3.1997 |
| GPT3-6.7B + muP | 2022 | 0.9325 | −1.8268 | −0.2300 | 2.9892 |
| Segatron-XL large, M=384 + HCP | 2022 | 1.2304 | −1.6045 | −0.2300 | 3.0650 |
| Transformer Large + HCP | 2022 | 1.4031 | −1.5517 | −0.2300 | 3.1848 |
| Segatron -XL base, M=150 + HCP | 2022 | 1.3444 | −1.5070 | −0.2300 | 3.0814 |
| MemSizer | 2022 | 1.3181 | −1.5583 | −0.2300 | 3.1063 |
| Chinchilla | 2022 | 0.8549 | −1.9631 | −0.2300 | 3.0480 |
| NoPos | 2022 | 1.3216 | −1.6698 | −0.2300 | 3.2214 |
| Monarch-GPT-2-Medium | 2022 | 1.3075 | −1.7053 | −0.2300 | 3.2428 |
| Monarch-GPT-2-Small | 2022 | 1.3160 | −1.6756 | −0.2300 | 3.2215 |
| LaMemo | 2022 | 1.3760 | −1.5591 | −0.2300 | 3.1651 |
| B2T connection (16L) | 2022 | 1.2833 | −1.6070 | −0.2300 | 3.1203 |
| DITTO | 2022 | 1.3861 | −1.5731 | −0.2300 | 3.1893 |
| NMST+GPT-2 | 2022 | 1.3158 | −1.6591 | −0.2300 | 3.2048 |
| Decaying Fast Weights Transformer | 2022 | 1.3118 | −1.5789 | −0.2300 | 3.1207 |
| Transformer + GFM | 2022 | 1.3021 | −1.5624 | −0.2300 | 3.0945 |
| Hybrid H3-355M | 2022 | 1.2279 | −1.6566 | −0.2300 | 3.1145 |
| Hybrid H3-125M | 2022 | 1.3747 | −1.6186 | −0.2300 | 3.2233 |
| Hybrid H3-2.7B | 2022 | 1.0253 | −1.7293 | −0.2300 | 2.9846 |
| Hybrid H3-1.3B | 2022 | 1.0969 | −1.7029 | −0.2300 | 3.0298 |
| Sparse Wide GPT-3 Small | 2023 | 1.3096 | −1.6483 | −0.3536 | 3.3116 |
| added up over 86 rows | 265.053 | ||||
| ÷ 86 rows | aWT103 = 3.0820 | ||||
| WT2 · 64 rows | |||||
| Zoneout + Variational LSTM (WT2) | 2016 | 2.0039 | −1.3410 | −0.0252 | 3.3701 |
| Pointer Sentinel-LSTM | 2016 | 1.9074 | −1.3385 | −0.0252 | 3.2711 |
| AWD-LSTM - 3-layer LSTM (tied) + continuous cache pointer (WT2) | 2017 | 1.7160 | −1.4451 | −0.1107 | 3.2718 |
| GL-LWGC-AWD-MoS-LSTM + dynamic evaluation (WT2) | 2017 | 1.6070 | −1.4608 | −0.1107 | 3.1785 |
| AWD-LSTM+WT+Cache+IOG (WT2) | 2017 | 1.7135 | −1.2823 | −0.1107 | 3.1065 |
| Fraternal dropout + AWD-LSTM 3-layer (WT2) | 2017 | 1.8069 | −1.4327 | −0.1107 | 3.3503 |
| AWD-LSTM-MoS + dynamic evaluation (WT2, 2017) | 2017 | 1.6094 | −1.4575 | −0.1107 | 3.1776 |
| RNNLM + Dynamic KL Regularization (WT2) | 2018 | 1.9385 | −1.3501 | −0.0698 | 3.3584 |
| LSTM+Noise(Beta) | 2018 | 1.9186 | −1.4129 | −0.0698 | 3.4012 |
| Dropout-LSTM+Noise(Laplace) | 2018 | 1.9143 | −1.3641 | −0.0698 | 3.3482 |
| Dropout-LSTM+Noise(Bernoulli) (WT2) | 2018 | 1.8854 | −1.4129 | −0.0698 | 3.3680 |
| aLSTM(depth-2)+RecurrentPolicy (WT2) | 2018 | 1.8096 | −1.3947 | −0.0698 | 3.2740 |
| AWD-LSTM-DOC (fin) (37M) | 2018 | 1.7637 | −1.4162 | −0.0698 | 3.2496 |
| LSTM+NeuralCache | 2018 | 1.8209 | −1.2402 | −0.0698 | 3.1308 |
| GPT-2 (1542M) | 2019 | 1.2634 | −1.7500 | −0.1361 | 3.1495 |
| AWD-LSTM-DRILL + dynamic evaluation† (WT2) | 2019 | 1.6232 | −1.4566 | −0.1361 | 3.2160 |
| AWD-LSTM + MoS + Partial Shuffled | 2019 | 1.5806 | −1.4476 | −0.1361 | 3.1643 |
| LSTM(medium)+Sememe+cell | 2019 | 1.9502 | −1.2972 | −0.1361 | 3.3835 |
| AWD-FWM (WT2) | 2020 | 1.7899 | −1.4765 | −0.1558 | 3.4222 |
| CT-MoS + DynamicEval (WT2) | 2020 | 1.6124 | −1.4666 | −0.1558 | 3.2347 |
| CT-MoS (WT2) | 2020 | 1.7939 | −1.4666 | −0.1558 | 3.4162 |
| Selfish-RNN (AWD-LSTM-MoS) | 2021 | 1.7997 | −1.4285 | −0.1951 | 3.4234 |
| GPT-Neo-2.7B (finetuned) | 2021 | 1.0326 | −1.8020 | −0.1951 | 3.0297 |
| GPT-Neo-125M(finetuned) | 2021 | 1.3416 | −1.6921 | −0.1951 | 3.2289 |
| GPT-Neo-125M | 2021 | 1.5091 | −1.6921 | −0.1951 | 3.3963 |
| GPT-Neo-2.7B | 2021 | 1.0565 | −1.8020 | −0.1951 | 3.0536 |
| GPT-Neo-1.3B (finetuned) | 2021 | 1.0824 | −1.8020 | −0.1951 | 3.0795 |
| GPT-J-6B | 2021 | 1.0366 | −1.8309 | −0.1951 | 3.0627 |
| GPT-2 (fine-tuned with HYDRA) | 2021 | 1.1810 | −1.3451 | −0.1951 | 2.7212 |
| GPT2+CoreLM+Fine-Tuning | 2021 | 1.5024 | −1.3633 | −0.1951 | 3.0608 |
| GPT-NeoX-20B | 2022 | 0.9638 | −1.8797 | −0.2300 | 3.0734 |
| OPT-66B | 2022 | 0.9703 | −1.9069 | −0.2300 | 3.1073 |
| OPT-2.7B (finetuned on WT2) | 2022 | 1.0116 | −1.7921 | −0.2300 | 3.0336 |
| OPT-6.7B | 2022 | 1.0358 | −1.8243 | −0.2300 | 3.0901 |
| OPT-1.3B (finetuned) | 2022 | 1.0871 | −1.7657 | −0.2300 | 3.0827 |
| OPT-2.7B | 2022 | 1.0959 | −1.7921 | −0.2300 | 3.1179 |
| OPT-175B | 2022 | 0.9217 | −1.9516 | −0.2300 | 3.1032 |
| OPT-13B | 2022 | 1.0056 | −1.8483 | −0.2300 | 3.0839 |
| OPT-125M (finetuned) | 2022 | 1.2978 | −1.6814 | −0.2300 | 3.2091 |
| OPT-1.3B | 2022 | 1.2151 | −1.7657 | −0.2300 | 3.2107 |
| OPT-30B | 2022 | 1.0282 | −1.8780 | −0.2300 | 3.1362 |
| OPT-350M | 2022 | 1.4052 | −1.6814 | −0.2300 | 3.3165 |
| BLOOM-1.7B | 2022 | 1.3047 | −1.7813 | −0.2300 | 3.3160 |
| BLOOM-1B | 2022 | 1.3747 | −1.7640 | −0.2300 | 3.3687 |
| BLOOM-560M | 2022 | 1.4778 | −1.7409 | −0.2300 | 3.4487 |
| BLOOM-3B | 2022 | 1.2448 | −1.8012 | −0.2300 | 3.2759 |
| BLOOM-7.1B | 2022 | 1.1679 | −1.8317 | −0.2300 | 3.2296 |
| Mogrifier RLSTM (WT2) | 2022 | 1.5798 | −1.4079 | −0.2300 | 3.2177 |
| GPT-2+Active-SGD | 2023 | 1.3137 | −1.4451 | −0.3536 | 3.1124 |
| LLaMA-33B | 2023 | 0.8388 | −1.9400 | −0.3536 | 3.1325 |
| LLaMA-13B | 2023 | 1.1458 | −1.8945 | −0.3536 | 3.3940 |
| LLaMA-7B | 2023 | 0.9773 | −1.8722 | −0.3536 | 3.2031 |
| LLaMA-65B | 2023 | 0.6955 | −1.9648 | −0.3536 | 3.0139 |
| Pythia-12b | 2023 | 1.0228 | −1.8450 | −0.3536 | 3.2215 |
| Pythia-6.9b | 2023 | 1.0573 | −1.8251 | −0.3536 | 3.2361 |
| Pythia-160m | 2023 | 1.5241 | −1.6905 | −0.3536 | 3.5682 |
| Pythia-1b | 2023 | 1.2162 | −1.7566 | −0.3536 | 3.3264 |
| Pythia-1.4b | 2023 | 1.1679 | −1.7681 | −0.3536 | 3.2897 |
| Pythia-410m | 2023 | 1.3034 | −1.7243 | −0.3536 | 3.3814 |
| Pythia-2.8b | 2023 | 1.1035 | −1.7929 | −0.3536 | 3.2500 |
| MPT-7B | 2023 | 0.9983 | −1.8689 | −0.3536 | 3.2208 |
| LLaMA-13B (LoRA finetuned) | 2023 | 0.7435 | −1.8945 | −0.3536 | 2.9917 |
| LLaMA-65B (LoRA finetuned) | 2023 | 0.6304 | −1.9648 | −0.3536 | 2.9488 |
| LLaMA-7B (LoRA finetuned) | 2023 | 0.7917 | −1.8722 | −0.3536 | 3.0176 |
| added up over 64 rows | 205.628 | ||||
| ÷ 64 rows | aWT2 = 3.2129 | ||||
The year gains are the same division the other way round. Take the compute term off a row and that test's baseline off too, and what is left is that year's gain, so average one year's rows instead of one test's:
| year | rows | added up, log10 P − b * log10 C − atest | ÷ rows | = gain |
|---|---|---|---|---|
| 2015 | 4 | 0.0000 | 4 | g2015 = 0.0000 |
| 2016 | 8 | −0.2015 | 8 | g2016 = −0.0252 |
| 2017 | 14 | −1.5501 | 14 | g2017 = −0.1107 |
| 2018 | 22 | −1.5353 | 22 | g2018 = −0.0698 |
| 2019 | 33 | −4.4923 | 33 | g2019 = −0.1361 |
| 2020 | 26 | −4.0508 | 26 | g2020 = −0.1558 |
| 2021 | 40 | −7.8046 | 40 | g2021 = −0.1951 |
| 2022 | 41 | −9.4293 | 41 | g2022 = −0.2300 |
| 2023 | 20 | −7.0730 | 20 | g2023 = −0.3536 |
Both divisions land back on the numbers the fit started from, which is what the two tables were for. g2015 comes out at exactly zero rather than being set there. Every 2015 row is on the same test, so that test's baseline absorbs their level and there is nothing left for the year to carry, which is what makes every later gain a reading against 2015. Those nine are the next step.
Step 3 · each year's gain, in compute
Divide every gyear by b
The data
Both are in score. The answer wants compute.
The equation
A minus over a minus gives a plus, which is right: the score fell, so the compute needed to reach a fixed score fell with it.
The result
607×
less compute in 2022 than in 2015 for the same score: 102.784.
Every year below, same division.
| year | rows behind it | gyear | ÷ b | = compute saved |
|---|---|---|---|---|
| 2015 | 4 | 0.0000 | −0.0826 | 0.000 |
| 2016 | 8 | −0.0252 | −0.0826 | 0.305 |
| 2017 | 14 | −0.1107 | −0.0826 | 1.340 |
| 2018 | 22 | −0.0698 | −0.0826 | 0.845 |
| 2019 | 33 | −0.1361 | −0.0826 | 1.648 |
| 2020 | 26 | −0.1558 | −0.0826 | 1.886 |
| 2021 | 40 | −0.1951 | −0.0826 | 2.362 |
| 2022 | 41 | −0.2300 | −0.0826 | 2.784 |
| 2023 | 20 | −0.3536 | −0.0826 | 4.280 |
Step 4 · the line through those savings
How steeply the saving rises is the answer
The data
The line starts at zero in 2015, because nothing had been saved yet. Only the steepness is unknown.
The equation
Σ means "add up the column". Both columns are worked out below.
The result
0.438
powers of ten of compute saved, every year
| year | x = years since 2015 | y = compute saved | x * y | x² |
|---|---|---|---|---|
| 2015 | 0 | 0.000 | 0.000 | 0 |
| 2016 | 1 | 0.305 | 0.305 | 1 |
| 2017 | 2 | 1.340 | 2.680 | 4 |
| 2018 | 3 | 0.845 | 2.534 | 9 |
| 2019 | 4 | 1.648 | 6.590 | 16 |
| 2020 | 5 | 1.886 | 9.428 | 25 |
| 2021 | 6 | 2.362 | 14.169 | 36 |
| 2022 | 7 | 2.784 | 19.485 | 49 |
| 2023 | 8 | 4.280 | 34.242 | 64 |
| added up | 89.43 | 204 | ||
Not very, and the range is worth knowing
What the answer would be if slightly different papers had been written
0.438 is one number worked out from one table, and nobody designed that table. It is the results that happened to get published and collected. A handful more in one year, or a few fewer in another, and the same arithmetic comes back with a different number. How different is worth knowing before anyone quotes this one.
There is no second table to check against and no way to get one, so the 208 rows have to stand in for the wider pool they came from. Draw 208 of them at random, putting each one back before drawing the next, and what comes out is a table the same size with a slightly different mix: some rows twice, some missing altogether. That is the nearest thing available to another 208 results that might have been published. Run the whole calculation on it and it gives its own answer. That is one draw:
rows drawn
208
same size as the real table
different rows in it
126
82 never got picked
its b
−0.0778
the real one is −0.0826
its answer
0.549
the real one is 0.438
That second card is why each row goes back in before the next is drawn: without that, a draw would hand back the same 208 rows every time in a different order, and every draw would give the same answer. The draw is the same size as the real table for the matching reason, that a smaller one would wobble more for a reason that has nothing to do with the data.
Do that 400 times and count where the answers fall. Each bar is one band of answers, and its length is how many of the 400 landed in it:
8 draws fell below this range and 8 above it, the furthest at −1.73 and 1.03. A draw that happens to leave one year nearly empty can return a wild rate. They are counted here rather than drawn, because bars wide enough to reach them would leave the shape above as a single column.
Sort those 400 answers smallest to largest and the range is two positions in the list. The 10% mark sits one tenth of the way along, the 90% mark nine tenths, and each is read between the two answers either side of it:
- position in 400
- 40.9
- the one below
- #40 · 0.2411
- the one above
- #41 · 0.2416
- halving time
- 15.0 months
- position in 400
- 200.5
- the one below
- #200 · 0.4299
- the one above
- #201 · 0.4299
- halving time
- 8.4 months
- position in 400
- 360.1
- the one below
- #360 · 0.6637
- the one above
- #361 · 0.6644
- halving time
- 5.4 months
So eight draws out of ten land between 0.242 and 0.664, a halving time anywhere from 5 to 15 months. The middle draw comes to 0.430, close to the 0.438 the real data gives. The average of all 400 is lower, at 0.409, because the few wild draws pull it down; the middle one ignores them, which is why the range above is quoted from positions in the list rather than from an average and a spread.
So treat 0.438 as the middle of a fairly wide range, not as a precise measurement. Epoch AI, who published the data and fit it a more careful way, report a halving every 8.4 months with a range of 4.5 to 14.3. Our answer and theirs sit inside each other's ranges, which is about as much agreement as this kind of estimate allows.
That range covers one thing only: how much the answer moves when the rows move. It does not ask whether the equation is the right shape, whether three language tests speak for robotics or driving, or what the years too thin to use would have done to it. Those are larger doubts than this one, and they are the four notes below.
Four things worth knowing before quoting this
- The people who published the data do it a harder way
- Epoch AI split the compute into two parts, the size of the model and the amount of text it read, because those are not interchangeable. They looked at the simpler version used here and set it aside as too crude. Our answer landing near theirs is agreement, not proof.
- It does not change any date on this site
- This rate is added to the rate at which training runs grow, and the capability map is then fitted against that total. Make this number bigger and the map's slope flattens by exactly as much. The forecast comes out the same either way.
- The early years are missing, and they matter
- 2012 to 2014 hold too few results to use. This is every row they have, and the gain each year would be credited with: 20121 row0.00020132 rows1.06420141 row1.461 4 rows in total, and they would carry 3 of the 12 points the line is fitted through. Counting them gives 0.549 instead of 0.438, which is a halving every 6.6 months instead of 8.2. A single 2014 model would be setting that year's whole gain. Filling those years is the most useful thing anyone could add.
- It is about language models, applied everywhere
- The results behind it are all language tests. The forecast uses the same rate for every domain, because nobody publishes a separate one for robotics or driving. That is an assumption, and it is written down here rather than buried.