Cheaper is not the same as more capable

Falling cost is necessary, not sufficient. Where efficiency does and doesn't move the needle.

The TowardSingularity Team · Oct 2026 · 3 min read

The price of a capable model keeps falling. This site reads it as the cheapest model on OpenRouter that also rates at least 1450 on LMArena, and on the latest pipeline run that was Gemma 4 31B at $0.09 per million tokens. Falling prices are a real and important trend. On their own, they do not mean AGI is closer.

To put that price in context: a million tokens is several novels’ worth of text. At $0.09, a model good enough to rate in the upper ranks of a public leaderboard will read or write that much for less than the cost of a stamp.

Why the reading needs a floor

“Cheapest model” sounds like a simple reading. It is not, and the first version of it on this site was close to useless.

Without a quality bar, the cheapest model on any given week is whatever tiny model a provider has decided to serve at cost. When the reading had only a context-length filter, the qualifying pool included a one-billion-parameter Llama. The number fell every week and told you nothing about what capable inference cost, only that small models are cheap.

So the reading is gated. A model only counts if it also rates at or above 1450 on LMArena’s leaderboard, which the pipeline fetches first for exactly this reason. Adding the gate moved the reading from $0.010 to $0.100 per million tokens, a tenfold change that came entirely from asking a better question.

The week the gate failed

On 24 August 2026 the LMArena leaderboard was down when the pipeline ran. With no ratings to check against, the cost reading fell back to the ungated cheapest price, which that day was $0.017 from a one-billion-class model.

That fallback is the designed behaviour, and the reading was flagged as a fallback. The bug was elsewhere. At the time, each run overwrote the cost file instead of adding to it, so the good gated reading from the previous run, $0.100 at a rating of 1451, was replaced by the bad one and lost. The fix was to make every daily reading append a row and keep one per date. A fallback now lands beside the real reading, marked as ungated, instead of on top of it.

It is a small story with a general point. A single price is easy to quote and easy to get wrong. The history of a price, with each point labelled by how it was taken, is much harder to fool yourself with.

Commodity, not ceiling

Falling cost mostly means today’s frontier becomes tomorrow’s commodity. It democratises capability and accelerates deployment, but a cheaper path to last year’s ceiling does not, by itself, raise the ceiling.

“Efficiency compounds access. Frontier capability is a different axis.”

Where falling cost does matter:

  • It lets labs run more experiments per dollar, raising the odds of a real breakthrough
  • It pulls capable models into agentic, always-on use cases
  • It frees compute budget to push the frontier rather than serve it

The second of those is easy to underrate. An agent that works through a long task makes many model calls, often hundreds, and each one reads back much of what came before. At a dollar per million tokens that is an expensive experiment. At nine cents it is something you leave running. Several of the capabilities the site measures, such as long software tasks done unattended, only become practical to use, as opposed to possible to demonstrate, once the price falls this far.

Why cost stays out of the date

None of which puts cost into the arrival model. Nothing in that chain reads a price at all: the efficiency trend it might be mistaken for is fitted on published test results instead, worked through step by step on its own page. Cost is published because how fast the frontier diffuses is worth knowing, not because it moves a date.

The reading is gated on a 1450 Elo floor, so that “cheapest” means the cheapest capable model rather than the cheapest tiny one. That floor will need raising. As small models improve, more of them will clear 1450, and the reading will slide back towards measuring the cheapest model that is merely good enough. When that happens the fix is a stricter bar, and the methodology says so in advance, so the change is a decision on the record and not a quiet adjustment.

ALL FIELD NOTES

See the data
behind
the notes.

Ten domains, each on its own measure, with the gaps published beside the measurements.