Historical source and event dates are not site publication dates. Product plans, policies and availability may have changed since retrieval.

The setup
Epoch AI runs a public database of machine-learning models, benchmark scores and training-compute figures that other reports, including some cited elsewhere in this publication, draw on as a shared reference. As retrieved on 16 September 2026, the database's overview page describes a collection of more than two thousand models spanning benchmarks, compute, chips and companies, published under a Creative Commons license and updated daily as a downloadable CSV. Using the tracker productively means knowing not just the listed number for a model's training compute, but how confidently Epoch AI stands behind it.
What the documents show
Epoch AI's estimation methodology page is unusually specific about this. When training compute is not directly reported, the methodology describes two main routes: working from known hardware type, quantity and training time, or counting floating-point operations implied by a model's architecture and training data. The hardware route requires assuming a utilization rate, typically 30 to 50 percent for large runs, when the developer has not published one; the page walks through a worked example estimating total compute from a stated chip-time figure. A separate section defines three confidence tiers, Confident, Likely and Speculative, corresponding to roughly 3x, 10x and 31x uncertainty bounds, stored alongside each figure in the dataset.
The friction
The same page is candid about how far some estimates lean on assumption: it names GPT-4 as a Speculative-tier estimate, built almost entirely from secondhand reporting about training duration and hardware rather than any figure the developer published directly. A third route, inferring compute from a model's benchmark scores against models of known compute, is usable only when enough comparison models exist, and the documentation states any figure derived this way should be excluded from research studying the relationship between benchmark performance and compute. None of this makes the dataset unreliable, but a headline number pulled without its confidence tier is missing information Epoch AI itself considers part of the figure.
What changed in the work
For a reader comparing two models' training scale, Epoch AI's documentation supports treating its dataset as an unusually transparent public source, precisely because it discloses its own estimation method rather than presenting every figure as equally measured. What it does not support is quoting a Speculative-tier figure with the confidence of a Confident one; the methodology page's own tiering exists to prevent that flattening, and dropping the tier when citing a number is a real loss of information the source supplies.
- Does the specific figure being cited carry a Confident, Likely or Speculative tag in Epoch AI's dataset?
- Was the figure derived from hardware, an architecture count, or a benchmark-score inference, and does that matter for the claim being made?
- Has the dataset been re-pulled recently, given it is described as updated daily?
Epoch AI's own methodology turns what could be a black-box number into a checkable one, at the cost of a reader looking past the headline figure to the confidence tier and estimation route the organization publishes alongside it.
Sources & verification
Preserved from the earlier archive. These sources have not all been freshly rechecked for this expansion.
- Data on the Trajectory of AISource date: not stated · Retrieved: 2026-09-16
Describes the scope of Epoch AI's public database of models, benchmarks and compute, updated daily and downloadable as CSV, as maintained on 16 September 2026.
- AI Models Documentation – EstimationSource date: not stated · Retrieved: 2026-09-16
Describes Epoch AI's methods for estimating training compute from hardware or architecture details, defines the Confident/Likely/Speculative confidence tiers, and gives GPT-4 as a Speculative-tier example.
Continue the workflow
- Run a workflow pilot that can answer a real question
Decide whether a proposed workflow deserves wider use using a modest, honest pilot.
- METR's benchmark tracks AI task length, not messy real work
METR's paper found the length of software tasks AI agents can finish has doubled roughly every seven months, with real-world caveats attached.
- Stanford's AI Index reports numbers it credits to other sources
The annual AI Index compiles labor-market and economy figures from named outside providers rather than measuring them itself.
- A model maker's own usage index counts conversations, not the economy
Anthropic's first Economic Index sorts a million Claude.ai conversations by occupation but says it cannot show AI use off its platform.