This is the cheat sheet from the carousel. Every number in the timeline comes from one source: Stanford's AI Index, the closest thing AI has to an official annual report card. It runs roughly 3,000 pages across 8 editions. I read them so you don't have to. Below: every stat with its source edition, plus the five numbers people cite wrong.
How to read the sources
Each edition covers the previous year's data, so the "2026 report" (published April 2026) is about 2025. There was no 2020 edition; the report skipped a year when it moved from a December to a spring release, so the 2021 edition covers the GPT-3 year. Training costs are compute-only estimates (via Epoch AI), not full program budgets.
The timeline
| Year | The number | What happened | Source |
|---|---|---|---|
| 2017 | ~$900 | Eight Google researchers publish "Attention Is All You Need", the nine-page Transformer paper. Stanford's estimate of the original training run's compute cost. GPT, Claude, Gemini, Llama, and DeepSeek all descend from it. | AI Index 2024 (first edition to publish cost estimates) |
| 2018 | 3.4 months | OpenAI measures compute in the largest training runs doubling every 3.4 months, more than 300,000x since 2012. Moore's Law over the same span delivered about 7x. | OpenAI "AI and Compute", cited in AI Index 2019 |
| 2019 | 9 months | OpenAI calls GPT-2 "too dangerous to release", then releases it anyway nine months later. Nothing bad happens. The staged-release playbook is born. | OpenAI announcements, Feb and Nov 2019 |
| 2020 | 175B params | GPT-3 proves that scale keeps working, which turns compute into a moat. The same year, AlphaFold 2 wins CASP14 and settles a 50-year protein-structure problem. The hinge year nobody watched. | AI Index 2021 |
| 2021 | 90.3 vs 89.8 | Microsoft's DeBERTa passes the human baseline on SuperGLUE, a benchmark built specifically because models had saturated the previous one. The report warns AI has "outpaced the benchmarks to test for them." | AI Index 2022 |
| 2022 | Nov 30 | ChatGPT ships. The model wasn't new; the interface was. Enterprise adoption had been flat at 50 to 58% for three straight years before it. | AI Index 2023 |
| 2023 | ~$78M | GPT-4's estimated compute bill, up from $900 six years earlier. Gen-AI funding jumps about 8x, from roughly $3B to $25.2B, while the aggregate headline said investment fell. | AI Index 2024 |
| 2024 | 4.4% to 71.7% | SWE-bench, which is real GitHub issues, in a single year. The steepest capability jump in the report's history. Adoption breaks its plateau, 55% to 78%, and US federal AI regulations go from 25 to 59. | AI Index 2025 |
| 2025 | 17.5 to 0.3 | The US-China gap on the main benchmark collapses to near-zero as DeepSeek's R1 reaches near-parity, with the US outspending China about 23 to 1. Same period: GPT-3.5-level inference falls from $20 to $0.07 per million tokens, a 280x drop. | AI Index 2025 and 2026 |
| 2026 | 88% vs ~5% | 88% of companies use AI, but only about 5 to 6% report material profit impact. The same models that win math-olympiad gold read an analog clock right 50.6% of the time (humans: 90.1%). Capability got cheap; trust is the scarce thing now. | AI Index 2026 |
The five numbers everyone cites wrong
These stats are real but usually quoted with the caveat stripped off. If you repeat one at work, repeat the fine print too.
- "DeepSeek trained a frontier model for $5.6M."That figure is the final successful training run only, on 2,048 GPUs. It excludes R&D, failed runs, and the hardware itself. Independent estimates put the full program between $500M and $1.6B. The real story is efficiency, not that frontier AI got cheap.
- "AI investment crashed in 2022-23." The aggregate did dip, mostly for macro reasons. Underneath the dip, generative-AI funding was multiplying 8x. The money was re-allocating, not leaving.
- The 280x price drop took about two years, not 18 months.The report's own dates run November 2022 to October 2024. Still absurd, just quote it right.
- "Gemini Ultra reached human parity on MMLU." That result used a non-standard evaluation setup. Under the standard setting it scored 83.7%, below the human baseline. If you want a clean human-parity stat, use DeBERTa on SuperGLUE (2021) or the SWE-bench jump instead.
- "Compute grew 7x faster than Moore's Law."Not what the analysis says. Compute grew more than 300,000x over six years; roughly 7x is what Moore's Law alone would have delivered in that span. The 7x is the comparison baseline, not a multiplier.
Go to the source
Every edition is free at aiindex.stanford.edu. Editions: 2018, 2019, 2021, 2022, 2023, 2024, 2025, 2026 (no 2020 edition). One warning if you dig in: the report occasionally revises its own historical totals when it switches data providers, so within-edition year-over-year comparisons are more reliable than comparing absolute numbers across editions.
This page is the companion to my AI Index carousel on Instagram (@nicnonac). If a stat here doesn't match something you've seen elsewhere, check whether the other source stripped one of the caveats above.