The State of the AI Economy
June 25, 2026
Azeem Azhar, William Gildea, Hannah Petrovic, PhD, Nathan Warren & Marija Gavrilov
1
Exponential View
www.exponentialview.co
Independently produced by Exponential View. Built from public disclosures and Exponential View’s own models; all conclusions are our own. �© Epiiplus1 Ltd 2026
2
There is a visibility problem in the AI economy. Until now, it has been impossible to deconstruct real customer demand.�
The supply side of the AI economy is well-documented. Most semiconductor companies and hyperscalers are public �and disclose their activities in some detail. Sell-side analysts have done a great job decomposing their performance.
The demand side, what customers are actually paying for and if the revenues are real, has been obscure.
The largest labs are private, and even public companies bury AI revenue inside segment totals.
Without understanding genuine demand, it is impossible to judge the health of the AI economy that underpins $22.7 trillion of stock market valuation and has driven US GDP growth in the past six quarters.
We hope that this report serves as a reference source on the current state of play, free of hype and fear, while helping us all have a more informed conversation about the gravitational pull AI is exerting on the economy and the world at large.
Special thanks to those who kindly reviewed an early draft of this presentation and gave us feedback:�Alex Imas, Shanu Mathew, Patrick Rutherford, Jaime Sevilla and Amy Sutter.
– Azeem and the Exponential View team
Why we’ve done this
3
We flag, investigate and maintain our datasets using analyst research, augmented by a proprietary system that scans, crawls and synthesizes insights.
Every revenue line traced to primary filings, audited accounts, transcripts and credible reporting; plus cloud-attribution where a private firm’s revenue surfaces in a public firm’s accounts (OpenAI via Azure, Anthropic via Bedrock).
We build full per-company models specifically for GenAI financials (split out from top-line reporting), covering key drivers of revenue, profitability and cost in the P&L, cash flows and balance sheets.
A proprietary line-level revenue model: Sourced, scored, triangulated, deduplicated
1
Source
Bottom-up, 1,000+ firms
›
2
Confidence score
Each line carries a rigorous confidence score before it enters any model, so weak inputs can’t inflate the number.
Confidence-scored before �it counts
›
3
Model and triangulate
Company financial models checked against top down
›
4
Deduplicate
Spend only counted once
Revenue is counted at every layer but never summed across them: attributed by value-add so the same dollar isn’t double- or triple-counted.
Audit trail: Sample source table
$40�value-add
$30�value-add
$30�value-add
e.g. $100 app spend that sends $60 to a model provider, which spends $30 on inference hosting, is counted as $100, not $190:
Scope: Global ex-China · App, model & infrastructure revenue counted · Excludes chips, AI ad-uplift, legacy-software features and financing.
$100 rev.
$60 rev.
$30 rev.
Apps
FM labs
Hosting
We grade filed figures highest, above other primary sources, corroborated 3rd-party estimates and single-sourced claims.
All derived numbers inherit the lowest grading from input sources.
These models are reconciled against independent proxies: silicon �(chip-maker revenue), build cost, �segment mix, industry research, �traffic and capacity.
Additional soft signals we use include unofficial sources such as:
Audit trail: Sample company revenue model
AI demand is more clearly validated by realized revenue than previous platform shifts. Generative AI ecosystem revenue has already surpassed $175 billion annualized (after removing double-counting from provider revenues).
CapEx intensity is growing well above historical large-cap technology norms to deliver the AI buildout. And third-party financing is increasingly entering the financing mix.
The open question is whether cheapening artificial intelligence can create enough volume and margin to service the buildout.
4
The top line
5
1 | Demand | Real, big and fast. External customers, real revenues, unprecedented growth. | 06-18 |
2 | Economy | Big is still small, and early. Gains exist, but they’re uneven and not measured. | 19-26 |
3 | CapEx | The biggest buildout in tech history is paying back (for now). | 27-38 |
4 | Tokens | The unit of value for the AI economy, or is it? | 39-51 |
5 | Stack | Where the value is captured. The stack turns capital and energy into cognition. | 52-62 |
Contents
1 | Demand:�It’s real, big & fast
6
Revenues are driven by real external customers. �The sector is growing 3x faster than any IT wave before it.
This demand has created a compute supercycle: 10x more compute, new energy generation, larger data centers and mounting backlogs where supply cannot keep up.
$110bn trailing 12-month revenues – now at a $175bn pace
Source: Exponential View analysis.�Note: Global ex-China. Deduplicated app, foundation model, and infrastructure hosting revenues. Excludes chip manufacturing.
Generative AI economy revenue, deduplicated
$bn/year, Jan 2023 - Jun 2026
1 | Demand: It’s real, big & fast
7
1.6x gap�the spread reflects the growth rate
$175bn�annualized run rate
$110bn�banked trailing�12-month revenue
Real, external demand drives AI revenues
1 | Demand: It’s real, big & fast
8
Illustration of how we model and deduplicate revenues between providers
App layer
Standalone GenAI apps
e.g. Cursor, OpenRouter, Harvey
Value add: Revenue minus token costs
Customer
pays app license/fees, or buys tokens directly
Hosting layer
$: AI revenues (full)
e.g. Azure, CoreWeave, Nebius
Foundation model layer
Closed-weight models
Value add: Revenue minus inference costs
e.g. Opus 4.8, GPT-5.5
Open-weight models
Value add: $0 (revenue counted in hosting)
e.g. Deepseek-v4, MiniMax M3
Chip sales (CapEx from hosting layer)
Ad uplift (Google/Meta AI ad revenue)
Non-AI-native apps (Counted via token spend)
Not counted:
Value add: Revenue minus token costs
e.g. Claude Code, Codex
Model lab’s apps
Revenues flow down the stack
We source, triangulate, model & audit to verify & deduplicate:
Rigorous company-by-company financial modeling to build a bottom-up deduplicated revenue model.
Continuous scanning and crawling across hundreds of sources to maintain and adjust the dataset.
Sourced from official filings, 1st-party disclosures, leaks, government stats, 3rd-party analysts, and proxy metrics; all sources quality-graded.
AI is scaling three times faster than any IT wave
Sources: Exponential View analysis; US Commerce; company filings; UBS; US Bureau of Labor Statistics.�Note: We use the first full year of revenues, so have started GenAI measurement from January 2023.
Realized revenue trajectory time-aligned to year zero
$bn/year, adjusted for inflation
1 | Demand: It’s real, big & fast
9
3x faster�than any prior wave
Indexing GenAI growth to past technology waves risks understating its speed when modeling:
Each new $1 billion of revenue arrives faster than the last
Source: Exponential View analysis.�Note: Global ex-China. Deduplicated app, foundation model, and infrastructure hosting revenues. Excludes chip manufacturing.
Time to add $1bn additional cumulative revenue
days, log scale
1 | Demand: It’s real, big & fast
10
In 2023, the AI industry needed 180 days to add �$1 billion in cumulative revenue.
It now needs less than �two days.
90x faster
Growth has held across each adoption phase
Source: Exponential View analysis.
Revenue growth quarter-on-quarter
% change since prior quarter
1 | Demand: It’s real, big & fast
11
Agentic Coding Era
Chatbot Subscription Era
35% QoQ
equivalent to
3.2x annually
ChatGPT Team launches
Claude Enterprise launches
Claude Code released
Codex launch
OpenClaw �goes viral
This rapid demand growth is showing as a contract backlog �for hyperscalers
Sources: Exponential View analysis; company filings.
Note: Microsoft = total RPO (incl. M365/Dynamics); Amazon = total company RPO (mostly AWS); Google = revenue backlog (mostly Cloud). Oracle quarter ends one month earlier.
Combined hyperscaler backlog (remaining performance obligations)
$bn
1 | Demand: It’s real, big & fast
12
Demand has launched a compute supercycle
Sources: Exponential View analysis; World Semiconductor Trade Statistics; Japanese Semiconductor History Museum of Japan.
Note: Nominal $US
Global semiconductor market revenues
$bn/year
1 | Demand: It’s real, big & fast
13
WSTS 2026 projection:
AI has spurred an uptick in a 50-year trend of compute growth
Source: Exponential View analysis.
Note: Includes mainframes, minicomputers, PCs, servers, smartphones, IoT and AI compute. AI-server FLOPS derived from the installed base of Nvidia GPUs by generation, using FP8 from Hopper onward.
Global compute since 1971
FLOPS
1 | Demand: It’s real, big & fast
14
66% CAGR
80% CAGR
AI demand is reigniting a moribund US power sector
Sources: Exponential View analysis; US Energy Information Administration.
US electricity net generation
TWh/month
1 | Demand: It’s real, big & fast
15
2008-2024: ±0 growth
1950-2008: +6 TWh/month annual growth
2024-today:�+9 TWh/month annual growth
The size of the largest data centers has grown 50x in four years
Sources: Exponential View analysis; TOP500 / Green500; Epoch AI; OpenAI; Oracle.
Note: Rainier full build is a target.
1 | Demand: It’s real, big & fast
Power of the most powerful computers over time
MW, log scale
16
Memory and compute now take a majority of every dollar spent on the data center buildout
Sources: Exponential View analysis; Goldman Sachs; Epoch AI; Semi Analysis. �Note: Figures may not sum to 100% because of rounding.
Share of total data center build cost by component, 2021 vs 2026E
%
1 | Demand: It’s real, big & fast
17
Leading to growing commitments for compute and energy
Sources: Exponential View analysis; Grid Strategies; Nvidia filings.
Compute & power commitments
Nvidia supply commitments ($bn, left axis) & US load-growth (GW, right axis)
1 | Demand: It’s real, big & fast
18
2 | Economy:�Big is still small, and early
19
Even for the highest corporate spenders, AI is a rounding error in the P&L. �It still looks early. Initiatives have focused on efficiency & cost savings, although the mix is changing. And measured revenue may understate the social gains, as consumers report benefits that don’t yet show up in the data.
Against GDP, AI revenue is still a rounding error
Sources: Exponential View analysis; St. Louis Fed.
Global AI revenues (ex. China), relative to US GDP, labor costs & corporate profits
%
2 | Economy: Big is still small, and early
20
3.0%�Profits
0.8%�Labor
0.4%�GDP
At a company level, AI spending is still relatively small:
e.g. Uber’s $1.5k per engineer barely dents the P&L
Sources: Exponential View analysis; Ramp Economics Lab (n = 70,000 US businesses), Uber filings.�Note: Top 1% / top 10% / median defined by level of AI spend compared across Ramp’s customer base. Uber figure is a max per-engineer cap, benchmarked against Ramp per-employee AI spend.
AI spend per employee
$/month, log scale, Ramp customers vs Uber cap
Uber AI spend (maxed cap) vs P&L line items
$/year (AI spend %), log scale, vs FY2025
2 | Economy: Big is still small, and early
21
$1.5k per-engineer cap roughly puts Uber in the top 10% of per-employee AI spend
$90m
$52bn(0.2%)
$8.7bn(1.0%)
$14bn(0.6%)
$3.4bn(2.6%)
$720m(12%)
Like previous general-purpose technologies, some gains may escape GDP measurement
2 | Economy: Big is still small, and early
22
Consumer surplus
Value that reaches people directly at a near-zero price. Little is sold, so GDP under-represents consumer benefit, e.g.:
Producer surplus
Value embedded in sold goods and services. More is transacted and recorded in GDP, e.g:
1980-2000: Automation
Programmable machine tools cut labor in every unit, raising output per worker.
Graetz & Michaels (2018)
+0.37
pp/yr GDP
+0.4
pp/yr GDP
1850-1870: Steam
Mechanized factories and railways, so one worker could produce and move far more to market.
Crafts (2004)
1880-1920: Electric lighting
Light became ~99.97% cheaper: �an hour’s wage buys ~40,000x more. Prices didn’t record this gain.
Nordhaus (1996)
≈ $0
direct GDP impact
2000-2020: Free digital goods
Free search, encyclopaedias and maps displaced paid services. Search alone �is worth ~$17.5k/yr/person.
Brynjolfsson et al. (2019)
≈ $0
direct GDP impact
AI impacts
Historic cases
GDP knows the price of everything but the value of nothing.�AI’s economic value exceeds measured revenue
Sources: Exponential View analysis; Stanford Digital Economy Lab
Note: Welfare value determined from responses to “Would you give up access to all AI tools like ChatGPT, Gemini, Claude, or Copilot �for one month starting tomorrow morning in exchange for [$US]?”. Revenue includes global (ex-China) consumer and enterprise spend.
Monthly GenAI revenues vs US consumer welfare
$bn/month, Jul 2024 - Mar 2026
2 | Economy: Big is still small, and early
23
Revenue: How much is spent on AI?
$3-4bn consumer surplus
(+30% on revenues)
Consumer welfare: What value do consumers place on AI?
Public companies are reporting increased impact of GenAI
Sources: Exponential View analysis; earnings calls.
2 | Economy: Big is still small, and early
Companies making claims of AI impact on earnings calls
S&P 500, Q4 2022 - Q1 2026
24
Seven in ten GenAI claims focus on cost savings or efficiency
Why initial projects prioritize efficiency, an illustrative example:
The same pure $1m impact has 4pp better profit margin from savings vs sales growth
Sources: Exponential View analysis; earnings calls. �Note: Figures may not sum to 100% because of rounding.
Claimed AI outcomes
S&P 500, Q4 2022 - Q1 2026
2 | Economy: Big is still small, and early
Note: For illustrative purposes, revenue is added at $0 cost. Adding costs-of-sales would dampen margin growth further.
25
$11m revenue • $7m cost
$4m profit
-$1m costs
Cost reduction: 25%
Time savings: 23%
Throughput increase: 22%
Quality improvement: 18%
Conversion improvement: 7%
Revenue gain: 6%
TODAY
30% net margin
$10m revenue • $7m cost
$3m profit
COST SAVINGS
40% net margin
$10m revenue • $6m cost
$4m profit
SALES GROWTH
36% net margin
+$1m revenue
As in prior waves, early adopters are outgrowing their peers
Sources: Exponential View analysis;�Bessen, Goos, Salomons & Van den Berge (2020):�“What Happens to Workers at Firms that Automate?”.
Sources: Exponential View analysis; Ramp Economics Lab �(n = 70,000 US businesses).�Note: High intensity = Top 25% AI spenders by share of revenue.
The historic case study:
Firm-level employment with/without automation
% change 2000-2016
The AI economy today:
Revenue growth with high vs no AI usage
% change since Nov 2022
2 | Economy: Big is still small, and early
26
Δ37%
Δ92%
3 | CapEx:�The largest buildout in tech history is paying back (for now)
27
Hyperscalers & neoclouds have committed to $2 trillion of cumulative CapEx to 2026, putting pressure on growing revenues to pay back, especially as more is funded by external capital. These economics set the tone for data center and token production finances.
Hyperscaler and neocloud CapEx reaches $2T cumulatively through 2026E
3 | CapEx: The largest buildout in tech is paying back (for now)
Sources: Exponential View analysis; company filings.
Note: 2026 is based on guidance values. Oracle uses a 5/12:7/12 split based on FY2026 reported and FY2027 guidance values.
Hyperscaler and neocloud CapEx
$bn, PP&E + leases
28
Total CapEx ≠ AI CapEx.
Announced numbers include pre-planned CapEx for existing cloud & SaaS businesses, metaverse (Meta), and logistics (Amazon)
AI-linked CapEx adds $535bn above the pre-AI trend by 2026E
Sources: Exponential View analysis; Hyperscaler earnings & guidance; US telecom & carrier guidance; The International Energy Agency.
3 | CapEx: The largest buildout in tech is paying back (for now)
Annual CapEx per industry
$bn
29
$535bn above pre-AI trend
Forecasts have chased the CapEx curve higher
3 | CapEx: The largest buildout in tech is paying back (for now)
Sources: Exponential View analysis; Barclays, Citi, Goldman Sachs, JP Morgan, Morgan Stanley, New Street Research, SemiAnalysis, UBS; company filings.
Note: Forecasters use slightly different baskets (the Big Five hyperscalers vs broader AI infrastructure). Actuals here correspond to the cash CapEx (purchases of property and equipment) for Microsoft, Alphabet, Amazon, Meta, Oracle, CoreWeave, and Nebius.
Hyperscaler / AI infra CapEx forecasts by analyst and forecast date
$/year
30
The marginal AI-infra dollar is increasingly externally financed
Sources: Exponential View analysis; company filings.�Note: Debt is net of repayments (not gross issuance) and includes all debt instruments (bonds, commercial paper, etc.).
3 | CapEx: The largest buildout in tech is paying back (for now)
Hyperscaler and neocloud CapEx by funding source
$, 2020-2026E
CapEx by funding source
%, 2020-2026E total
31
Hyperscaler CapEx primarily cash, but risk to wider economy from their market cap weight �vs rest of market
Leases
Debt (net new)
Equity
Cash
Operating cash flow
Paying cash keeps a bad bet inside the firm, only denting profits
External funding moves risk outside firms, �as third-parties expect repayment
Neocloud CapEx primarily debt-funded
The 2026E depreciation charge approaches $111 billion
3 | CapEx: The largest buildout in tech is paying back (for now)
Source: Exponential View analysis.�Note: IT equipment is depreciated over 6 years, and buildings over 14 years. Required revenue values exclude OpEx.�Headroom is the portion of revenue beyond that required to meet the depreciation expense.
32
CapEx is expensed through depreciation over the assets’ useful life. So the cost is spread and doesn’t need to be recognized instantly
CapEx is spent throughout the year (not all on 1st Jan), so the full depreciation charge doesn’t hit fully in year 1
Revenues cover the ongoing expense, not yet the cumulative bill
3 | CapEx: The largest buildout in tech is paying back (for now)
Sources: Exponential View analysis; company filings.
Note: Meta contributes to industry CapEx but initiatives are focused on ad uplift, so not recognized as pure GenAI revenue, or currently have minimal direct monetization (e.g. Meta AI assistant, Muse Spark).
Quarterly AI revenues & CapEx depreciation
$bn/quarter, hyperscalers & neoclouds only
Cumulative AI revenues & CapEx depreciation
$bn, hyperscalers & neoclouds only
33
Still ~half-covered: cumulative revenue has nearly covered cumulative depreciation, but still has to cover the expected headroom
Q4 2025: Quarterly revenues first exceed CapEx depreciation
AI infra revenue now just clears today’s depreciation hurdle
3 | CapEx: The largest buildout in tech is paying back (for now)
Source: Exponential View analysis.
Headroom after quarterly CapEx depreciation
% = (Revenue – Depreciation) ÷ Revenue
34
All GenAI revenues�32%
19%�Hyperscaler &�neocloud revenues only
Q4 2025: Quarterly revenues first exceed CapEx depreciation
Rental rates suggest demand is absorbing existing supply
3 | CapEx: The largest buildout in tech is paying back (for now)
Sources: Exponential View analysis; SemiAnalysis.
H100 1-year rental contract price
$/hour/GPU. H100 is the most liquid, most-traded GPU in the merchant market: a useful signal.
35
$3.05: Launch-era scarcity premium
$1.70: Peak overbuild fears as Blackwell ships
$2.40: Surging inference demand
Data center economics set the hurdle for token pricing
Sources: Exponential View analysis; Epoch AI; SemiAnalysis.��Note: Illustrative model. 1GW of IT capacity (6.7k GB200 NVL72 systems, 480k GPUs), cost of ownership per Epoch AI (May 2026), annualized incl. cost of capital, 6-year IT life. Token output from SemiAnalysis InferenceX: FP4, 8k-in/1k-out, 50 tokens/sec/user, 65% utilization (Apr 2026). Token output range reflects +10-25% throughput uplift from speculative decoding. Open-weight models incur no model-licensing fee. Closed-weight column adds a 25% licensing fee of Kimi’s $1.29 blended price.
3 | CapEx: The largest buildout in tech is paying back (for now)
36
REQUIRED CUSTOMER VALUE FOR 25% ROI
● Kimi K2.5 (1T) ● Kimi-class model under closed licensing (illustrative)
DIVIDE $7.9bn/yr BY TOKEN OUTPUT:
÷ TOKENS PER GW / YEAR
75-85 quadrillion
= COST FOR INFERENCE PROVIDER
$0.10
$0.42
REQUIRED PRICE FOR 50-75% GROSS MARGIN:
per 1m tokens
$0.20 – $0.40
$0.84 – $1.68
$0.25 – $0.50
$1.05– $2.10
$7.9bn annual cost to own and operate 1 GW of AI capacity
Capital costs $7.0bn/year · 89%
65%
20%
Servers $4.6bn
480,000 GPUs in 6,700 systems · 6-year life
Facility $1.4bn
Building shell, power & cooling · 14-year life
DC network $1.0bn
Switching & interconnect fabric · 6-year life
Land + utility $33m
Land at cost of capital · utility 14-year life
OpEx $900m/year · 11%
66%
Energy $594m
Electricity to run the fleet · 66% of OpEx
Other OpEx $308m
Staff, maintenance & overhead
per 1m tokens
per 1m tokens
$0.32 licensing fee (~25% of the blended selling price of $1.29)
Gross rental yields suggest useful lives extend past six years
3 | CapEx: The largest buildout in tech is paying back (for now)
Sources: Exponential View analysis; Silicon Data.�Note: Yield = (On-demand rate x 50% utilization x 8760 hours) ÷ original list price.
GPU yield at 50% utilization
%, excludes OpEx
37
Older GPUs�earn yields long beyond their six-year depreciation life
Years since release:
1yr
6yr
7yr
8yr
9yr
Depreciation charge (6-year): 17%
Newer GPUs earn yields well above depreciation charge
3yr
4yr
After 6 years, chip CapEx fully depreciated
Longer GPU useful life stretches headroom
3 | CapEx: The largest buildout in tech is paying back (for now)
Sources: Exponential View analysis; Meta.�Note: Buildings depreciated over 14 years as constant.
Range of headroom after CapEx per chip depreciation schedule
Q1 2026, 3-9-year schedules, % = (Revenue – Depreciation) ÷ Revenue
38
If chip life is shorter, revenues don’t repay CapEx (this would require initial H100 purchases becoming obsolete today)
Mark Zuckerberg
Meta Q3 2025 earnings call
“... the kind of very worst case would be that we effectively have just prebuilt for a couple of years, in which case, of course, there would be some loss and depreciation, but we’d grow into that and use it over time.”
Overbuild can be a bet on longer chip depreciation
Headroom (infrastructure revenues only)
Headroom (whole market revenues)
Chip depreciation schedule | ||||||
3y | 4y | 5y | 6y | 7y | 8y | 9y |
At standard 6-year chip life, Q1 2026 infrastructure revenue headroom is 19%
Extending useful chip life boosts margins: Using chips for 8 years (as old as T4s) raises infrastructure headroom to 36%
4 | Tokens:�The unit of value for the AI economy?
39
Token volumes are growing 14x annually, propelled by agentic workloads and highly elastic demand. Token-based pricing has made this especially pertinent, but it also represents an opportunity for the industry to attribute and evaluate the output from token consumption.
Is the GenAI economy a token economy?�Sort of.
40
“The input is electrons, the output is tokens. In the middle is Nvidia.”
– Jensen Huang
“Tokens, the fundamental units of data our models process…”
– Sundar Pichai
Global token volumes exceed 30Q/month, growing 14x YoY
4 | Tokens: The unit of value for the AI economy?
Sources: Exponential View analysis. �Note: Global, inc. China
Inference tokens processed
Quadrillion tokens per month (left axis), growth rate multiple (right axis)
41
Year-on-Year change
Total (API + subscription + internal)
API-only
Jan 23 | Apr 23 | Jul 23 | Oct 23 | Jan 24 | Apr 24 | Jul 24 | Oct 24 | Jan 25 | Apr 25 | Jul 25 | Oct 25 | Jan 26 | Apr 26 |
The transition from chat to agents is multiplying token use
Sources: Exponential View analysis; OpenRouter; Bai, Huang, Wang, Sun, Mihalcea, Brynjolfsson and Pentland 2026.
4 | Tokens: The unit of value for the AI economy?
Agent coordination density
% tool use per prompt, OpenRouter
Token consumption per task
Average tokens consumed per task, log scale
42
Cheaper tokens and better models amplify demand
Sources: Exponential View analysis; Epoch AI.
4 | Tokens: The unit of value for the AI economy?
43
Tokens processed per output token:
12 → 36
Blended price per million tokens:
$17 → $2
Epoch Capabilities Index:
112 → 158
Token demand appears elastic: As prices fall, usage grows faster
4 | Tokens: The unit of value for the AI economy?
Sources: Exponential View analysis; Google; OpenAI; ByteDance.�Note: Time-series, not cross-sectional: price and usage both trend with time, so β may overstate pure price-elasticity.
Google price elasticity (avg. price vs volume)
$/million tokens vs trillion tokens/month, log-log scale, 2023-2026
44
Across providers, magnitude of elasticity ≈ 1.2-1.8:�every 10% price cut → 12–18% more tokens → total token spend still rises
Sam Altman OpenAI: “Three Observations”, 2025
“The cost to use a given level of AI falls about 10x every 12 months, and lower prices lead to much more use”
Price −90% | Volume ↑
Sundar Pichai Google: I/O 2025
“…we were processing 9.7 trillion tokens a month. Now, over 480 trillion — 50x more.”
Price −97% | Volume 50x
Tan Dai Volcengine / ByteDance, 2025
“Doubao’s daily token usage exceeded 50 trillion this month, up from �4 trillion in Dec 2024.”
Price −50% | Volume 12x
trend β = -1.7
R2 = 0.93
Increasing usage
Decreasing prices
Token-based pricing is AI’s ‘pay-per-click’ moment
4 | Tokens: The unit of value for the AI economy?
Sources: Exponential View analysis; IAB/PwC Internet Advertising Revenue Report.
Evolution of AI pricing models
Annual digital ad revenue
$bn/year, log scale, 1996-2024
45
Attribution with CPC enabled a sustainable and�profitable market. Spend could be linked to ROI.
Untracked banner ads had a limited market size and crashed with the dot-com bust
2002: Pay-per-click (Google AdWords)
1
FREE
Unrecognized in budgets
2
SUBSCRIPTION
Seat-based pricing without tracing to specific value
3
TOKEN-BASED PRICING
Metered usage enables (and requires) attribution to projects
$500m
Each generation lifts token output per gigawatt of capacity
Source: Exponential View analysis. Global ex-China.
4 | Tokens: The unit of value for the AI economy?
Tokens produced per gigawatt of data center capacity
Trillion tokens per GW per month
46
This efficiency is increasing monetization per GW of capacity while revenues per token fall
Source: Exponential View analysis. Global ex-China. Annualized revenue.
4 | Tokens: The unit of value for the AI economy?
Revenue generated from tokens & data center capacity
$m/TTok (left axis), $bn/GW (right axis)
47
Tokens are AI’s billing metric, but not yet a unit of value
4 | Tokens: The unit of value for the AI economy?
48
INTERNET
Pageviews
Sessions & clicks
Advertisers pay on attention and conversion:
CPM → CPC → CPA
MOBILE INTERNET
Megabytes
Active users
Apps valued on engagement:
DAU / MAU, retention
GEN AI
Tokens
“Intelligence”?
The value-producing unit is still undefined
ELECTRICITY
Lamps
Kilowatt-hours
Edison’s first customers paid per light bulb installed. Metering came later.
Quality-adjusted tokens come closest to a usable unit of value
4 | Tokens: The unit of value for the AI economy?
49
Output tokens
Input tokens
Reasoning tokens
Input & reasoning tokens are a cost of production: exclude them
More consumption here often buys more capability, but it’s not the user-facing output
Capability
×
Scale for quality of output
Output from “better” models is more valuable to users
=
Intelligence
Quality-adjusted �output tokens
Volume of useful output
Users draw on intelligence to complete tasks of increasing complexity
Token volume
Token value
Regardless of the measure you pick, the trend is on the up
Sources: Exponential View analysis; Epoch AI; METR; Artificial Analysis.�Note: METR tasks primarily consist of software engineering, machine learning, and cybersecurity tasks.
4 | Tokens: The unit of value for the AI economy?
Epoch Capabilities Index
Score: GPT-5 = 150,�Claude 3.5 Sonnet = 130
METR Task Horizon
Human task duration with 50% model success rate, log scale
Artificial Analysis Intelligence Index
Score: 0-100
50
Quality-adjusted output kept pace with raw volume growth
Sources: Exponential View analysis; OpenRouter; Epoch AI; METR; Artificial Analysis.�Note: ECI, METR & AAII indexed to average Jan 2025 scores for comparative purposes.
4 | Tokens: The unit of value for the AI economy?
Tokens: Total, output & quality-adjusted
Trillion tokens per month, Jan 2025 vs Apr 2026
51
Total tokens
Output�tokens
+39x
+30x
+33x
+59x
+35x
Total tokens
Output tokens
ECI-adjusted
METR-adjusted
AAII-adjusted
5 | Stack:�Where the value is captured
52
Revenue is concentrated, but apps and models are gaining share. Labs retain pricing only while they hold the frontier; the economic value of yesterday’s frontier diminishes quickly into open weights. Where margin accrues across the stack will depend on technical advancements, which cut across competitive dynamics.
The stack turns capital and energy into cognitive work
5 | Stack: Where the value is captured
53
Chips
e.g. Nvidia, Huawei,�AMD, Cambricon, Cerebras
Convert energy to�tokens
Hosting
e.g. CoreWeave, Nebius,�Azure, AWS�
Convert chip CapEx�to token OpEx
Foundation Models
e.g. OpenAI, Anthropic
�
Convert tokens to intelligence
Apps
e.g. Cursor, Harvey, OpenRouter
Convert intelligence to consumer value
Revenue is concentrated today, but the mix is shifting
Sources: Exponential View analysis; company filings. �Note: Revenues not subject to deduplication adjustment.
5 | Stack: Where the value is captured
Annualized GenAI revenue
$million, Q1 2024
Annualized GenAI revenue
$million, Q1 2026
54
Apps
Foundation Models
Hosting
Chips
Value is moving up the stack, towards apps and models
5 | Stack: Where the value is captured
Source: Exponential View analysis.�Note: Global ex-China. Excludes chips (which are capitalized as CapEx by the hosting layer). �Figures may not sum to 100% because of rounding.
Quarterly GenAI revenues, deduplicated by layer to reduce double-counting
$bn per quarter
55
2.95x in one year
7%
11%
82%
4%
9%
87%
3%
8%
89%
Pricing power follows competitive pressure, not stack position
Source: Exponential View analysis. �Note: Hosting & Apps are defined by share of revenues. Foundation Models are defined by share of tokens (due to open-weight competition). Chips are defined by share of compute (H100-eq) due to vertical integration of dedicated hyperscaler chips. Herfindahl–Hirschman index is a measure of market concentration.
5 | Stack: Where the value is captured
Market share of leading companies & Herfindahl–Hirschman Index (HHI)
% per layer (left axis) & HHI (right axis)
56
Largest provider
Others
HHI
Frontier labs can defend premium pricing, for now
Sources: Exponential View analysis; Epoch AI.
5 | Stack: Where the value is captured
Epoch Capabilities Index
Score at release, OpenAI/Anthropic/open-weight
Output price
$/million output tokens
57
OpenAI
Anthropic
Others
Labs must outrun open-weight commoditization �to hold margin
5 | Stack: Where the value is captured
58
Lightweight open-weight
Summarize an article
Leading open-weight
Draft working code
Closed-weight frontier
Multi-day agentic work
Last year’s frontier is commoditizing fast
Sources: Exponential View analysis; Epoch AI.
5 | Stack: Where the value is captured
Price per capability frontier
Blended price $, grouped by performance at GPQA Diamond (PhD-level science)
59
Labs earn a time-limited premium on the current frontier
Among self-selecting OpenRouter users, token share is moving to open-weight
Sources: Exponential View analysis; OpenRouter.�Note: While OpenRouter is not a cross-section of the market, its data shows the behavior of self-selecting “model-routing” users.
5 | Stack: Where the value is captured
Weekly OpenRouter token share
% per model author
60
72%
33%
Google + OpenAI + Anthropic:
Under pricing pressure, labs push into apps and infrastructure
5 | Stack: Where the value is captured
Sources: Exponential View analysis; TechCrunch; OpenAI; Anthropic.
61
Hosting
Apps
Foundation�models
Consumer�surplus
Today
Building vertical apps: Law
Claude for Legal
Codex for Legal
Compress every layer, and consumers capture the surplus
5 | Stack: Where the value is captured
62
Hosting
Apps
Foundation�models
Consumer�surplus
Today
Scenario 2: General-purpose models
Scenario 3: Reduced compute needs
Scenario 4: All 3 combined
Scenario 1: Open-weight models catch up
Genuine demand and price elasticity of demand:�Cost reductions grow the market and result in increased consumer surplus. |
Tokens are cheaper for apps to consume & �hosting to serve
Foundation model labs build application layer�and their own infrastructure
Frontier model labs absorb integration work; generic "wrapper" pricing power is squeezed
Models and apps cost less to run, growing the applicable market
Cost reductions grow the applicable market
Significant consumer surplus as benefits delivered beyond price
Some local hosting, with chips distributed�to edge devices
Apps defend with proprietary data, domain-specific workflows & own models
Labs unable to charge license fee for �frontier models
Frontier models can replace AI application/integration layer
Models become smaller, and compute �more efficient
The perfect storm of efficiency improvements
AI demand is more revenue-validated than �any prior platform shift.
The investment case comes down to whether falling prices can move enough token volume to earn a return on CapEx.
63
64
What we count in, and what we exclude
How we deduplicate, and why
Sources
All figures are built bottom-up from primary and specialist sources, and triangulated against top-level estimates and proxies. Revenue and CapEx are grounded on company filings (SEC 10-K, 10-Q and 8-K) and executive disclosures, cross-checked through cloud attribution (e.g. Azure and Bedrock) where a private company’s revenue appears in a public firm’s accounts. Our systems daily scan and crawl available sources to create, maintain and improve the breadth and depth of our data, with source attribution and confidence grading. This high-quality dataset is used to build up full financial models (including P&L) for major companies, and driver-based models for smaller companies.
CapEx’s AI share is carved per company and reconciled against silicon (chip providers’ revenue), build cost, segment composition, and sell-side research. �Token volumes are reconciled from executive statements, third-party sampling and analysis, traffic volumes, and triangulated against revenues and available compute capacity.
Methodology: How we count revenues, CapEx, tokens
Authors
Marija Gavrilov
Azeem Azhar
Hannah Petrovic, PhD
Nathan Warren
William Gildea
Managing Director
Founder
Senior Researcher
Senior Researcher
Product Manager
We welcome feedback and contributions �at aieconomy@exponentialview.co
For advisory requests and institutional inquiries, please contact helen@exponentialview.co
Follow our analysis on
Exponential View
66
Subscribe to Exponential View to receive our research in your inbox each week