AI Compute Supply Can't Keep Up With Demand

a16z argues AI demand is accelerating while compute supply is capped through 2028, keeping paybacks inside a year.

5 min read
Data center racks and power infrastructure illustrating AI compute supply constraints
a16z discussion argues demand is outrunning supply through 2028· a16z
Visual TL;DR
Compute supply cappedDriver
a16z sees demand accelerating while supply stays capped through 2028
From the articleAI compute supply is now the binding constraint, and a16z argues demand accelerated in July and August anyway.
Sub one-year paybacksOutcome
spot pricing and SpaceX cluster speed cut payback to nine or ten months
From the articleSpaceX compresses that further by building large clusters fast, which is why the guests see sub one-year paybacks on tens or hundreds of billions of deployable capex.
Demand keeps inflectingDriver
usage data kept inflecting up through July and August despite constraints
From the articleThe conversation frames the moment as the age of Elon and Jensen, with markets pricing fear while usage data keeps inflecting up.
Internal usage explosionCore
a16z firm token consumption rose 100x from March to August
Compute supply cappedDriver
a16z sees demand accelerating while supply stays capped through 2028
From the articleAI compute supply is now the binding constraint, and a16z argues demand accelerated in July and August anyway.
Demand keeps inflectingDriver
usage data kept inflecting up through July and August despite constraints
From the articleThe conversation frames the moment as the age of Elon and Jensen, with markets pricing fear while usage data keeps inflecting up.
Frontier labs monetize inferenceCore
From the articleFrontier labs monetize inference at roughly 60 billion dollars per gigawatt, and some are pushing toward 100 billion per gigawatt today.
Demand still tiny vs baseContext
From the articleDemand looks tiny against the base, with perhaps sub 10 million heavy token payers versus 1.5 billion knowledge workers.
Internal usage explosionCore
a16z firm token consumption rose 100x from March to August
Capex math compressedContext
customers prepay 50 to 60 percent and gigawatts cost about 50 billion
Sub one-year paybacksOutcome
spot pricing and SpaceX cluster speed cut payback to nine or ten months
From the articleSpaceX compresses that further by building large clusters fast, which is why the guests see sub one-year paybacks on tens or hundreds of billions of deployable capex.
Neoclouds benefitEffect
Nebius and CoreWeave see paybacks inside a year on tens of billions deployed
From the articleCustomers often prepay 50 to 60 percent, and spot pricing can cut payback to nine or ten months for neoclouds like Nebius and CoreWeave.

AI compute supply is now the binding constraint, and a16z argues demand accelerated in July and August anyway.

AI Compute Supply Can't Keep Up With Demand - a16z
AI Compute Supply Can't Keep Up With Demand, from a16z

The conversation frames the moment as the age of Elon and Jensen, with markets pricing fear while usage data keeps inflecting up.

How the crunch actually works

Frontier labs monetize inference at roughly 60 billion dollars per gigawatt, and some are pushing toward 100 billion per gigawatt today.

A gigawatt costs about 50 billion to bring online. Customers often prepay 50 to 60 percent, and spot pricing can cut payback to nine or ten months for neoclouds like Nebius and CoreWeave.

SpaceX compresses that further by building large clusters fast, which is why the guests see sub one-year paybacks on tens or hundreds of billions of deployable capex.

Demand looks tiny against the base, with perhaps sub 10 million heavy token payers versus 1.5 billion knowledge workers.

Internal token consumption at the firm rose 100x from March to August. Access to Grokbot Enterprise points to another 10 to 20x once agentic automation turns summarizers into recommended actions.

Financing masks the strain. Low-cost capital from Blackstone, KKR, and Apollo is betting useful lives extend as token ROI per gigawatt rises.

Why this matters, and what still breaks

Builders should price for scarcity, not abundance. No capacity is available through 2028, and political and regulatory delays point to underbuild, not overbuild.

Token prices could rise 10x if supply stays fixed. That would create compute inequality, where large firms can pay and everyone else waits. Advertising will take years to make a free tier viable.

The training-versus-inference trade is the hidden volatility driver. A shift from eight gigawatts on inference to eight on training would collapse annualized revenue from 480 billion to 120 billion in the hypothetical.

Orbital compute gets presented as swing capacity, not science fiction. Solar and radiators replace 15 billion in power and cooling, and launch falls below one billion per gigawatt if Starship reusability hits two launches per day per pad.

For enterprises, the diffusion signal is the power law inside companies. Top engineers spend 100x the median, and AI-native firms already spend high single digits to 10 percent of compensation on tokens.

For policymakers the pitch is industrial, not abstract China competition. Data centers use natural gas, negligible water, and 10x local tax revenue that revives small towns.

No patch exists for copper, wafers, power, or permitting. Every transformational tech has bubbled and overbuilt when debt demands immediate ROI.

The gap to watch is disclosure. Anthropic is in a quiet period, and OpenAI is holding the next checkpoint until Astra shows, which will set pricing and allocation.

© 2026 StartupHub.ai. All rights reserved. You may not republish this article in full without a license. Search engines and AI research tools may crawl and summarize for reference. Bulk reproduction or model training requires a license. See our terms.