← Back to Horse Energy

So Who's Going to Buy All These Tokens?

Assume every AI model works exactly as promised and every demo ships. The spending still doesn't pencil out, because in a competitive market the buyer never gets to keep what the tokens save.

August 10, 2026 | AI Economics | Data Analysis

Every AI bear case you've read is a bet against the technology. The models are overhyped, the agents don't work, the pilots quietly die. Ed Zitron has built a whole beat on it.

This piece makes the opposite bet. Assume the models work. Assume they keep getting better. Assume every demo ships. The problem isn't the technology. It's the customer's math: all this spending only makes sense if somebody buys roughly $1.2 trillion of tokens a year, climbing toward $2.5 trillion by 2031.

So who's the buyer? And do they ever get their money back?

Short answer: the buyer is payroll, and no.

Two channels pay a token bill: replacing wages already being paid, or replacing wages that would otherwise have been hired. Either way, competition pushes the savings through to customers, not to the company that bought the tokens. That's what a competitive market is supposed to do. It just means the compute layer needs the buyers to keep losing the very gains that were supposed to fund the hardware.

The tab

Start with what's being spent. Goldman Sachs models AI capex[1] at $765 billion this year, rising to $1.6 trillion a year by 2031, for $7.6 trillion all in. That's a model, not a promise, but company guidance is in the same neighborhood. Microsoft[2] is guiding to about $175 billion of reported capex this year (roughly $190 billion before a lease-accounting change), Meta[3] to $130–145 billion, Amazon[4] to about $220 billion, and Alphabet[5] to $195–205 billion. That's roughly $733 billion for the four of them at the midpoints, or call it $748 billion on Microsoft's old basis. Oracle[6] runs on its own fiscal calendar, but its next-year indication works out to as much as $95 billion gross[7]. Stack them all and you're in the $830–840 billion range, with the caveat that the fiscal periods and definitions don't line up perfectly.

It's not spread evenly. Oracle just finished a fiscal year spending $55.7 billion on capex against $67.4 billion of revenue[6]. That's 83 cents of every dollar it brought in. The big four are nowhere near that, but they're all at levels that would have looked insane five years ago.

2026 Capital Expenditure Guidance, by Company
Capex is single-year guidance as disclosed by each company; fiscal periods and definitions vary, and ranges are shown at their stated midpoint. None of the five gives forward free-cash-flow guidance, so the second series shows what each actually reported most recently[2][6][8][57][58]: full fiscal-year totals for Microsoft (FY26, ended June 2026) and Oracle (FY26, ended May 2026), trailing twelve months through the most recent quarter for Meta, Amazon, and Alphabet.

Can cash flow cover this? For some, yes: Alphabet generated $53 billion of free cash flow over the past year even after $132 billion of capex[8]. For others, no: Oracle just ran a fiscal year about $24 billion free-cash-flow negative[6]. Sector-wide, the Bank of England estimates the buildout needs roughly $1.5 trillion of outside money[9] under current plans, including about $800 billion from private credit (a forward-looking estimate for the whole sector, refreshed this July[10], not a claim that every checkbook is empty). Some of it is plain debt; a lot of it is leases, capacity deals, and private-credit structures.

Now the payback math. AI chips don't earn forever. Alphabet books its servers over about six years[11], but a chip can stop earning premium rates long before the accounting says so[12], because the next generation shows up and undercuts it. And once you're buying new hardware every single year, the question stops being "when does this batch pay off" and becomes "what does the whole machine need to earn, every year." That's just annual capex divided by the margin on compute. Say the sellers keep 65 cents of each revenue dollar after the direct cost of serving it. That number is my assumption; move it and the hurdle moves. Then $765 billion of capex needs about $1.18 trillion of revenue a year just to cover the hardware. Not profit. Hardware. Before salaries, buildings, interest, or a single dollar of return for anyone. At 2031's spending rate, the number is about $2.5 trillion. Call this what it is: a stress test, not an accounting identity. But it's the right stress test, because the hardware really does have to be bought again and again.

The Hardware Alone Needs a Trillion-Dollar Customer
Annual capex versus revenue required to cover just the hardware, assuming sellers keep 65¢ of every compute revenue dollar after direct serving costs (author's assumption; move it and the hurdle moves).

One caveat before someone raises it: not all of this comes back as literal token sales. Meta's data centers pay off through better ads[3]. Google's through search[8]. Fine, but the return still has to land on someone's income statement, as new revenue or lower cost, at the same scale. Moving the meter around doesn't shrink the bill.

First, untangle the demand

A lot of today's "demand" is the supply side buying from itself. Microsoft's money flows to OpenAI, and OpenAI's compute runs on Microsoft's cloud[13], and the FTC found these partnerships came with requirements to spend big chunks of the investment right back on the partner's cloud[14]. Amazon has put $8 billion into Anthropic[15], Google another $2.55 billion[16], and Anthropic buys enormous amounts of compute from both[17]. Nvidia owns a piece of CoreWeave[18], added $2 billion more[19], and commits to buy up to $6.3 billion of CoreWeave's unsold capacity through 2032[20]. Real money moves, real revenue gets booked, and some of it may even reflect real demand. But entangled revenue can't prove the thing this essay needs proven: that an outside customer, spending only its own money, will pay.

Three Closed Loops Behind “Demand”
Each investor's cash comes back to it as a customer payment for cloud or compute: the same dollars, changing hats. This doesn't make the money fake; it just means it can't prove an outside buyer exists.
invests pays for Azure compute Microsoft OpenAI $8B invests pays AWS $2.55B invests pays Google Cloud Amazon Google Anthropic equity stake + $2B buys $6.3B capacity thru '32 Nvidia CoreWeave Investment Compute / cloud purchase
Microsoft, Amazon, Google, and Nvidia each invest in an AI lab or infrastructure company that turns around and spends a large share of that money buying compute or cloud capacity back from the same investor.[13][14][15][16][17][18][19][20]

We've seen this movie. In the late-90s telecom bubble, Lucent extended about $8 billion of financing to its own customers, and Nortel about $3 billion[12]. Vendor financing made demand look structural when a lot of it was the sellers funding their own order books[21]. And keep the ending of that story in mind: much of the fiber was real and eventually useful. It got lit, it carried the internet, consumers won huge. The overbuild destroyed enormous amounts of investor capital anyway. Useful infrastructure and destroyed capital are not opposites.

That doesn't make the labs' revenue fake. It means you have to count it carefully: money from actual outside companies and actual consumers counts. Money from your own investors, partners, and suppliers doesn't.

Consumers are real, just small. AI apps pulled in over $4 billion of in-app purchases in the first half of 2026[22], a number that leaves out web subscriptions and API deals, and is growing fast. It's still $4 billion per half-year against a hurdle of a trillion per year.

The only pool big enough

So where does a trillion a year come from? A company can pay a token bill, directly, in two main ways: replace work it already pays for, or build something new that earns more than the tokens cost. (There are softer justifications like quality, compliance, and keeping the board calm, but those live in the insurance bucket, and we'll get there.) Follow each one.

The first is simple math: fire people, hire tokens. And the pool is enormous. US employee compensation runs about $15 trillion a year[23]. Worldwide, labor's share of GDP is a bit over half[24], so against a world economy of about $118 trillion[25], call it $60 to 62 trillion. For scale, estimates of the entire global software market run from about $700 billion[26] to $1.4 trillion a year[27], depending on what you count. And you can already see the first drops moving: one study, "Payrolls to Prompts"[28], found the companies most exposed to AI increasing their spending with model providers while cutting spending on hiring marketplaces, relative to everyone else. Tokens in, contractors out.

Pool of money Annual size
Today's global software market$0.7–1.4 trillion[26][27]
Token revenue needed to cover hardware today~$1.2 trillion
US employee compensation~$15 trillion[23]
Token revenue needed to cover hardware by 2031~$2.5 trillion
Global labor income (labor's share of world GDP)~$60–62 trillion[24][25]

The second way is more interesting. Say Uber spends a dollar on tokens and stands up a new bike-courier business that brings in a dollar fifty. Great trade, until Lyft copies it, because Lyft can buy the same tokens. Now they're competing on price, the margin gets squeezed toward zero, and token prices are falling too, which just gives everyone more room to cut. The savings cascade downhill, layer by layer, until they reach the first company with a moat. And if nobody has one, they land in the customer's pocket as a service that's worth more than it costs.

But notice what didn't stop: the token bill still gets paid. Even at zero margin, the bike business pays its token supplier every month; the customers fund it, the way passengers fund jet fuel. So the second channel doesn't rescue the buyer. It just changes who writes the check.

And look at where that check really comes from. Who would have run dispatch, support, and routing for that bike business in 2019? People. The new-revenue channel doesn't escape the wage bill. It raids the payroll of jobs that never get posted. Either way (and this is my claim, the spine of the piece, not a measured fact), the price a token can charge is anchored to the cost of the human who would have done the task. Wages paid, or wages avoided.

One honest exception: work with no human alternative at all, like an agent grinding through a million-step problem overnight, something no team could do at any price. There, the anchor is what someone will pay, not what a person costs. That category is real. Nobody has measured it; my judgment is that today it's small.

The bigger pushback: AI won't just do old work cheaper, it'll grow the pie with new drugs, new products, and new revenue, and the token bill gets paid out of the growth. Three problems. First, the bulls don't get to count growth that was already coming: at roughly 3% real growth[29], the world economy adds output equivalent to about $3 to 4 trillion in today's dollars in a normal year, with no help from AI. Second, scale: paying the 2031 bill out of genuinely new growth means adding an annual flow roughly the size of Italy's entire economy[30], above trend, and keeping it there. Third, timing: the chips need paying back inside a three-to-six-year window, and the new-revenue stories run on longer clocks, since a new drug can still take roughly a decade to develop and approve[31]. The growth is possible. The depreciation is scheduled.

Who buys the most, and why they can't keep the savings

Here's my model of who buys. Two things decide how many tokens an industry uses: how big a share of the product's cost AI can become, and how hard competitors force you to hand the savings along.

The first is already visible in software itself. BCG's work on AI-native software[32] puts these companies at 50 to 60% gross margins out of the gate, against the 80 to 90% classic SaaS enjoys. It's a consulting framework built from client work, not a census, but the direction is hard to argue with. Running the models is a real cost of the product now.

Here's the uncomfortable part. The industries that buy the most tokens are the ones least able to keep the gains. Heavy AI cost-share means big savings; fierce competition means those savings go straight into lower prices. The best customers for tokens are the worst positioned to profit from them.

It runs the other way too. A company with real pricing power (think Visa, or anyone with a lock on its market) can keep whatever AI saves it. But nobody's forcing that company to spend, so it buys relatively few tokens. The firms that could keep the gains don't have to buy. The firms that have to buy can't keep the gains.

Why the buyer never gets paid

This is the crux, so slow down here.

There's a difference between AI making an industry more productive and AI making a company more money. When everyone gets the same tool, competition decides who keeps the gains. The usual winner is the customer: lower prices, faster service, more stuff.

This isn't a hunch. The economist William Nordhaus studied a half-century of American innovation[33] and found the companies doing the innovating kept about 2.2% of the value they created; the rest flowed through to everyone else. It's a macro estimate, not a law for every firm. But it's what new technology normally does, and AI would have to be the exception. Nobody has explained why it would be.

Jamie Dimon said the quiet part out loud on JPMorgan's Q1 2026 earnings call[34]: everyone is going to adopt this, and the benefits "will be passed on to the marketplace." That's the CEO of America's biggest bank telling you, as I read him, that his AI gains are temporary.

The standard comeback: "we won't cut people, we'll ship five times as much." Sure. Your competitor ships five times as much too. Picture Facebook cloning whatever works at TikTok (not proof, but you know the dynamic), or just replay the bike-courier race from two sections ago. Everybody runs faster; nobody moves up. A Red Queen race.

To be exact about the claim: defensive spending isn't worthless. It keeps you from losing revenue to a rival, and that's worth something. What it doesn't do is earn anything extra. Same revenue as before, new cost line under it.

What the numbers say so far

The honest summary of the evidence: mixed.

My read of the pattern, and it is a read, not a measurement: companies are buying insurance, not making investments. They pay so they don't fall behind, not because they can point to the return. And insurance budgets get cut the moment the fear fades.

The price problem

Careful here, because there are two different prices in this story, and right now they're moving in opposite directions.

The price of capability is collapsing. One research team tracked tokens at fixed quality falling 600-fold[42], with budget-tier prices halving about every 13 months. The cost of hitting a given benchmark score has been falling 5 to 10x a year[43], with the usual caveats about benchmarks[44]. Gartner predicts the cost of running trillion-parameter models drops more than 90% by 2030[45], and Nvidia itself promises up to 10x cheaper tokens[46] going from its Blackwell chips to Rubin (a vendor claim, but a telling one).

600× Fall in price for fixed-quality tokens tracked by one research team[42]
5–10× / yr Annual decline in cost to hit a given benchmark score[43]
>90% Gartner-predicted drop in trillion-parameter inference cost by 2030[45]

The price of raw compute is rising. SemiAnalysis tracked GPU rental prices up roughly 40% between last October and this spring[47]. Google signed a deal with SpaceX worth about $920 million a month[48] for a bundle built around 110,000 Nvidia GPUs plus the networking, memory, and services wrapped around them. And the supply of frontier compute is growing an estimated 3x a year[49], which is fast by any normal standard and slow against what everyone is trying to do with it.

+40% GPU rental price increase, last October to this spring[47]
$920M / mo Google's compute deal with SpaceX, ~110,000 Nvidia GPUs[48]
3× / yr Estimated growth in frontier compute supply[49]

The bulls have an answer for this. Two answers, actually, and they point in opposite directions. The first is the Jevons paradox: when something useful gets cheaper, people don't spend less on it, they find new uses and spend more. It happened with electric light, it happened with bandwidth, and it's plausibly what has happened with tokens so far, though nobody has built the clean price-times-volume index that would prove it. The second is Dwarkesh Patel, who argued in July that compute gets 10 to 15x more expensive[50]: a human-level software engineer running on a single H100 would justify renting that H100 for $250,000 a year, because that's what the engineer it replaces costs. (He labels his own key assumptions speculative, and to be precise, the two stories can coexist for a while: capability prices falling while scarce frontier compute rents rise.) But as accounts of where the money ends up, they point opposite ways, and notice that neither one rescues the token buyer. In the Jevons world you consume more without earning more. In Dwarkesh's world you pay engineer prices for the engineer's replacement, and the savings round to zero. The strongest bull case out there quietly agrees with this essay about the buyers.

My bet against the Dwarkesh camp comes down to one thing: you can't photocopy an engineer, but you can copy a model. Chinese open-weight models sit near the frontier on the benchmarks Stanford tracks[51], at a fraction of the posted prices[52], and a worker you can copy gets priced at the cost of running it, not at the wage it replaces. That last step is a prediction, not a fact; moats made of data, distribution, and reliability can hold markups up for a while. But gravity is gravity.

Either way, what decides whether the data centers get paid is neither price on its own. It's the spread: what a GPU's output sells for, minus what it costs to produce, and whether a new chip generation guts the rates and workloads of the old one before it's paid off. Price minus cost, batch by batch. That number works in both camps' worlds, and it's the one to watch.

The Spread That Decides Who Gets Paid
Illustrative framework, not measured data: as each new GPU generation launches, it drags down what the previous generation's output can still sell for, while build cost per generation keeps climbing. The gap between the two lines, not either line alone, is what pays back the buildout.

How it ends

By this point, token spending comes in two shapes, and they don't behave the same.

Some of it is cost of goods, meaning tokens baked into a product that sells, like the bike business. That spending is durable. Nobody cuts the ingredients of a product that's selling: the airline industry has famously failed to earn its cost of capital across its history[53] (it still doesn't[54]), and the fuel bill got paid through all of it. The rest is insurance: defensive spending with no product attached, bought so nobody can say you fell behind. That's the fragile part.

So, two endings.

One: the race just continues. The cost-of-goods spending compounds, fear keeps the insurance renewed, and the compute layer gets paid out of the adopters' hides.

Two: enough CFOs ask "where's the return?" in the same quarter, and the insurance gets cancelled. What synchronizes them? The usual things: a rate spike, an earnings recession, one bad quarter that makes every board ask why the AI line doubled. Insurance is an expense you cut the moment money gets tight, and this is the most expensive insurance ever sold. The cost-of-goods spending survives that moment; the insurance doesn't. So the size of the crash depends on which shape dominates when the trigger comes, and my read of the receipts above is that most of today's spending is still insurance-shaped. If that's right when the trigger comes, a large piece of that $7.6 trillion never gets earned back by anyone.

Notice what both endings share: the companies buying the tokens don't come out ahead. In a competitive market, lasting extra returns for adopters trend toward zero. The only live question is how much the compute layer keeps before its own prices get competed down too.

Yes, there are exceptions. A company with data nobody else has, locked-in distribution, or a regulator keeping rivals out can hold onto its gains. The question was never whether exceptions exist. It's whether they're anywhere close to big enough to cover a $7.6 trillion tab.

The test is already on the calendar

None of this has to stay an argument. It's checkable, and the first check has a date.

Alphabet booked $6.482 billion of depreciation in Q1 2026[56], on about $110 billion of quarterly revenue. If the buildout pushes depreciation to $12 billion by Q1 2027 (my assumption, not their guidance), that's roughly $5.5 billion of new cost in a single quarter. At the same 65-cent margin, covering it takes about $8.5 billion of new revenue: 7.7 points of growth consumed before operating income moves at all. Hold everything else equal, which reality won't, and growth that would look great in any normal year reads as running in place. Watch how that call goes.

Depreciation drag model Q1 2026 (actual) Q1 2027 (author's hypothetical)
Quarterly depreciation$6.482B[56]$12B (assumption)
Quarterly revenue~$110Bn/a
Incremental depreciation vs. Q1 2026n/a~$5.5B
New revenue needed to cover it at a 65¢ marginn/a~$8.5B
Growth consumed before operating income movesn/a~7.7 points

Longer term, five signals would tell you this thesis is breaking:

  1. Adopters' margins expand and management credits AI, on the record, with numbers.
  2. Token revenue from real outside customers grows at stable or rising prices.
  3. Old GPUs stay busy at yesterday's rates after new chips launch.
  4. AI-built apps show retention, paying users, and normal software margins, not just download counts.
  5. Labor cost per unit falls without prices falling or defensive spending rising to match.

None of these is courtroom proof, and each has an innocent explanation on its own. Together, they're the shape my being wrong would take. If they show up, the tab gets paid. Until then, here's the answer to the title: the tokens get bought with wages, paid or avoided, and in a competitive market, most of what that money buys ends up with your customers, not with you.

Sources & References

  1. Goldman Sachs, "Tracking Trillions: The Assumptions Shaping the Scale of the AI Build-Out," Goldman Sachs Insights, 2026. Link
  2. Microsoft, FY2026 Q4 earnings materials and call, Microsoft Investor Relations, 2026. Link
  3. Meta Platforms, "Meta Reports Second Quarter 2026 Results," Meta Investor Relations, 2026. Link
  4. Associated Press, coverage of Amazon's capital-expenditure guidance, AP News, 2026. Link
  5. Investing.com, "Google Quarterly Cloud Revenue Growth Beats Expectations" (Alphabet capex guidance), 2026. Link
  6. Oracle Corporation, "Oracle Announces Record Q4 and FY2026 Results Driven by Cloud Infrastructure & Cloud Applications," Oracle Investor Relations, 2026. Link
  7. Benzinga, "Full Transcript: Oracle Q4 2026 Earnings Call," 2026. Link
  8. Alphabet Inc., Q2 2026 Earnings Release, 2026. Link
  9. Bank of England, Financial Stability Report, December 2025. Link
  10. Bank of England, Financial Stability Report, July 2026. Link
  11. Alphabet Inc., Investor Relations FAQ (server depreciation schedule). Link
  12. Princeton CITP Blog, "The Lifespan of AI Chips: The $300 Billion Question," October 15, 2025. Link
  13. Microsoft, "The Next Phase of the Microsoft-OpenAI Partnership," Official Microsoft Blog, April 27, 2026. Link
  14. Federal Trade Commission, "FTC Issues Staff Report on AI Partnerships & Investments Study," press release, January 2025. Link
  15. Anthropic, "Anthropic and Amazon Expand Partnership with Trainium," Anthropic News. Link
  16. Federal Trade Commission, 6(b) staff report on AI partnerships and investments (redacted). Link
  17. Anthropic, "Expanding Our Use of Google Cloud TPUs and Services," Anthropic News. Link
  18. Nvidia Corporation, Schedule 13G filing disclosing stake in CoreWeave, SEC EDGAR. Link
  19. Nvidia Corporation, press-release exhibit on additional CoreWeave investment, SEC EDGAR. Link
  20. Investing.com, "CoreWeave, Nvidia Sign New $6.3 Billion Deal for Cloud Computing Capacity," 2026. Link
  21. MIT Center for Transportation & Logistics, "Telecom, Cisco and Lucent" (vendor-financing case study). Link
  22. Sensor Tower, "State of AI 2026 Report: Global Time Spent on Generative AI Apps Projected to More Than Double Year-Over-Year." Link
  23. U.S. Bureau of Economic Analysis, "National Economic Accounts Annual Update," Survey of Current Business, November 2025. Link
  24. UN Statistics Division, "Extended Report 2025: SDG Goal 10." Link
  25. Federal Reserve Bank of St. Louis, FRED series "World GDP" (NYGDPMKTPCDWLD). Link
  26. World Intellectual Property Organization, "Global Software Spending," Global Innovation Index blog, 2025. Link
  27. Gartner, "Gartner Forecasts Worldwide IT Spending to Grow 13.5% in 2026, Totaling $6.31 Trillion," April 22, 2026. Link
  28. "Payrolls to Prompts," arXiv preprint, 2026. Link
  29. International Monetary Fund, World Economic Outlook, April 2026. Link
  30. Federal Reserve Bank of St. Louis, FRED series "Italy GDP" (MKTGDPITA646NWDB). Link
  31. U.S. Food and Drug Administration, "Drug Development and Review Definitions." Link
  32. Boston Consulting Group, "Executive Perspectives: AI and the Future of Software," 2026. Link
  33. William D. Nordhaus, "Schumpeterian Profits in the American Economy: Theory and Measurement," NBER Working Paper No. 10433, 2004. Link
  34. JPMorgan Chase & Co., Q1 2026 earnings call transcript. Link
  35. Los Angeles Times, "Uber Caps Staff Use of AI Coding Tools After Blowing Its Budget," June 2, 2026. Link
  36. TechCrunch, "Uber Caps Employee AI Spending After Blowing Through Budget in Four Months," June 2, 2026. Link
  37. Morgan Stanley, "Economic Signals Emerging from Founders Institute 2026," Morgan Stanley Insights. Link
  38. The Corner, "25% of S&P Companies Already Reporting Measurable Benefits from AI." Link
  39. NBER Working Paper No. 34836, executive survey on AI adoption and productivity. Link
  40. Project NANDA (MIT), "The State of AI in Business 2025" (archived). Link
  41. 9to5Mac, "App Store Sees 84% Surge in New Apps as AI Coding Tools Take Off," April 6, 2026. Link
  42. arXiv preprint tracking inference price-per-quality trends. Link
  43. arXiv preprint tracking benchmark-score cost trends. Link
  44. Epoch AI, "How Persistent Is the Inference Cost Burden?" Gradient Updates. Link
  45. Gartner, "Gartner Predicts That by 2030, Performing Inference on an LLM with 1 Trillion Parameters Will Cost GenAI Providers Over 90% Less Than in 2025," March 25, 2026. Link
  46. Nvidia Corporation, "Rubin Platform: AI Supercomputer," Nvidia Newsroom. Link
  47. SemiAnalysis, "The Great GPU Shortage: Rental Capacity," newsletter. Link
  48. Alphabet Inc., agreement with SpaceX, Form FWP exhibit, SEC EDGAR. Link
  49. Epoch AI, "AI Chip Production," Data Insights. Link
  50. Dwarkesh Patel, "Why Compute Might Get 10x More Expensive," July 2026. Link
  51. Stanford HAI, "2026 AI Index Report: Technical Performance." Link
  52. J.P. Morgan Private Bank, "AI Use Is Exploding, So Are the Bills." Link
  53. McKinsey & Company, "The Six Secrets of Profitable Airlines." Link
  54. International Air Transport Association, press release, December 9, 2025. Link
  55. JPMorgan Chase & Co., 2026 Company Update, full event transcript. Link
  56. Alphabet Inc., Form 10-Q for the quarter ended March 31, 2026, SEC EDGAR. Link
  57. GuruFocus, "Meta Platforms Free Cash Flow," trailing-twelve-month data, accessed August 2026. Link
  58. Amazon.com, Inc., "Amazon.com Announces Second Quarter Results," Q2 2026 earnings release, SEC EDGAR. Link