🟪 Friday Charts

The cost of compute

"The data center is the new unit of computing."
— Jensen Huang

Friday charts: The cost of compute

Graphics processing units (GPUs) were first used to develop AI in 2012, when Alex Krizhevsky bought two Nvidia GTX 580 graphics cards to train AlexNet, a pioneering computer vision model.

The GTX 580 was built for high-end PC games like Skyrim and Portal 2 — games whose graphics were rendered with the same kind of parallel matrix math Krizhevsky used to train his neural networks. They cost $499 each.

The first GPU designed for AI was the P100, released in 2016. They cost about $6,000 each.

Nvidia bundled eight of these, along with CPUs, memory, and software, into an “AI supercomputer” called the DGX-1. That cost about $129,000.

(Jensen Huang personally delivered the first one to Elon Musk at OpenAI.)

Supercomputers have been getting bigger and more expensive ever since.

This week, customers began running Nvidia's Vera Rubin Pod in production — a system Jensen Huang describes as “five purpose-built racks operating as one massive AI supercomputer for agentic workloads.”

As shown above, miles of fibre optic cabling, sheathed in yellow, transform those racks into a single, super-fast, super-powerful computer. There’s miles of copper wiring in the back, too. 

From left to right, the five cabinets house GPUs, LPUs (a new addition, to speed things up), CPUs, memory, and networking equipment. The yellow-sheathed cables tie everything together into a single system.

Unlike a GTX 580, you can’t order one of these from the internet, so we don’t know exactly what they cost. But just the cabinet for GPUs is estimated to cost $9 million. A pod of five cabinets (or racks) likely costs tens of millions of dollars.

To build a new data center, you might need 1,000 of these, all networked together.

To train a frontier language model, you might need a network of such data centers. Perhaps ten of them, combining the computing power of nearly a million GPUs into a single supercomputer.

It’s a long way from the two off-the-shelf gaming GPUs Alex Krizhevsky used to train AlexNet for $1,000. 

It’s also a lot more expensive. The network of data centers OpenAI said this week it will build in Georgia is expected to cost $30 billion.

For investors, it’s getting to be a bit much.

Alphabet shares fell 7% this week after reporting negative free cash flow of $5.9 billion for the second quarter of the year. It’s the first time since going public in 2004 that the company spent more money than it earned.

There’s nothing inherently wrong with being free-cash-flow negative. Often, it’s encouraged. We spent a decade complaining about big tech hoarding cash and buying back their shares instead of investing.

Now, they’re investing.

But we’re not sure how we feel about that, so we’ve de-rated their stocks. Alphabet reported quarterly earnings up 30%, but its shares are just unchanged on the year. Roughly speaking, that means the stock has gotten 30% cheaper.

The de-rating is even more dramatic at Microsoft and Oracle, where earnings are significantly higher and the shares are significantly lower (by 19% and 40%, respectively).

The problem is simply that the cost of generating AI compute is spiraling higher — from $1,000 in 2012 to billions of dollars now — and we still don’t know exactly how useful it will be.

Technically, Vera Rubin pods are yet another extraordinary accomplishment from Nvidia. But is buying 1,000 of them for a data center a good investment? 

The market is starting to have its doubts.

Let’s check the charts.

Trend change:

One of these is not like the others. Shareholders may need some time to adjust.

The investment that needs a return:

By the end of the year, hyperscalers and the neoclouds will have spent $2 trillion on capex.

Agentic AI requires a lot of memory.

The price of DRAM is up 10x in a year. The price of Micron shares is up less than 2x, because surely this can’t go on?

The picks and shovels:

Makers of semiconductors are the big winners so far. A report from Exponential View cites an estimate of $1.5 trillion for 2026, more than double 2025.

Model revenue:

Callum Williams estimates revenue earned from LLMs at a $120 billion annualized run rate, with Anthropic taking by far the biggest share. That number will have to grow rapidly for everyone to make a return on their investments.

Growing rapidly:

Revenue growth at the three cloud providers is accelerating, at Google, especially.

The machines keep getting better:

Per Exponential View, one gigawatt of data center capacity now produces nearly 500 trillion tokens. That’s up from a number that rounded to zero as recently as 2023.

Demand keeps growing:

All those data centers now produce roughly 35 quadrillion tokens per month. I can’t imagine how they measure that number, but it’s growing at the exponential rate of 14x a year.

Compute is on back order.

The hyperscalers have an unprecedented $2 trillion backlog of orders. Whether the customers ordering all this compute will be able to pay for it when it arrives remains to be seen.

Another big number:

Nikkei research estimates the hyperscalers now have $1.65 trillion of debt that does not show up on their balance sheets. It’s not a secret. The numbers above are taken from the footnotes of the hyperscalers’ financial statements. But who reads the footnotes?

Future earnings will have to exceed these debts — and more. Because the stuff it buys from Nvidia isn’t getting any cheaper.

Have a great weekend, cash-flow-positive readers.

— Byron Gilliam