# The Entire AI Data Center Explained — From Electricity to ChatGPT

- Source: https://www.youtube.com/watch?v=ckoi0RTEgcY (YouTube)
- Creator: Leo Cui, Ph.D., CFA
- Published: 2026-07-23T17:17:04.000Z
- Transcribed by Memora: 2026-08-07T05:01:51.009Z
- Canonical page: https://media-pilot-nine.vercel.app/youtube/90

> Transcript and summary produced by Memora from the publicly available
> video linked above. The original video belongs to its creator.

## Summary

This video presents a comprehensive breakdown of the AI data center, framing it as a factory that converts electricity into tokens. It begins by contrasting traditional search with generative AI's compute-heavy inference, explaining key concepts like tokens, flops, training, and inference. The scaling law—where larger models predictably improve with more compute and data—turned AI into a capital expenditure race, driving massive infrastructure investment. The narrative then traces a user query's journey through networking, tokenization, prefill, and decode, introducing the data center's ten-layer stack. The first layer, power, faces explosive demand: Nvidia racks are projected to reach 600 kW, overwhelming grids and leading hyperscalers to behind-the-meter solutions like nuclear, gas turbines, and fuel cells. Cooling follows, with air cooling failing above 30–50 kW per rack, forcing a shift to liquid cooling (direct-to-chip). The server layer examines Nvidia's dominance through GPUs, high-bandwidth memory, and its CUDA software moat, while competitors like AMD and Broadcom (custom chips) carve niche positions. Server integrators struggle with slim margins, while networking—NVLink for scale-up, Ethernet defeating InfiniBand for scale-out—and optical components become high-margin essentials. Memory and storage are covered: HBM is sold out years ahead for inference's bandwidth needs, and hard drives see a surprise revival for training checkpoints. The software stack layer highlights Linux, Kubernetes, CUDA lock-in, and inference-serving engines that slash costs 3–10x via batching, caching, and quantization, making API pricing viable. Finally, the video traces the money flow: hyperscalers, neoclouds, and AI labs spend $725B annually, with Nvidia’s own investments recycling back, but end-user payments remain the smallest piece, raising sustainability questions. The closing frames the AI data center buildout as potentially the largest infrastructure project in history, yet its immediate purpose can be mundane, and its ultimate value remains deeply uncertain.

## Key points

- [object Object]
- [object Object]
- [object Object]
- [object Object]
- [object Object]

## Chapters

- [0:00](https://media-pilot-nine.vercel.app/youtube/90?t=0) Why AI?
- [6:44](https://media-pilot-nine.vercel.app/youtube/90?t=404) The Build
- [21:07](https://media-pilot-nine.vercel.app/youtube/90?t=1267) Hardware
- [33:39](https://media-pilot-nine.vercel.app/youtube/90?t=2019) Economics

## Transcript

**[0:00]** Last night, sometime around

**[0:01]** 7 PM, You pulled out your

**[0:03]** phone, you typed a question.

**[0:04]** Maybe it was, "What should I make

**[0:06]** for dinner with chicken and rice?"

**[0:08]** And about two seconds later,

**[0:09]** A machine wrote you an answer.

**[0:11]** Two seconds.

**[0:12]** That's what I want to do in this video.

**[0:14]** I want to slow those two

**[0:15]** seconds down, way down.

**[0:16]** Because in those two seconds, your

**[0:18]** questions left your phone, traveled

**[0:20]** hundreds of miles through strands of

**[0:22]** glass thinner than a human hair, and

**[0:24]** arrived at a building the size of several

**[0:25]** football fields, a building that drinks

**[0:27]** as much electricity as a small city.

**[0:29]** Inside that building, your question

**[0:31]** passed through a machine that costs

**[0:33]** as much as a house, got translated

**[0:35]** into pure math, was processed by

**[0:37]** chips running so hot they have to be

**[0:39]** liquid cooled like a race car engine.

**[0:41]** And then the answer came back to

**[0:43]** you letter by letter before you

**[0:44]** had time to lower your thumb.

**[0:46]** and here's the part that should

**[0:47]** get your attention as an investor.

**[0:49]** To make those two seconds possible,

**[0:51]** the largest companies on Earth are

**[0:53]** spending this year alone roughly

**[0:55]** seven hundred and twenty-five

**[0:56]** billion dollars on infrastructure.

**[0:58]** That's just four companies, Amazon,

**[1:00]** Microsoft, Google, and Meta.

**[1:02]** That's more in one year than the

**[1:04]** inflation adjusted cost of the

**[1:05]** entire US interstate highway system.

**[1:08]** Goldman Sachs projects a total build out at

**[1:10]** seven point six trillion dollars between

**[1:12]** twenty twenty-six and twenty thirty-one.

**[1:14]** Jensen Huang, the CEO of Nvidia,

**[1:16]** stood on stage at Davos this

**[1:18]** January and called it his own

**[1:20]** words, "The largest infrastructure

**[1:22]** build out in human history."

**[1:23]** So the question this whole

**[1:24]** video hangs on is simple: where

**[1:26]** does all the money actually go?

**[1:28]** By the end of this video, you're

**[1:29]** going to be able to answer that.

**[1:31]** you'll understand every single step

**[1:32]** your question takes, the power plants,

**[1:34]** the cooling systems, the chips, the

**[1:36]** memory, the fiber optics, and software.

**[1:39]** You'll know which companies sit at every

**[1:40]** step, which companies are printing money,

**[1:43]** which companies are telling stories,

**[1:44]** and where the whole thing could crack

**[1:46]** I'm Leo, a VC investor.

**[1:48]** This is educational purposes,

**[1:50]** not financial advice

**[1:51]** So before we trace your question

**[1:52]** across the country, we have to

**[1:54]** answer something more basic.

**[1:55]** The internet has existed for thirty years.

**[1:58]** Google has answered

**[1:59]** trillions of questions.

**[2:00]** Why did nobody need to spend

**[2:01]** three-quarters of a trillion

**[2:02]** dollars a year until now?

**[2:04]** What changed?

**[2:05]** The answer comes down to a

**[2:06]** difference between two people,

**[2:08]** A librarian and a writer.

**[2:10]** Google is a librarian.

**[2:11]** When you search chicken rice recipe,

**[2:14]** Google doesn't cook anything.

**[2:15]** It walks into a giant library

**[2:17]** it has already organized.

**[2:18]** It indexed the whole internet years

**[2:20]** ago and keeps updating it, and it

**[2:22]** hands you pages that already exist.

**[2:24]** The expensive work happened in advance.

**[2:26]** Answering you is just a lookup.

**[2:28]** Fast, cheap, done.

**[2:30]** A Google search costs

**[2:31]** a fraction of a cent.

**[2:32]** ChatGPT is a writer.

**[2:33]** When you ask it the same question,

**[2:35]** there's no answer sitting on a shelf.

**[2:37]** There's no database entry that says,

**[2:39]** "Here's what to tell this person." The

**[2:40]** model composes your answer from scratch,

**[2:43]** one word at a time, every single time.

**[2:45]** Even if a million people ask the same

**[2:47]** question today, it doesn't retrieve, it

**[2:49]** generates. And generation is expensive.

**[2:52]** A single ChatGPT query can cost 10 to 100

**[2:55]** times more compute than a Google search.

**[2:57]** Now multiply that by 900

**[2:59]** million weekly users.

**[3:00]** That's the entire reason

**[3:01]** this video exists.

**[3:03]** Search retrieves, AI generates, and

**[3:05]** generation is a manufacturing process

**[3:07]** Which brings me to the analogy I'm going

**[3:09]** to use for the rest of this video.

**[3:11]** I want you to think of an AI data center

**[3:13]** as a factory, a very strange factory.

**[3:15]** Raw material goes in

**[3:16]** one side, electricity.

**[3:18]** A product comes out the other end, words.

**[3:20]** And like any factory, it has departments,

**[3:23]** a power plant, a cooling system,

**[3:25]** assembly lines, a shipping department.

**[3:27]** We are going to tour each one.

**[3:29]** But first, three terms you need.

**[3:31]** Term one, the token.

**[3:32]** A token is the product this factory makes.

**[3:35]** Language models don't actually

**[3:36]** read words, they read tokens,

**[3:38]** which are chunks of text, roughly

**[3:40]** three-quarters of a word each.

**[3:42]** Chicken and rice is about four tokens.

**[3:44]** Your question gets chopped

**[3:45]** into tokens on the way in, and

**[3:47]** the answer gets manufactured

**[3:48]** token by token on the way out.

**[3:50]** And here's why investors care.

**[3:52]** Tokens are the unit of

**[3:53]** revenue in the AI economy.

**[3:55]** OpenAI and Anthropic literally price

**[3:57]** their product per million tokens.

**[3:59]** When you hear token, think widget

**[4:01]** comes out the assembly line

**[4:03]** Term two, the flop.

**[4:04]** A flop is one floating point operation.

**[4:07]** one single arithmetic calculation.

**[4:09]** One multiply or one add.

**[4:11]** It's a unit of labor in this factory.

**[4:13]** Manufacturing a single token requires a

**[4:15]** model to do hundreds of billions of these

**[4:17]** calculations, not per answer, per word.

**[4:20]** When people say a chip does a thousand

**[4:22]** trillion flops per second, they

**[4:23]** are telling you how many workers

**[4:25]** that chip has on the factory floor

**[4:27]** Term three, and this is a big

**[4:29]** one, training versus inference.

**[4:31]** Training is building the factory.

**[4:33]** You take a model.

**[4:34]** Think of it as a machine with

**[4:35]** over one trillion adjusting

**[4:36]** knobs called parameters.

**[4:38]** And you show it a huge portion of the

**[4:40]** reading internet, adjusting these knobs

**[4:42]** until it gets good at predicting language.

**[4:44]** This takes months.

**[4:45]** Tens of thousands of chips

**[4:46]** running around the clock, and on

**[4:48]** the order of a hundred million

**[4:49]** dollars or more per frontier model.

**[4:51]** It happens once per model.

**[4:53]** Inference is running the factory.

**[4:55]** Every time you ask ChatGPT

**[4:56]** anything, that's inference.

**[4:57]** The trained model

**[4:58]** manufacturing an answer for you.

**[5:00]** And here's the misconception I

**[5:01]** most want to kill in this video.

**[5:03]** People assume training

**[5:04]** is where the money goes.

**[5:06]** Wrong.

**[5:06]** By twenty twenty-six, roughly two

**[5:08]** thirds of all AI compute is inference.

**[5:11]** Because training happens once, but

**[5:12]** inference happens billions of times a day.

**[5:15]** OpenAI's inference bill alone is projected

**[5:17]** around fourteen billion dollars this year.

**[5:19]** The factory was expensive to build.

**[5:21]** It's even more expensive to run.

**[5:23]** Okay, so why did all of this

**[5:25]** suddenly explode after 2022?

**[5:27]** One discovery.

**[5:28]** The most economically important discovery

**[5:30]** of this decade, and most people have never

**[5:32]** heard of it: scaling law. Around 2020,

**[5:35]** researchers at OpenAI found something

**[5:37]** almost embarrassing in its simplicity.

**[5:39]** If you make the model bigger, give

**[5:41]** it more data, and spend more compute,

**[5:43]** altogether, the model gets smarter.

**[5:45]** Not sometimes. Predictably, on a

**[5:47]** chart, it's nearly a straight line.

**[5:49]** If you spend ten times more,

**[5:51]** you get a reliably better model.

**[5:53]** And stop and think about

**[5:54]** what that means for a CEO.

**[5:55]** For fifty years, better software

**[5:57]** meant hiring smarter programmers.

**[5:59]** Scaling law turned intelligence

**[6:00]** into something you could purchase.

**[6:02]** It converted AI from a research problem

**[6:04]** into a capital expenditure problem.

**[6:06]** And big companies know exactly how

**[6:08]** to compete on capital expenditures.

**[6:10]** Outspend everyone.

**[6:11]** That is the moment software

**[6:12]** stopped being about code and

**[6:14]** started being about concrete.

**[6:16]** That's why we suddenly need factories.

**[6:18]** Because here's the closing

**[6:19]** thought for this act.

**[6:20]** For the entire history of

**[6:21]** Silicon Valley, software was an

**[6:23]** escape from the physical world.

**[6:24]** zero marginal cost, infinite

**[6:26]** copies, no factory needed.

**[6:28]** AI reversed that.

**[6:29]** The frontier of software is now

**[6:31]** poured in concrete, measured in

**[6:33]** megawatts, and cooled with water.

**[6:35]** Every additional smart answer

**[6:36]** requires physical machines, physical

**[6:38]** electricity, physical heat removed.

**[6:41]** Software became heavy industry,

**[6:42]** and that changes who makes money

**[6:44]** Let's un-freeze your question.

**[6:46]** It's 7:00 PM.

**[6:47]** You have typed, "What should I

**[6:48]** make for dinner with chicken and

**[6:49]** rice?" And your thumb hits send.

**[6:51]** Here's the actual complete

**[6:53]** no-step-skips journey

**[6:54]** Step one, the trip.

**[6:55]** Your question leaves your phone as radio

**[6:57]** waves, hitting a cell tower or your

**[6:59]** Wi-Fi router, and within a few miles

**[7:01]** becomes pulses of light inside fiber

**[7:04]** optic cable, glass strands carrying

**[7:06]** data at two-thirds the speed of light.

**[7:08]** and it got routed to the nearest

**[7:09]** entry point of the AI company's

**[7:11]** network, then travel often

**[7:13]** hundreds of miles to a data center.

**[7:15]** Total time so far, a

**[7:17]** few hundred seconds

**[7:18]** Step two, the front door.

**[7:20]** Your question arrive at API gateway.

**[7:22]** think of it as a factory receiving desk.

**[7:24]** It checks who you are, checks you

**[7:26]** are not sending a thousand requests a

**[7:27]** second, runs a safety screen, and staple

**[7:30]** together everything the model needs,

**[7:31]** the system instructions, your past

**[7:33]** conversation, and your new question.

**[7:35]** Step three, tokenization.

**[7:37]** That full text get chopped into tokens.

**[7:39]** The puzzle pieces from act one,

**[7:41]** your dinner question, plus context,

**[7:43]** maybe a few hundred tokens.

**[7:45]** This get converted into numbers, because

**[7:47]** from here on, everything is math.

**[7:49]** Step four, prefill.

**[7:51]** The model reads.

**[7:52]** Here's something almost nobody knows

**[7:54]** The model process your question

**[7:55]** in two totally different phases.

**[7:57]** the first is called prefill.

**[7:59]** The model read your entire prompt

**[8:00]** all at once in parallel and build

**[8:03]** an internal understanding of it.

**[8:04]** This is a burst of raw computation,

**[8:06]** billions of calculations,

**[8:08]** And it produced something

**[8:09]** called the KV cache.

**[8:10]** don't let the name scare you.

**[8:12]** The KV cache is simply the model's

**[8:14]** working memory of your conversation.

**[8:16]** Its notes on everything said

**[8:17]** so far, held in super fast

**[8:19]** memory right next to the chip.

**[8:20]** Ever notice ChatGPT pause for a little

**[8:22]** bit before the first word appears?

**[8:24]** That pause is prefill.

**[8:25]** The factory is reading the

**[8:27]** work order. Step five, decode.

**[8:29]** The model arrives.

**[8:30]** Now the assembly line starts.

**[8:31]** The model generate the

**[8:32]** answer one token at a time.

**[8:34]** It look at your answer plus everything

**[8:36]** it has written so far, runs the

**[8:37]** entire training neural network,

**[8:39]** hundreds of billions of calculations,

**[8:41]** and produces one word, "try."

**[8:43]** Then it does the whole thing again

**[8:44]** for the next word, A. Again, one pass.

**[8:47]** Again, every single word of

**[8:49]** every ChatGPT answer on Earth is

**[8:51]** manufactured this way, one at a time.

**[8:54]** Full network pass each time.

**[8:55]** When you watch the answer type itself onto

**[8:58]** your screen, that's not a design flourish.

**[9:00]** You are literally watching an

**[9:01]** assembly line run in real time.

**[9:03]** Each word appears the

**[9:04]** moment it's manufactured.

**[9:05]** Step six, the trip home.

**[9:07]** Each token streams back through

**[9:09]** the same fiber, and two seconds

**[9:10]** later, after you hit send, you are

**[9:13]** reading the dinner ideas. One more

**[9:14]** thing happening behind the curtain.

**[9:16]** You are not alone in there.

**[9:18]** The factory will batch everything

**[9:20]** you sent with hundreds of other

**[9:21]** people's questions on the same chip

**[9:23]** simultaneously, like a delivery

**[9:25]** driver grouping orders on one route.

**[9:27]** That batching is the difference

**[9:28]** between your question costing

**[9:30]** cents and costing dollars.

**[9:32]** Now zoom all the way out because

**[9:33]** here's the whole factory in layers.

**[9:36]** This is the map for

**[9:37]** the rest of this video.

**[9:38]** 10 layers.

**[9:39]** And here's a one-sentence

**[9:40]** version of this entire video.

**[9:42]** Electricity comes in one

**[9:43]** side, flows through silicon,

**[9:45]** becomes computation and heat.

**[9:46]** The heat gets carried away by water.

**[9:48]** The computation gets coordinated by light,

**[9:51]** and what ships out the door is words.

**[9:53]** Electrons in, tokens out.

**[9:55]** That's the factory.

**[9:56]** So let's start a tour where every

**[9:57]** factory tour starts, the power plant.

**[10:00]** Because, and this surprised me the most

**[10:02]** when I first dug into this ecosystem,

**[10:04]** the story of AI in twenty twenty-six is

**[10:06]** not mainly a story about chips anymore.

**[10:08]** It's a story about electricity.

**[10:10]** Let me give you the number that

**[10:11]** framed this whole industry for me.

**[10:13]** A traditional rack of servers, the kind

**[10:15]** that ran the internet for the last twenty

**[10:17]** years, draws about five to ten kilowatts.

**[10:19]** Think of a kilowatt as ten old-fashioned

**[10:22]** 100-watt light bulbs burning at once.

**[10:24]** Nvidia's flagship AI rack, one rack,

**[10:27]** one refrigerator-sized cabinet, draws

**[10:29]** one hundred and twenty kilowatts.

**[10:31]** And the next generation coming later this

**[10:32]** year, the Vera Rubin racks, are projected

**[10:35]** to approach six hundred kilowatts per rack.

**[10:38]** That's sixty to a hundred times jump

**[10:40]** in power density in under a decade.

**[10:42]** The electrical demand of an entire

**[10:43]** neighborhood packed into a phone booth.

**[10:45]** This is called power density, and it's

**[10:47]** the root cause of nearly everything

**[10:49]** in the next two acts. Now scale up.

**[10:51]** A large AI campus today

**[10:52]** wants a gigawatt or more.

**[10:54]** A gigawatt is a thousand megawatts,

**[10:56]** roughly the output of a full-size

**[10:58]** nuclear reactor, enough electricity

**[11:00]** for about a million homes.

**[11:01]** Individual companies are now planning

**[11:03]** multiple campuses of that size.

**[11:05]** Data centers consume about

**[11:06]** four to five percent of US

**[11:07]** electricity going into this boom.

**[11:10]** The credible projections put it at nine

**[11:12]** to seventeen percent by twenty thirty.

**[11:14]** And here's the collision.

**[11:15]** The US electrical grid was built

**[11:17]** brilliantly decades ago for demand

**[11:20]** that grew one or two percent a year.

**[11:22]** AI showed up asking for

**[11:23]** tens of gigawatts right now.

**[11:25]** The grid physically cannot say yes.

**[11:27]** Two bottleneck numbers, and they are

**[11:29]** the most important numbers in this act

**[11:31]** Number one, the interconnection queue.

**[11:33]** to plug a big new facility into

**[11:35]** the grid, you file a request and

**[11:37]** wait in line while utilities study

**[11:39]** whether the grid can handle you.

**[11:41]** That line is currently

**[11:42]** four to five years long.

**[11:43]** This April, there are about

**[11:45]** four hundred and ten gigawatts of

**[11:46]** large projects waiting to connect.

**[11:49]** Eighty-seven percent of

**[11:49]** They are data centers.

**[11:51]** That's nearly five times the entire

**[11:52]** Texas grid's peak demand waiting in line.

**[11:55]** Number two, the transformer.

**[11:57]** A large power transformer, the giant

**[11:59]** gray box that steps voltage down,

**[12:00]** used to take about a year to order.

**[12:02]** Today, two and a half to four years,

**[12:04]** with prices up nearly eighty percent.

**[12:06]** You can have your chip in six months.

**[12:08]** The gray box that powers them, twenty

**[12:10]** twenty-nine. So what do you do if you're

**[12:12]** Microsoft or Meta and every month

**[12:14]** of waiting costs you the API race?

**[12:16]** You stop waiting for the grid.

**[12:17]** You go around it.

**[12:19]** The industry calls this behind-the-meter

**[12:20]** power, generating electricity

**[12:22]** on site or next door, so you

**[12:24]** never touch the public queue.

**[12:26]** And that decision, thousands of

**[12:27]** companies making it simultaneously,

**[12:29]** is what lit a fire under an entire

**[12:31]** forgotten sector of the stock market,

**[12:33]** boring old industrial power companies.

**[12:35]** Let me introduce the players, from

**[12:37]** most dramatic to most dependable.

**[12:39]** The nuclear resurrection.

**[12:41]** In 2024, Microsoft signed a

**[12:43]** deal that would have sounded

**[12:44]** like satire a decade ago.

**[12:46]** A 20-year agreement with Constellation

**[12:48]** Energy to restart Three Mile Island.

**[12:51]** The undamaged reactor next to

**[12:52]** the one from the 1979 accident.

**[12:54]** Constellation is spending about

**[12:56]** one point six billion, backed by

**[12:58]** a one billion dollar federal loan

**[12:59]** to bridge the 835 megawatt unit

**[13:02]** back in the second half of 2027.

**[13:05]** And Microsoft will buy every

**[13:06]** megawatt it produces for twenty years.

**[13:08]** Constellation operates the

**[13:09]** largest nuclear fleet in America,

**[13:11]** about twenty-two gigawatts.

**[13:12]** And suddenly those aging

**[13:13]** reactors became some of the most

**[13:15]** valuable energy assets on Earth.

**[13:17]** Why?

**[13:18]** Because AI factories run twenty-four

**[13:19]** seven, and nuclear is the only

**[13:21]** carbon-free power source that

**[13:22]** also runs twenty-four seven.

**[13:24]** Constellation stock tells the story.

**[13:26]** It's now roughly a ninety

**[13:27]** billion dollar company.

**[13:28]** Its pair, Vistra, with a thirty-seven

**[13:30]** gigawatt fleet mixing nuclear

**[13:32]** and gas, rode the same wave.

**[13:34]** The moat here is beautiful in simplicity.

**[13:36]** You cannot build a new conventional

**[13:38]** nuclear plant in America this decade.

**[13:40]** Existing reactors are replaceable.

**[13:42]** The small modular reactor lottery tickets.

**[13:44]** You have heard the tickers.

**[13:46]** Oklo, backed by Sam Altman, with

**[13:48]** over 14 gigawatts of signed pipeline.

**[13:50]** NuScale, the only SMR design

**[13:52]** actually certified by US regulators.

**[13:54]** Here's my analytical skeptical

**[13:56]** framing, and I'll be blunt.

**[13:57]** These are pre-revenue companies whose

**[13:59]** first commercial electricity arrives

**[14:01]** around twenty thirty at the earliest.

**[14:02]** Oklo doesn't yet have final

**[14:04]** regulatory approval for its design.

**[14:06]** NuScale booked about thirty-one

**[14:07]** million in revenue against a three

**[14:09]** hundred and fifty-six million loss.

**[14:11]** Both stocks are down sixty-five

**[14:12]** percent to seventy-eight percent from

**[14:14]** their late twenty twenty five peaks.

**[14:16]** That is not a business yet.

**[14:17]** That's an option on the twenty thirties.

**[14:19]** Next, the fastest power in the West.

**[14:21]** If the grid takes four years

**[14:23]** and nuclear takes 10, what can

**[14:25]** you get in 12 to 18 months?

**[14:26]** Fuel cells.

**[14:27]** Bloom Energy makes solid oxide fuel

**[14:29]** cells, boxes that convert natural

**[14:31]** gas into electricity chemically, no

**[14:33]** combustion, and you can park them behind

**[14:35]** the meter next to a data center fast.

**[14:37]** The stock nearly quadrupled

**[14:39]** in 2025, then doubled again in

**[14:41]** the first half of this year.

**[14:42]** And Bloom announced seven point six

**[14:44]** five billion in data center contracts

**[14:46]** in a single nine-day stretch.

**[14:48]** But,

**[14:48]** this July, Hunterbrook, an

**[14:50]** investigative outlet whose affiliated

**[14:53]** funds short a stock it covers.

**[14:55]** So weigh the source accordingly.

**[14:56]** Publish a report

**[14:57]** Saying that Bloom's marketed twenty

**[14:59]** billion backlog is more than forty

**[15:01]** times its binding contract obligations

**[15:03]** versus about two X for typical peers,

**[15:06]** and that scaling to its stated ambitions

**[15:08]** would consume nearly the entire global

**[15:10]** supply of scandium, a metal China

**[15:12]** now requires export licenses for.

**[15:14]** Bloom formally rejected the

**[15:15]** claims as false and misleading

**[15:17]** When a backlog number and the

**[15:19]** SEC filing disagrees by 40X, the

**[15:21]** burden of proof is on a company

**[15:22]** Next, the arms dealer

**[15:24]** setting out through 2030.

**[15:25]** My favorite business in this act is

**[15:27]** the least glamorous, GE Vernova the

**[15:29]** power spin-off of General Electric.

**[15:31]** They make the giant gas turbines

**[15:33]** that are realistically the number

**[15:35]** one near-term power source for AI

**[15:37]** because gas is the only thing you

**[15:38]** can build at scale before 2030.

**[15:40]** GE Vernova's turbine slots are sold

**[15:42]** out through the end of this decade.

**[15:44]** Their backlog is around 163 billion.

**[15:46]** In the first quarter of 2026 alone,

**[15:49]** they booked 2.4 billion in data

**[15:50]** center electrification orders.

**[15:52]** More than all of 2025.

**[15:54]** The stock is up so much it's now

**[15:55]** a nearly $300 billion company.

**[15:57]** Their only real global rival

**[15:59]** at scale, Siemens Energy.

**[16:00]** Between the substation and the

**[16:01]** chips sits a layer of equipment

**[16:04]** most people never think about.

**[16:05]** Switchgear, busways, and interruptible

**[16:08]** power supplies, the UPS, essentially

**[16:09]** a giant battery that catches the load

**[16:11]** instantly if the grid blinks, because even

**[16:14]** a half second outage can cause a training

**[16:16]** run that's been running for a month.

**[16:18]** Three companies own this layer, and

**[16:19]** remember their names because two of

**[16:21]** them show up again in the next act.

**[16:22]** Vertiv, Schneider Electric, and Eaton.

**[16:25]** Eaton's electric backlog grew

**[16:26]** forty-eight percent year over year.

**[16:27]** Vertiv's backlog more than doubled

**[16:29]** to fifteen billion dollars, and the

**[16:31]** company joined S&P 500 in March.

**[16:33]** These are the companies selling shovels

**[16:35]** to every miner, regardless of who wins.

**[16:38]** And the last line of defense,

**[16:39]** Rows of backup generators

**[16:41]** from Caterpillar and Cummins.

**[16:43]** Diesel engines the size of school

**[16:45]** buses idling in wait for the

**[16:47]** one hour a year the grid fails.

**[16:49]** Analysts think Caterpillar's data center

**[16:50]** generator business could triple by

**[16:52]** twenty-thirty. Before we move on, the

**[16:54]** uncomfortable part, because this act has

**[16:56]** one and it's showing up in your inbox.

**[16:58]** Because data centers bid for scarce

**[17:00]** power, they bid against you.

**[17:02]** In the PJM market, the grid covering 13

**[17:04]** states from Illinois to Virginia, data

**[17:07]** center demand added over nine billion

**[17:09]** dollars to the latest capacity auction,

**[17:11]** translating to residential bills rising

**[17:13]** sixteen dollars to eighteen dollars a

**[17:14]** month in parts of Ohio and Maryland.

**[17:17]** Communities are noticing.

**[17:19]** Moratoriums are being proposed.

**[17:21]** This is becoming a genuine political

**[17:22]** risk to the build-out, and any honest

**[17:25]** map of this industry has to include it.

**[17:27]** So the factory has power.

**[17:28]** one hundred kilowatts are now

**[17:29]** flowing into a single rack of chips,

**[17:32]** which create an immediate problem.

**[17:34]** Physics 101.

**[17:34]** Every one of those watts becomes heat.

**[17:37]** The factory is running a fever.

**[17:38]** A single flagship AI chip today dissipates

**[17:41]** over one thousand watts of heat.

**[17:43]** A chip the size of a postcard

**[17:44]** puts out the heat of

**[17:45]** a full-size space heater.

**[17:47]** Now stack seventy-two of them into one

**[17:49]** rack, plus their memory and networking.

**[17:51]** You have got one hundred

**[17:52]** and twenty kilowatts of heat.

**[17:53]** The output of about eighty space

**[17:55]** heaters in a cabinet you could hang.

**[17:57]** Why?

**[17:57]** Because computation is heat.

**[17:59]** Every one of those trillions of

**[18:00]** calculations pushes electrons through

**[18:02]** microscopic wires and electrical

**[18:04]** resistance turns into warmth.

**[18:06]** The factory's raw material, electricity,

**[18:08]** doesn't get consumed making tokens.

**[18:10]** It gets converted almost

**[18:12]** entirely into heat.

**[18:13]** Cooling isn't a support

**[18:14]** function of AI data center.

**[18:16]** Cooling is half the job.

**[18:17]** For thirty years, the answer was air

**[18:19]** conditioning, genuinely just fancy AC.

**[18:22]** cold air pushed up through the floor,

**[18:24]** hot air sucked out the back, giant

**[18:26]** chillers and cooling towers on the roof.

**[18:28]** And air worked fine up to about

**[18:30]** thirty to fifty kilowatts per rack

**[18:32]** But we just passed that line permanently.

**[18:34]** Air physically cannot carry

**[18:36]** heat away fast enough from a one

**[18:37]** hundred and twenty kilowatt rack.

**[18:39]** You need hurricane force

**[18:40]** winds through the servers.

**[18:41]** So the industry is undergoing its

**[18:43]** biggest plumbing change in its history.

**[18:45]** The switch from air to liquid.

**[18:47]** Water carries heat about 3,000 times

**[18:49]** more effective than air per unit volume.

**[18:51]** The technology ladder in one breath.

**[18:53]** Rear door heat exchangers, a water-cooled

**[18:56]** radiator bolted to the back of the rack.

**[18:58]** A transitional patch.

**[18:59]** Direct to chip cooling, the 2026

**[19:01]** mainstream, a metal plate with

**[19:03]** liquid channels sits directly on top

**[19:05]** of each chip, connected by hoses to

**[19:07]** a CDU, a coolant distribution unit.

**[19:09]** Think of it as the rack's heart, pumping

**[19:11]** coolant to every chip and carrying

**[19:13]** the heat to the building's water loop.

**[19:15]** Nvidia's flagship racks don't

**[19:17]** offer this as an option.

**[19:18]** They require it.

**[19:19]** And at extreme immersion cooling,

**[19:21]** literally dunking entire servers into

**[19:24]** tanks of non-conductive fluid, like

**[19:26]** deep-frying a computer that never burns

**[19:28]** two quick vocabulary items

**[19:30]** investors will encounter.

**[19:31]** PUE, power usage effectiveness.

**[19:34]** It's a factory efficiency score.

**[19:36]** Total power in divided by power

**[19:38]** that actually reaches the computer.

**[19:39]** A perfect score is one point zero. Old

**[19:41]** data centers run about 2.0, a watt of

**[19:44]** cooling for every watt of computing.

**[19:45]** Modern liquid cooled facility hits 1.1.

**[19:49]** That efficiency gap times a gigawatt

**[19:51]** times electricity prices is real money.

**[19:54]** And water.

**[19:54]** Many data centers cool by evaporating

**[19:56]** millions of gallons, which is

**[19:58]** becoming a genuine permitting and

**[20:00]** political fight in dry regions.

**[20:02]** closed loop liquid systems help.

**[20:04]** But watch this issue.

**[20:05]** It decides where facilities

**[20:07]** get built. Who gets paid?

**[20:08]** Largely the same names as the power

**[20:10]** room, Vertiv is the market leader.

**[20:12]** The rare company selling both the

**[20:14]** power gear and the liquid cooling.

**[20:16]** A one-stop shop growing revenue

**[20:17]** twenty-eight percent a year

**[20:19]** at twenty percent margin.

**[20:20]** And then something remarkable happened.

**[20:22]** The two electrical giants each spend

**[20:24]** billions to buy their way into liquid

**[20:26]** cooling within months of each other.

**[20:28]** Eaton paid about nine point

**[20:29]** five billion for Boyd Thermal.

**[20:31]** Schneider Electric bought Motivair.

**[20:33]** When the electrics more

**[20:34]** disciplined industry acquires,

**[20:36]** both pay up for the same niche.

**[20:38]** They are telling you what they think

**[20:40]** every future data center looks like.

**[20:41]** smaller pure players nVent

**[20:43]** including loops and enclosures

**[20:45]** and private cool IT systems.

**[20:47]** The specialists whose cold plates

**[20:49]** ship inside many brand name servers.

**[20:51]** The liquid cooling market was

**[20:52]** about five billion in 2025.

**[20:54]** Forecasts put it at fifteen to

**[20:55]** twenty-seven billion by the early 2030s.

**[20:58]** It's the single clearest picks and

**[21:00]** shovels growth lane in this entire

**[21:01]** ecosystem because it does not care

**[21:03]** whether NVIDIA or AMD or Google wins.

**[21:06]** Heat is heat.

**[21:07]** All right, the factory has power.

**[21:09]** The fever is under control.

**[21:10]** It's time to walk onto the factory

**[21:12]** floor and meet the machine your dinner

**[21:14]** question actually runs on and the three

**[21:16]** trillion dollar company that built it

**[21:17]** This is the machine your

**[21:18]** dinner question runs through.

**[21:20]** Nvidia's GB200, NVL72 seventy-two

**[21:23]** GPUs wired together so tightly they

**[21:25]** behave as a single giant computer.

**[21:27]** It weighs about a ton and a half,

**[21:29]** draws those one hundred and twenty

**[21:30]** kilowatts we discussed, and costs

**[21:32]** roughly three million dollars.

**[21:34]** So let's open it up and to keep the

**[21:35]** parts straight, come back to the

**[21:37]** factory, specifically its kitchen.

**[21:39]** The CPU is the head chef.

**[21:41]** The central processing unit runs

**[21:42]** the operating system, takes orders,

**[21:45]** coordinates everything, brilliant

**[21:47]** at complex sequential tasks.

**[21:48]** But there's only a handful of them.

**[21:50]** For decades, the CPU was the star

**[21:52]** of computing, Intel's kingdom.

**[21:54]** In the AI server, it has

**[21:55]** been demoted to management

**[21:57]** The GPUs are 10,000 line cooks.

**[21:59]** A graphics processing unit,

**[22:01]** originally invented to draw video

**[22:03]** game graphics, contains thousands

**[22:04]** of small, simple cores that all do

**[22:07]** the same operation simultaneously.

**[22:08]** It turns out the math inside a neural

**[22:10]** network is exactly that kind of work.

**[22:13]** Billions of identical

**[22:14]** multiply and add operations.

**[22:16]** One head chef cannot do that.

**[22:18]** 10,000 line cooks, each chopping

**[22:20]** one onion at the same instant, can

**[22:22]** That accident of history, gaming

**[22:24]** graphics and AI needing the same math,

**[22:26]** is the foundation of Nvidia's empire.

**[22:29]** HBM is the countertop,

**[22:30]** high bandwidth memory.

**[22:32]** Hold that thought.

**[22:33]** It gets its own act.

**[22:34]** The SSD is the pantry.

**[22:36]** The NIC, the network interface

**[22:37]** card, is the waiter carrying

**[22:39]** dishes between kitchens.

**[22:41]** And the power supplies and the motherboard

**[22:43]** are the plumbing and wiring holding

**[22:45]** it all together. Now the companies.

**[22:47]** Nvidia finished its last fiscal year with

**[22:49]** $215.9 billion in revenue, up 65%, of

**[22:53]** which about $194 billion was data center.

**[22:55]** It controls roughly 80 to 86%

**[22:58]** of the AI accelerator market.

**[23:00]** Its gross margin in the recent quarter

**[23:01]** is about 75%, 75% on hardware.

**[23:05]** Apple, the most admired hardware

**[23:07]** company in history, runs around 46.

**[23:09]** Nvidia became the first $5

**[23:10]** trillion company last October.

**[23:12]** And depending on the week,

**[23:13]** roughly seven cents of every

**[23:15]** dollar in the S&P 500 is Nvidia.

**[23:17]** How is that margin possible?

**[23:19]** Everyone says best chips, and sure,

**[23:21]** but the real answer is a word we'll

**[23:23]** unpack fully in act eight, CUDA.

**[23:25]** Twenty years of software that every

**[23:27]** AI developer on Earth was trained on.

**[23:29]** For now, the one line version,

**[23:31]** Nvidia doesn't only sell chips.

**[23:33]** It sells the only complete factory

**[23:34]** floor system the world's engineers

**[23:36]** already know how to operate.

**[23:37]** Buying a competitor's chip means

**[23:39]** retraining your whole workforce

**[23:41]** The bear case, about forty percent

**[23:43]** of Nvidia's revenue comes from just

**[23:44]** four customers, and all four are

**[23:46]** building their own chips to replace it.

**[23:48]** Next, the challenger, AMD.

**[23:50]** AMD's Instinct GPUs are genuinely

**[23:52]** competitive on inference, more memory per

**[23:54]** chip, and by some estimates, twenty-five

**[23:57]** to forty percent better tokens per dollar.

**[23:59]** Their problem was never silicon.

**[24:01]** It's software.

**[24:02]** Their CUDA alternative, called

**[24:03]** ROCm, now hits ninety to

**[24:05]** ninety-five percent of Nvidia's

**[24:07]** performance on standard workloads.

**[24:09]** But ninety percent as good with more

**[24:10]** friction is a hard pitch when your

**[24:13]** training can cost a hundred million.

**[24:14]** AMD holds maybe five to

**[24:16]** seven percent of this market.

**[24:17]** Watchable, but it's improving.

**[24:19]** still a distant second.

**[24:21]** Intel, painful to say,

**[24:22]** is barely in this race.

**[24:24]** Still selling plenty of head chefs, but

**[24:26]** the kitchen stopped being about head chefs

**[24:28]** Now the quiet assassin, Broadcom.

**[24:30]** Here's the plot twist most

**[24:31]** retail investors miss.

**[24:33]** Those four hyperscalers building their own

**[24:34]** chips, they cannot actually do it alone.

**[24:37]** Designing a frontier AI chip

**[24:38]** takes a decade of specialized IP.

**[24:40]** So they hire Broadcom, which co-designs

**[24:43]** Google's TPU, Meta's training

**[24:44]** chip, and reportedly OpenAI's.

**[24:47]** Broadcom's books over sixty

**[24:49]** percent of custom chip market.

**[24:50]** Posted AI revenue up one hundred

**[24:52]** and six percent last quarter, with

**[24:54]** a seventy-three billion backlog.

**[24:55]** And management says it has lines

**[24:57]** of sight to a hundred billion of

**[24:58]** AI revenue in twenty twenty-seven.

**[25:00]** In our factory analogy, Nvidia

**[25:02]** sells finished kitchens.

**[25:03]** Broadcom helps the biggest

**[25:05]** restaurant chains build their

**[25:06]** own and takes a cut either way.

**[25:08]** It also dominates the

**[25:09]** switch silicon in ASIC.

**[25:11]** One company, both sides of the war

**[25:13]** Now, who actually built those racks?

**[25:15]** Not Nvidia.

**[25:16]** Nvidia designs.

**[25:17]** Super Micro integrate full liquid

**[25:19]** cooled racks faster than anyone.

**[25:21]** Revenue up 123% last quarter.

**[25:23]** Over 90% of it AI.

**[25:25]** Gross margin between six and

**[25:26]** 10%, depending on the quarter.

**[25:28]** Dell has taken over 64 billion in

**[25:30]** AI server orders with a 43 billion

**[25:33]** backlog, a staggering number as

**[25:35]** server segment margins under 9%.

**[25:37]** HPE writes supercomputing

**[25:39]** and sovereign AI deals.

**[25:40]** and beneath the brands sit the true

**[25:42]** invisible giants, the Taiwanese ODMs,

**[25:44]** original design manufacturers, Foxconn,

**[25:47]** which assembles roughly 40% of the

**[25:49]** world's AI racks, Quanta, Wiwynn

**[25:51]** Celestica.

**[25:52]** The same rack passed through many hands.

**[25:54]** Nvidia captures seventy-five

**[25:55]** percent margin on the silicon.

**[25:57]** The company that physically screws

**[25:58]** it all together keeps six to 10.

**[26:00]** In hardware, the profit

**[26:02]** lives in whatever is scarce.

**[26:03]** Chips and software are scarce.

**[26:05]** Assembly is not.

**[26:06]** But I've been hiding something from you.

**[26:08]** I said seventy-two GPUs

**[26:09]** behave as a single computer.

**[26:11]** Nobody hit that with a magic wand.

**[26:13]** Making ten thousand line cooks work

**[26:15]** as one brain is arguably the hardest

**[26:17]** engineering problem in the entire

**[26:19]** building, and it's where some of the

**[26:20]** best businesses in the ecosystem hide

**[26:22]** So why can't one GPU do the job?

**[26:25]** Simple.

**[26:26]** The model doesn't fit.

**[26:27]** A Frontier model has over

**[26:28]** one trillion parameters.

**[26:30]** Those adjacent knobs require

**[26:32]** terabytes of ultra-fast memory.

**[26:34]** The biggest GPU carries

**[26:35]** a few hundred gigabytes.

**[26:36]** so the model gets sliced

**[26:37]** across thousands of chips.

**[26:39]** And here's the consequence.

**[26:40]** To produce every single token, those

**[26:42]** chips must exchange intermediate results

**[26:45]** constantly at unimaginable speed.

**[26:48]** Back to the kitchen.

**[26:49]** Ten thousand line cooks

**[26:50]** preparing one dish together.

**[26:52]** every chef needs ingredients

**[26:53]** from other chefs every second.

**[26:55]** If passing ingredient is slow,

**[26:57]** your ten thousand chefs stand

**[26:59]** around waiting, and these are the

**[27:00]** most expensive chefs in history.

**[27:02]** at cluster scale, a network

**[27:04]** that's ten percent slower can idle

**[27:06]** billions of dollars of silicon.

**[27:08]** That's why networking is roughly

**[27:09]** forty to sixty percent of

**[27:11]** spending for every dollar of GPU.

**[27:13]** Two words to define.

**[27:14]** Bandwidth is how much

**[27:16]** data moves per second.

**[27:17]** the width of the conveyor belt.

**[27:19]** Latency is the delay for one handoff.

**[27:22]** How long a single pass takes.

**[27:23]** AI needs both everywhere at once.

**[27:26]** The wiring comes in two flavors.

**[27:28]** Scale up inside the rack is Nvidia's

**[27:31]** proprietary NVLink, an extreme speed web

**[27:33]** that makes seventy-two GPU one machine.

**[27:35]** Scale out rack to rack across the

**[27:38]** building is where a war is being fought.

**[27:40]** For years, serious AI clusters ran on

**[27:43]** InfiniBand, a specialized ultra-low

**[27:45]** latency networking that NVIDIA

**[27:47]** acquired in 2019 with Mellanox.

**[27:50]** Premium performance,

**[27:51]** premium price, one vendor.

**[27:53]** Against it, Ethernet, the open

**[27:55]** universal standard, the same family

**[27:57]** of technology as your home network.

**[27:59]** Historically slower, but backed by

**[28:01]** literally everyone who isn't NVIDIA.

**[28:03]** By early 2026, about two-thirds of

**[28:05]** new AI cluster networking is Ethernet.

**[28:08]** Open standards, given time, usually win.

**[28:10]** They just did.

**[28:11]** This is a proprietary versus

**[28:13]** open story as old as tech

**[28:15]** who profit from the nervous system?

**[28:17]** Broadcom again.

**[28:18]** Its Tomahawk chips are the

**[28:20]** merchant silicon inside most

**[28:22]** high-end Ethernet switches.

**[28:23]** Arista Networks build the switches

**[28:25]** themselves, the best in class boxes and

**[28:27]** software hyperscalers standardize on.

**[28:29]** Roughly six percent gross margin, guiding

**[28:31]** to eleven point five billion this year,

**[28:33]** with the risk that its two biggest

**[28:35]** customers are forty percent of revenue.

**[28:37]** Cisco, the incumbent, still huge,

**[28:40]** fighting to stay relevant in AI backend.

**[28:42]** Marvell plays the Broadcom

**[28:43]** playbook one tier down.

**[28:45]** custom chips for Amazon and Microsoft,

**[28:47]** plus leadership in the digital signal

**[28:49]** processor inside optical modules.

**[28:52]** And Astera Labs, one of the most

**[28:54]** remarkable margin stories in ecosystem,

**[28:56]** Makes tiny retimer chips that clean up

**[28:59]** electrical signals degrading over mere

**[29:01]** inches of circuit board at this speed.

**[29:04]** Boring, invisible, seventy-six

**[29:06]** percent gross margins.

**[29:07]** Revenue up ninety-three

**[29:09]** percent last quarter.

**[29:10]** When data moves this fast, even the

**[29:12]** space between two chips become a market

**[29:14]** And then the plot

**[29:15]** literally turned to light.

**[29:16]** copper wires can carry this speed only

**[29:18]** at a few meters before signals degrade.

**[29:20]** Fine inside a rack.

**[29:22]** Use this across a football field building.

**[29:24]** So between racks, everything converts

**[29:26]** to light through glass fiber.

**[29:28]** The device doing the conversion is the

**[29:29]** optical transceiver, a thumb-sized gadget,

**[29:32]** an electrical to light translator, and you

**[29:34]** need one at each end of every fiber link.

**[29:37]** A single large AI cluster consumes

**[29:38]** hundreds of thousands of them at

**[29:40]** hundreds of thousands of dollars

**[29:41]** each, replaced every upgrade cycle.

**[29:44]** It's a razor blade for the data centers.

**[29:46]** The names?

**[29:47]** Coherent, the market leader, Lumentum

**[29:49]** Chinese volume champion Innolight.

**[29:51]** Fabrinet.

**[29:52]** The contract manufacturer that

**[29:54]** assemble for nearly all of them.

**[29:55]** The arms dealer's arms dealer.

**[29:57]** Corning, which draws

**[29:58]** the glass fiber itself,

**[29:59]** And Amphenol, whose connectors are

**[30:02]** the knuckles of the entire system.

**[30:04]** The frontier to watch is

**[30:05]** co-packaged optics, moving the

**[30:07]** light conversion directly onto

**[30:08]** the switch chip to slash power

**[30:10]** The nervous system is built.

**[30:11]** 10,000 chips thinking as one.

**[30:14]** But there's a dirty secret

**[30:15]** on the factory floor.

**[30:16]** Most of the time, the most expensive

**[30:18]** chips in the world are waiting, not for

**[30:20]** data from across the room, for data from

**[30:22]** two centimeters away. Here's a secret.

**[30:25]** During the decode phase, the one word at a

**[30:27]** time assembly line from act two, the GPU's

**[30:30]** math cores are often not the bottleneck.

**[30:32]** For every token, the chip must

**[30:34]** pull the model's parameters and the

**[30:36]** conversation's working memory, that

**[30:38]** KV cache, from memory into its cores.

**[30:40]** The math is fast.

**[30:42]** The fetching is slow.

**[30:43]** Modern inference is what engineers

**[30:45]** call memory bandwidth bound.

**[30:47]** The line cooks are lightning,

**[30:48]** but a countertop cannot feed them

**[30:50]** ingredients fast enough, which makes

**[30:51]** the countertop one of the most valuable

**[30:53]** pieces of real estate in technology

**[30:55]** The industry's answer is

**[30:57]** HBM, high bandwidth memory.

**[30:59]** Instead of laying memory chips flat on

**[31:01]** a board a few inches from the processor,

**[31:03]** HBM stacks them vertically, eight,

**[31:06]** 12 stories high, drills thousands of

**[31:08]** microscopic elevator shafts through the

**[31:10]** silicon, and glues the whole tower directly

**[31:12]** next to the GPU on the same package.

**[31:14]** A skyscraper of memory downtown instead

**[31:17]** of suburbs of memory across a highway.

**[31:19]** The result is five to six times the

**[31:21]** bandwidth of conventional memory

**[31:22]** at five to six times the price.

**[31:24]** Nvidia happily pays.

**[31:25]** Memory is now one of the biggest

**[31:27]** cost components inside every AI chip

**[31:29]** you have heard of, and only three

**[31:30]** companies on earth can make it

**[31:32]** SK Hynix, the Korean company, owns

**[31:34]** roughly 60% of HBM and got there by

**[31:37]** out-executing its giant neighbor.

**[31:39]** It bet on HBM years before it

**[31:41]** mattered and shipped each generation

**[31:43]** first, unlocking the lion's share of

**[31:45]** Nvidia's next generation allocation.

**[31:47]** Samsung, the largest memory company

**[31:49]** overall, was embarrassingly late.

**[31:51]** Micron, the American champion, went

**[31:53]** from afterthought to selling out

**[31:55]** its entire 2026 HBM capacity in

**[31:58]** advance. Here's the stat for the act.

**[32:00]** In May 2026, all three memory makers

**[32:02]** crossed a trillion dollars in market value,

**[32:04]** combined over four trillion, roughly

**[32:06]** sixteen times what they were a decade ago.

**[32:08]** Memory used to be the most brutal

**[32:10]** commodity business in tech.

**[32:11]** Boom, bust, bankrupt, repeat.

**[32:14]** HBM changed the psychology.

**[32:16]** It sold out more than a year in

**[32:17]** advance, allocated like a scarce

**[32:19]** resource, priced like a luxury good.

**[32:21]** The open question is whether that price

**[32:23]** survives the moment all three giants

**[32:25]** finish their capacity expansions at once.

**[32:27]** Memory cycles have broken hearts before

**[32:30]** Now walk out the back of the

**[32:31]** factory to the warehouse storage.

**[32:33]** The hierarchy in one line, the closer

**[32:35]** to the chip, the faster and pricier.

**[32:37]** Cache on the chip itself, HBM beside

**[32:40]** it, regular DRAM on the motherboard,

**[32:42]** SSDs, flash drives for hot data, and

**[32:45]** at the bottom, the technology everyone

**[32:47]** declared dead ten years ago, the

**[32:49]** spinning hard drive, still unbeatable

**[32:51]** per terabyte for cold bulk data.

**[32:53]** And AI turned out to be an AI hoarder.

**[32:55]** Training datasets, model

**[32:56]** checkpoints saved every few hours.

**[32:58]** And the part nobody predicted, the output.

**[33:01]** Every conversation, every log, every

**[33:03]** generated image retained forever.

**[33:05]** Seagate CEO calls it the

**[33:07]** inference inflection.

**[33:08]** AI doesn't just consume data,

**[33:10]** It produces it endlessly.

**[33:12]** Hard drives are a literal duopoly,

**[33:14]** Seagate and Western Digital, plus flash

**[33:16]** drive players Kioxia and Solidigm.

**[33:19]** After a decade of decline, both

**[33:20]** drive makers sold out their

**[33:22]** entire production into 2027.

**[33:24]** Western Digital now ships 89% of its

**[33:26]** revenue to cloud customers and was

**[33:28]** one of the S&P 500 top performers.

**[33:30]** Consumer hard drive prices jumped

**[33:32]** 50% because AI ate the supply.

**[33:34]** A dying industry resurrected by

**[33:36]** the factory next door needing

**[33:37]** somewhere to put infinity.

**[33:39]** The machine is complete,

**[33:40]** powered, cooled, wired, fed.

**[33:42]** But a pile of perfect hardware

**[33:44]** answers exactly zero questions.

**[33:46]** Something invisible has to run the place.

**[33:48]** Everything we have toured so

**[33:49]** far, you could theoretically buy.

**[33:51]** Hardware is purchasable, where

**[33:52]** does the durable competitive

**[33:53]** advantage actually live?

**[33:55]** The stack briefly bottom to top.

**[33:57]** Linux, the free operating system running

**[33:59]** effectively every server on Earth,

**[34:01]** commercialized by Red Hat and Canonical.

**[34:04]** Kubernetes, the invisible foreman.

**[34:06]** Open source software that schedules

**[34:07]** work across thousands of machines.

**[34:09]** restarts what crashes and keeps

**[34:11]** the factory floor humming.

**[34:12]** Then the serving layer, then the models.

**[34:15]** Two stories in this act matter to

**[34:16]** investors more than all the rest combined.

**[34:18]** First story, CUDA, the twenty-year

**[34:20]** trap. In 2006, Nvidia made a

**[34:22]** decision Wall Street hated.

**[34:24]** It spent billions building a

**[34:25]** programming platform so scientists

**[34:27]** can use gaming chips for general math.

**[34:29]** For a decade, this looked

**[34:31]** like an expensive hobby.

**[34:32]** Then deep learning arrived, and

**[34:34]** every AI researcher on Earth

**[34:35]** learned to build on CUDA because

**[34:37]** it was the only mature option.

**[34:39]** Twenty years later, CUDA has

**[34:40]** millions of developers, thousands

**[34:42]** of specialized libraries, and every

**[34:44]** framework optimized for it first.

**[34:46]** Understand what this means.

**[34:47]** When AMD ships a chip with better

**[34:49]** specs, and it sometimes does, the

**[34:51]** customer isn't comparing chips.

**[34:53]** They're comparing chips plus the

**[34:55]** cost of retaining their entire

**[34:56]** engineering organization and rewriting

**[34:58]** their code with their one hundred

**[35:00]** million training run on the line.

**[35:01]** That's why seventy percent gross

**[35:03]** margins survive competition.

**[35:04]** The moat was never the silicon.

**[35:06]** The moat is the muscle

**[35:07]** memory of a million engineers.

**[35:09]** Second one, why your question

**[35:10]** costs cents, not dollars?

**[35:12]** Raw, naive inference on trillion

**[35:13]** parameter model would be super expensive.

**[35:16]** The economics only work because of an

**[35:18]** unglamorous layer called a serving engine.

**[35:20]** Software like vLLM and NVIDIA's

**[35:23]** TensorRT doing three tricks.

**[35:25]** Batching, grouping hundreds of users'

**[35:27]** questions through the chip at once.

**[35:28]** The delivery route trick from Act Two.

**[35:30]** Caching, reusing the KV working

**[35:32]** memory instead of recomputing the

**[35:34]** conversation from scratch for every word.

**[35:36]** And quantization.

**[35:38]** Rounding the model's numbers to lower

**[35:39]** precision, like shipping a slightly

**[35:41]** compressed photo, nearly identical

**[35:43]** quality, fraction of the cost.

**[35:45]** Together a three to ten times cost

**[35:47]** reduction from software alone.

**[35:49]** When OpenAI or Anthropic cuts

**[35:51]** API price 80% in a year, it's

**[35:53]** mostly this layer, not new chips.

**[35:55]** Token manufacturing costs

**[35:56]** are collapsing our curve.

**[35:57]** Remember that for the finale,

**[35:59]** because it cuts both ways.

**[36:00]** One last floor of the stack.

**[36:02]** Your dinner question doesn't

**[36:03]** need it, but enterprise AI does.

**[36:05]** RAG.

**[36:06]** Retrieval augmented generation.

**[36:08]** The model is a brilliant

**[36:09]** writer with a fixed education.

**[36:10]** RAG hands it your company's documents

**[36:13]** at question time via vector database.

**[36:15]** search engine that finds text by

**[36:17]** meaning rather than keywords from

**[36:19]** players like Pinecone, plus the data.

**[36:21]** platform Databricks and Snowflake.

**[36:23]** The librarian and the

**[36:24]** writer working together.

**[36:26]** That's the enterprise AI

**[36:27]** pitch in one sentence.

**[36:28]** And with that, the tour is over.

**[36:30]** You have seen every layer: power,

**[36:32]** cooling, silicon, light, memory, software.

**[36:35]** Now let's follow the money, all of it

**[36:36]** in one map, and then ask one question:

**[36:39]** Does any of this actually pay for itself?

**[36:41]** So here's the whole board.

**[36:42]** On the demand side, the miners, three

**[36:44]** tiers, the hyperscalers, Microsoft,

**[36:46]** Amazon, Google, Meta, Oracle,

**[36:48]** spending their combined seven hundred

**[36:50]** and twenty-five billion this year.

**[36:51]** The neoclouds, specialized GPU landlords,

**[36:54]** CoreWeave, Nebius, Lambda, Crusoe

**[36:56]** And AI labs, OpenAI now

**[36:59]** valued around $850 billion.

**[37:01]** Anthropic, which passed it at

**[37:03]** roughly $965 billion on about $47

**[37:05]** billion of annualized revenue.

**[37:06]** And XAI folded into SpaceX.

**[37:09]** Both major labs filed for IPO this June.

**[37:12]** The private market has already

**[37:13]** priced them as two of the most

**[37:14]** valuable companies on Earth.

**[37:16]** The public market is about

**[37:17]** to vote on the costs.

**[37:19]** Bernstein estimates one gigawatt

**[37:20]** of AI data center, one campus,

**[37:23]** costs about $35 billion to build.

**[37:24]** Where it goes, roughly 39% is the

**[37:27]** chips, the single biggest line, which

**[37:29]** is why Nvidia's gross profit alone is

**[37:31]** estimated near 30% of total industry cost.

**[37:34]** And that thirty-five billion

**[37:35]** machine depreciates fast.

**[37:37]** most operators write chips off over

**[37:39]** four to six years, And skeptics argue

**[37:41]** even that flatters the accounting since

**[37:43]** a five-year-old GPU competes against

**[37:45]** chips 10 times better Every year of

**[37:47]** depreciation life added or removed

**[37:50]** swings billions in reported profit.

**[37:52]** Watch that debate

**[37:53]** so follow this.

**[37:54]** NVIDIA invests billions directly into

**[37:55]** CoreWeave, into Nebius, into OpenAI.

**[37:58]** those companies use the

**[37:59]** money to buy NVIDIA chips.

**[38:01]** revenue for NVIDIA.

**[38:02]** The hyperscalers sign enormous

**[38:03]** contracts with the neoclouds.

**[38:05]** Meta alone committed roughly

**[38:06]** twenty-one billion to CoreWeave and

**[38:08]** reportedly twenty-seven billion to

**[38:09]** Nebius, which conveniently moves

**[38:11]** data center spending off Meta's

**[38:13]** balance sheet into someone else's debt.

**[38:15]** The AI lab signs compute

**[38:16]** deals with hyperscalers.

**[38:17]** OpenAI with Microsoft and Oracle.

**[38:19]** Anthropic with Amazon and Google.

**[38:21]** paid partly with money those same

**[38:23]** hyperscalers invested into them.

**[38:25]** Let's be fair, because the analytical

**[38:27]** skeptic label cuts both ways.

**[38:29]** This is not fraud, and it's not new.

**[38:31]** Vendor financing built the

**[38:32]** railroads and the telephone network.

**[38:34]** The loop has exactly one opening

**[38:36]** where fresh money is supposed to

**[38:38]** enter, end users and enterprises.

**[38:40]** You're paying twenty bucks a month,

**[38:42]** companies paying for APIs and copilots.

**[38:44]** That is the only exit that

**[38:45]** isn't recycled capital.

**[38:47]** And today, that corner is the

**[38:48]** smallest number on the board.

**[38:50]** The two biggest labs combined annualized

**[38:52]** run rates now top seventy billion.

**[38:54]** Generally spectacular.

**[38:55]** The fastest revenue ramp

**[38:56]** in software history.

**[38:58]** But the revenue they'll actually book

**[38:59]** this calendar year is a fraction of

**[39:01]** that, set against seven hundred and

**[39:03]** twenty-five billion of annual spend,

**[39:04]** with OpenAI still expected to lose on

**[39:07]** the order of fourteen billion this year

**[39:09]** The entire $7 trillion machine is

**[39:10]** the bet that a small corner grows

**[39:12]** faster than the big circle spins.

**[39:15]** So let's watch it one more time.

**[39:16]** Same two seconds, but now you can see.

**[39:18]** Your thumb hits send, and the

**[39:20]** question becomes light in glass

**[39:21]** fiber crossing state lines.

**[39:23]** arrive at a building

**[39:24]** drawing the power of a city.

**[39:26]** power from a restarted nuclear

**[39:27]** plant, a Solar turbine, a fuel

**[39:29]** cell parked behind the meter.

**[39:30]** It's chopped into tokens and

**[39:32]** fed into a three million dollar

**[39:34]** rack assembled in Taiwan, sold

**[39:36]** at a seventy-five point margin,

**[39:38]** where ten thousand line cooks

**[39:39]** fetch a trillion parameters

**[39:41]** from memory skyscrapers.

**[39:42]** While liquid coolant carries away

**[39:44]** the heat of eighty space heaters,

**[39:46]** and light speed interconnects let

**[39:48]** ten thousand chips think one thought.

**[39:50]** The software batches you with one

**[39:52]** thousand strangers, and the answer

**[39:54]** streams back token by token, each

**[39:56]** word manufactured the instant you read it

**[39:58]** Two seconds, $700 billion a year,

**[40:00]** the largest infrastructure project

**[40:02]** our species has ever attempted.

**[40:04]** So a machine can suggest you make

**[40:05]** one-pan chicken and rice. Whether

**[40:07]** that's the most important investment

**[40:09]** in history or the most expensive,

**[40:11]** Honestly, nobody on Earth knows yet.

**[40:13]** That's the whole map.

**[40:14]** If you find this video helpful,

**[40:15]** please subscribe, and I'll

**[40:17]** see you in the next one.
