
NVIDIA has become one of the great corporate success stories of the AI boom.Its GPUs power much of the infrastructure behind today’s frontier AI models. Revenue has exploded. Profitability has exploded. Its market capitalization has reached almost incomprehensible levels.And because NVIDIA sells the “shovels” during the AI gold rush, the conventional wisdom seems pretty straightforward:Even if the AI bubble eventually bursts, NVIDIA wins.After all, somebody still has to sell the picks and shovels.I’m not convinced.In fact, I think there is a scenario in which NVIDIA becomes one of the biggest casualties of the next phase of the AI revolution.Not because its technology suddenly becomes bad.But because the economics of AI could fundamentally change.NVIDIA’s Moat Depends on One Big AssumptionThe bull case for NVIDIA ultimately rests on a simple proposition:AI has an essential dependency on NVIDIA GPUs.Today, that proposition looks pretty damn convincing.Frontier models require enormous amounts of computing power. Companies like OpenAI and Anthropic have traditionally trained their models using massive clusters of NVIDIA GPUs inside enormous data centers.And NVIDIA’s data-center business is now overwhelmingly important to the company.The logic therefore seems almost circular:AI gets bigger → AI needs more compute → more compute requires NVIDIA GPUs → NVIDIA makes more money.But what happens if the amount of compute required to produce useful AI falls dramatically?What happens if frontier models become increasingly commoditized?And, perhaps most importantly:What happens if AI inference moves out of the data center and onto the devices sitting on our desks?That’s where things get interesting.The First Problem: Frontier AI Is Becoming CommoditizedOne of the most interesting developments in AI isn’t happening in Silicon Valley.It’s happening in China.U.S. restrictions on advanced NVIDIA chips have forced Chinese AI companies to become extraordinarily creative with limited computing resources. They’ve developed alternative hardware and software stacks while finding ways to train increasingly capable models with less compute.The result is a strange paradox.The harder the United States tried to restrict China’s access to advanced AI hardware, the stronger the incentive became for Chinese companies to figure out how to build AI without it.And we’re now seeing highly capable open-weight models emerge that can compete surprisingly well with leading proprietary systems.The important point isn’t whether one particular Chinese model is better than Claude or ChatGPT.The important point is what happens when the model itself stops being scarce.If someone can download a highly capable frontier-class model for free, the economic value begins moving somewhere else.The model becomes a commodity.And once the model becomes a commodity, the question changes from:“Who has the best AI model?”to:“Where should we host all of this AI?”That distinction could be enormously important for NVIDIA.The AI Revolution Has Two Different ProblemsThere’s a distinction that often gets lost in the AI discussion:Training is not the same thing as inference.Training is the process of creating the model.Inference is what happens every time you actually use it.Every time you ask ChatGPT a question, summarize a document, generate an image, write some code, or run an AI agent, you’re performing inference.And I think inference could become NVIDIA’s Achilles’ heel.Why?Because inference has a very different economic profile from training.For inference, the bottleneck isn’t always raw computational power.It can be memory.Consider a hypothetical near-frontier model with hundreds of billions of parameters.A mixture-of-experts architecture might only activate a relatively small portion of those parameters for any individual token. The actual computation required can therefore be surprisingly manageable.The problem is that the entire model still needs to reside somewhere in memory.That’s where things get interesting.What If Your Mac Can Run Frontier AI?Imagine you want to run Deep Seek V4 Flash, a roughly 284-billion-parameter model locally.You might need around 90–100 GB of memory to hold the model.NVIDIA’s obvious solution is to use an expensive data-center GPU with enormous amounts of high-speed VRAM.And if the model gets even larger?Add more GPUs.Connect them using NVIDIA’s proprietary high-speed interconnect technology.Add network
Podzilla Summary coming soon
Sign up to get notified when the full AI-powered summary is ready.
Free forever for up to 3 podcasts. No credit card required.

LIVE Office Hours: The Tech Job Apocalypse Isn’t Over

Could AI actually destroy humanity within the next five years?

I Replaced Claude With Chinese AI?

AI Just Destroyed the Internet?
Free AI-powered recaps of AsianDadEnergy's Substack Podcast and your other favorite podcasts, delivered to your inbox.
Free forever for up to 3 podcasts. No credit card required.