
A 30B parameter model runs on a MacBook because only 3B parameters fire per token. Mixture of Experts splits memory cost from compute cost, and that changes everything about where AI can run.
Podzilla Summary coming soon
Sign up to get notified when the full AI-powered summary is ready.
Free forever for up to 3 podcasts. No credit card required.

The Evaluation Crisis: We Do Not Know How Good Our Models Actually Are

The Agent Interoperability Problem: Why Your AI Agents Can Not Talk to Each Other

KV Cache Compression: The Memory Wall Nobody Talks About

Context Rot: Why Million-Token Windows Quietly Fail
Free AI-powered recaps of The Practical AI Digest and your other favorite podcasts, delivered to your inbox.
Free forever for up to 3 podcasts. No credit card required.