Microsoft Mechanics Podcast

Tokenomics | The new AI currency & your options explained

July 9, 2026·14 min
Episode Description from the Publisher

Control what you're billed on. Tokens are the currency of AI, and how you design your app determines how many you spend. Compress conversation history instead of resending it raw, cap output tokens with matching prompt instructions, and cache static context so each reuse costs a fraction of the first request. Route each prompt to the right model by complexity, or set Model Router in Microsoft Foundry to handle that automatically — balanced, quality, or cost mode. Then optimize the whole stack. Run Agent Optimizer to test your prompt, model, and tool configurations together and surface better setups. Use Toolbox to dynamically select only the tools each request needs and cut input token overhead by 90%.  April Gittens, Microsoft Principal Cloud Advocate, joins Jeremy Chapman, Microsoft 365 Director, to share how to seize control of AI token spend through smarter app design.  ► QUICK LINKS:  00:00 - Tokenomics foundation 01:06 - Token cost basics 02:11 - Context Window creep 03:46 - Reduce unnecessary tokens 04:29 - Trim context costs 05:21 - Cap output tokens 06:27 - Cache for savings 07:54 - Model cost tradeoffs 09:15 - Model Router 09:42 - Toolbox in Microsoft Foundry Toolbox 11:18 - Agent Optimizer <span class="ytAttributedStringLinkInherit

Podzilla Summary coming soon

Sign up to get notified when the full AI-powered summary is ready.

Get Free Summaries →

Free forever for up to 3 podcasts. No credit card required.

Listen to This Episode

Get summaries like this every morning.

Free AI-powered recaps of Microsoft Mechanics Podcast and your other favorite podcasts, delivered to your inbox.

Get Free Summaries →

Free forever for up to 3 podcasts. No credit card required.