
Free Daily Podcast Summary
by The New Stack
The New Stack Podcast is all about the developers, software engineers and operations people who build at-scale architectures that change the way we develop and deploy software.
The most recent episodes — sign up to get AI-powered summaries of each one.
Harness Field CTO Martin Reynolds joins The New Stack to talk about what happens after coding agents start opening pull requests faster than anyone can review them. He explains how he first saw the bottleneck during early GitHub Copilot trials, the three ways enterprises are coping with the volume now, and why Harness rebuilt its Code Repository and launched AI Code Review for agent traffic. The conversation also covers GitHub's recent outages, the software delivery knowledge graph behind Harness's reviewer, and how much of the delivery pipeline should stay deterministic. Learn more from The New Stack around the latest in coding agents:AI coding agents can write code, Crafting wants to help them ship itGit real: AI agents aren't just for solo developers anymoreJoin our community of newsletter subscribers to stay on top of the news and at the top of your game.
Traces provide a detailed view of a request’s journey through data, microservices and applications, helping SREs pinpoint where failures occur and resolve issues faster. But while tracing can reduce downtime and developer burnout, collecting every trace creates its own problems. Storing massive volumes of data is expensive, can burden the systems being monitored and makes it harder to find the information that actually matters.The solution isn’t abandoning tracing, but being smarter about what gets retained. Head sampling captures only a portion of traces upfront, while tail sampling evaluates completed traces and keeps those most valuable for troubleshooting. Dynamic sampling goes further by filtering repetitive or nearly identical traces before they overwhelm storage.On The New Stack podcast, Sarah Hudspeth of Chronosphere, a Palo Alto Networks company, explains how teams can build a more effective tracing strategy. She breaks down how thoughtful sampling and observability design can turn tracing from a data-hoarding problem into a practical tool for production troubleshooting. Learn more from The New Stack around the latest in tracing:Sampling: the philosopher’s stone of distributed tracingHow OpenTelemetry Works: Tracing, Metrics and Logs on KubernetesWhy Synthetic Tracing Delivers Better Data, Not Just More DataJoin our community of newsletter subscribers to stay on top of the news and at the top of your game.
As AI evolves from conversational chatbots to autonomous agents, CPUs are becoming an increasingly important part of the infrastructure equation. In this episode, The New Stack speaks with Bhumik Patel of Arm and Mo Farhat of Google about how CPUs act as an “air traffic controller” for agentic workloads, handling orchestration, data preparation, semantic search, vector databases, code execution and API calls alongside GPUs and TPUs. Smaller AI models, including summarizers and evaluators, can also run effectively on CPUs for specialized tasks. As agents increasingly generate and execute code, secure sandboxing becomes critical. Google’s gVisor and GKE Agent Sandbox provide isolation and scalable environments, with the latter supporting up to 300 sandboxes per second per cluster. The discussion also explores efficiency and cost, with Google highlighting Axion’s price-performance and energy-efficiency advantages across different workload types. Ultimately, the shift toward agentic AI is creating a more diverse compute environment where CPUs, GPUs and TPUs each play complementary roles in delivering scalable, efficient AI applications. Learn more from The New Stack around the latest in CPUs in the world of AI agents: AI Agents Will Eat Enterprise Software, Just Not in One Bite How to ground AI agents in accurate, context-rich data Join our community of newsletter subscribers to stay on top of the news and at the top of your game.
Doist CTO Gonzalo Silva says AI is reshaping software development, but success depends on restraint rather than rapid feature expansion. Instead of chasing every AI capability, Doist prioritizes “subtraction over addition,” removing features that fail to deliver lasting value despite development investment. After experimenting with nearly 20 AI concepts, the company found success with Ramble, an AI-powered voice task capture feature, while remaining model-agnostic through rigorous testing and evaluations. Internally, developers use a variety of AI coding tools rather than standardizing on one platform, while Doist OS—a companywide AI assistant with nearly 100 shared skills—helps employees across all functions work more effectively. Silva also outlined Doist’s approach to AI-powered automations, separating AI-driven workflow generation from deterministic execution to improve reliability and reduce token costs. Throughout its AI strategy, the company emphasizes purposeful features, privacy, transparency, and continuous improvement, ensuring AI enhances user productivity without compromising product quality or trust. Learn more from The New Stack around developer productivity: Developer Productivity in 2025: More AI, but Mixed Results Optimizing for Developer Productivity Creates a Winning DevEx Join our community of newsletter subscribers to stay on top of the news and at the top of your game.
As AI coding agents accelerate software development, they also create new challenges for site reliability engineers (SREs), who are increasingly responsible for debugging systems that no single human fully understands. In this episode ofThe New Stackpodcast, Sam Farid and Nate Heinrich of Chronosphere argue that AI agents should also be used for root-cause analysis, helping teams diagnose failures more quickly as model capabilities continue to improve. Rather than immediately purchasing a commercial solution, they recommend organizations first build an in-house AI SRE. The process of documenting systems, dependencies, and operational knowledge creates valuable context that enables AI agents to troubleshoot effectively while improving institutional knowledge. Although Chronosphere offers its own AI SRE platform, the hosts emphasize that building an internal prototype helps teams understand their needs before evaluating vendor tools. As AI-generated code becomes more common, organizations that invest in mapping their systems and leveraging AI for operations will be better equipped to reduce downtime and support increasingly complex software environments. Learn more from The New Stack around AI SREs: 5 ways SRE AI agents are set to augment human capabilities The Future of AI in SRE: Preventing Failures, Not Fixing Them AI Reliability Engineering: Welcome to the Third Age of SRE Join our community of newsletter subscribers to stay on top of the news and at the top of your game.
In this episode with The New Stack Agents, Frederic Lardinois, NVIDIA’s Joey Conway says advances in AI over the past year have dramatically improved the capabilities of local models, making them practical for enterprise and personal use alongside frontier cloud models. Rather than replacing large models, Conway envisions a “system of models” where specialized local models handle routine, cost-sensitive, or privacy-focused tasks, while larger frontier models tackle more complex reasoning. He explains that organizations can fine-tune smaller open models using domain-specific data, creating expert AI agents that reflect the specialized roles found within businesses. NVIDIA supports this ecosystem through open models, training tools, and software such as NeMo, Dynamo, and Nemotron. Conway also highlights the growing importance of agentic harnesses, which give AI models access to tools, memory, and iterative workflows, significantly improving performance and reducing costs. Looking ahead, he expects AI orchestration to become increasingly important, with intelligent routing systems selecting the right model for each task based on complexity, cost, latency, and data governance requirements, enabling enterprises to balance performance, security, and efficiency. Learn more from The New Stack around NVIDIA's latest updates in AI: Palantir and Nvidia want to change who owns government AI Nvidia's best model is now live Join our community of newsletter subscribers to stay on top of the news and at the top of your game.
In this episode, Mark Russinovich, CTO of Microsoft Azure revealed Brain, the AI-powered AIOps system that continuously monitors Azure’s health, detects incidents, identifies root causes, and increasingly automates responses such as pausing problematic deployments and notifying affected customers. Built on Azure Resource Graph, Brain creates a real-time digital twin of Azure, mapping dependencies across hundreds of services, data centers, and regions. Although Brain predates the generative AI boom, years of data engineering, standardized service-level indicators (SLIs), and machine learning laid the foundation for today’s capabilities. Brain combines standardized SLIs, service-specific monitoring, and third-party signals to detect anomalies, while ML models dynamically establish service baselines and correlate outages with software rollouts. Microsoft says automated notifications have reduced customer support tickets by four to six times, with 80–90% of Brain-covered services receiving notifications within 15 minutes, often in under five. The company is also layering LLM-powered agents, called Triangle, on top of Brain to streamline incident routing and eventually enable AI agents to autonomously troubleshoot and remediate outages. Learn more from The New Stack around the latest in Microsoft Azure: Meet Brain, the AI that decides when Azure is officially down Microsoft's pitch to enterprises: Ditch Azure Repos for GitHub, despite its rocky reliability record Join our community of newsletter subscribers to stay on top of the news and at the top of your game.
Subquadratic is beginning to back up its ambitious claims with benchmarks and third-party validation for its SubQ 1.1 Small model, which uses its proprietary Sparse Attention (SSA) architecture to dramatically improve long-context performance. Rather than comparing every token to every other token, SSA selectively processes relationships, enabling near-linear scaling while maintaining high accuracy across context windows of up to 12 million tokens. The company reports near-perfect retrieval performance, competitive coding and reasoning benchmarks, and compute savings of up to 1,000x at maximum context lengths. Rather than targeting frontier models immediately, Subquadratic is focusing on enterprise customers that need efficient analysis of massive datasets. The current model was built by replacing the dense attention mechanism in an existing open-weight model and then continuing long-context pretraining. Looking ahead, the startup plans to release a larger mid-tier model while continuing research into "zero attention" architectures that could eliminate attention mechanisms altogether, with the long-term goal of surpassing today's transformer-based AI models in both efficiency and capability. Learn more from The New Stack around cloud spending: The context window has been shattered: Subquadratic debuts a 12-million-token window What comes after attention? This startup says it already knows. Join our community of newsletter subscribers to stay on top of the news and at the top of your game.
Free AI-powered daily recaps. Key takeaways, quotes, and mentions — in a 5-minute read.
Get Free Summaries →Free forever for up to 3 podcasts. No credit card required.
Listeners also like.

NVIDIA AI Podcast
Explores how artificial intelligence and emerging technologies are driving innovation across science, sustainability, and industry.

How I AI
A practical guide to using AI tools in work and life, featuring guests who share specific, actionable techniques and workflows.

Latent Space: The AI Engineer Podcast
Explores AI engineering breakthroughs in foundation models, code generation, and AI agents through interviews with researchers and developers.

Everyday AI Podcast – An AI and ChatGPT Podcast
Practical AI and ChatGPT tips for professionals to improve productivity and grow their careers.

The AI Daily Brief: Artificial Intelligence News and Analysis
A daily analysis of artificial intelligence news, exploring its creative potential, industry impacts, and ethical challenges.

Technology Now
A podcast exploring cutting-edge technology trends and innovations through interviews with industry leaders and HPE experts.

OpenAI Podcast
Conversations with OpenAI researchers and builders exploring how frontier AI models are developed and used in practice.

"The Cognitive Revolution"
Interviews with AI developers and researchers exploring the transformative impact of artificial intelligence on society and technology.

AI and I
Interviews with professionals who use AI tools in their work, exploring how AI affects creativity, thinking, and daily life through live demonstrations.

The Engineering Leadership Podcast
Discusses essential practices and insights from top software engineering leaders to advance leadership skills in tech.

Training Data
Experts discuss AI advancements and their impact on technology, business, and society with insights from leading researchers and builders.

AI For Humans: Weekly AI News, Tools & Trends
A weekly breakdown of major AI news, tools, and breakthroughs for both newcomers and seasoned enthusiasts.
The New Stack Podcast is all about the developers, software engineers and operations people who build at-scale architectures that change the way we develop and deploy software.
AI-powered recaps with compact key takeaways, quotes, and insights.
Get key takeaways from The New Stack Podcast in a 5-minute read.
Stay current on your favorite podcasts without falling behind.
It's a free AI-powered email that summarizes new episodes of The New Stack Podcast as soon as they're published. You get the key takeaways, notable quotes, and links & mentions — all in a quick read.
When a new episode drops, our AI transcribes and analyzes it, then generates a personalized summary tailored to your interests and profession. It's delivered to your inbox every morning.
No. Podzilla is an independent service that summarizes publicly available podcast content. We're not affiliated with or endorsed by The New Stack.
Absolutely! The free plan covers up to 3 podcasts. Upgrade to Pro for 15, or Premium for 50. Browse our full catalog at /podcasts.
The New Stack Podcast publishes weekly. Our AI generates a summary within hours of each new episode.
The New Stack Podcast covers topics including News, Technology. Our AI identifies the specific themes in each episode and highlights what matters most to you.
Free forever for up to 3 podcasts. No credit card required.
Free forever for up to 3 podcasts. No credit card required.