Best AI papers explained

Multi-Turn LLM Conversations under the Least-Recently-Used Policy: Mean-Field Asymptotics

September 17, 2026·22 min
Episode Description from the Publisher

This paper introduces a Mean-Field Asymptotic framework designed to estimate the hit ratio in multi-turn large language model (LLM) serving systems. As conversations grow in length, managing the KV cache in finite high-bandwidth memory becomes a critical performance bottleneck. The authors model these dynamics using the least-recently-used (LRU) eviction policy to determine which conversation histories are retained or discarded. By analyzing the system as memory capacity and arrival rates scale toward infinity, they derive a closed-form limit to accurately predict cache reuse. The study further proposes a practical estimator that accounts for partially filled, unhashable memory blocks common in real-world applications. Finally, the researchers validate their theoretical findings through experiments with the Qwen3-8B model, demonstrating that their model reliably predicts system performance under varying workloads.

Podzilla Summary coming soon

Sign up to get notified when the full AI-powered summary is ready.

Get Free Summaries →

Free forever for up to 3 podcasts. No credit card required.

Listen to This Episode

Get summaries like this every morning.

Free AI-powered recaps of Best AI papers explained and your other favorite podcasts, delivered to your inbox.

Get Free Summaries →

Free forever for up to 3 podcasts. No credit card required.