
This paper introduces a Mean-Field Asymptotic framework designed to estimate the hit ratio in multi-turn large language model (LLM) serving systems. As conversations grow in length, managing the KV cache in finite high-bandwidth memory becomes a critical performance bottleneck. The authors model these dynamics using the least-recently-used (LRU) eviction policy to determine which conversation histories are retained or discarded. By analyzing the system as memory capacity and arrival rates scale toward infinity, they derive a closed-form limit to accurately predict cache reuse. The study further proposes a practical estimator that accounts for partially filled, unhashable memory blocks common in real-world applications. Finally, the researchers validate their theoretical findings through experiments with the Qwen3-8B model, demonstrating that their model reliably predicts system performance under varying workloads.
Podzilla Summary coming soon
Sign up to get notified when the full AI-powered summary is ready.
Free forever for up to 3 podcasts. No credit card required.

Jailbreaking Jailbreaks: A Proactive Defense for LLMs

When Agents Slow Down: Understanding LLM Agents’ Test-Time Strategies via Elo-per-token Analysis

Thinking with Looped Flows

Breaking the Token Ceiling: Distilling Smaller, Stronger Byte Models
Free AI-powered recaps of Best AI papers explained and your other favorite podcasts, delivered to your inbox.
Free forever for up to 3 podcasts. No credit card required.