
今天的节目是一集学习播客,学习Kimi K3技术报告,希望和大家一起领略“技术之美”。K3是有效扩展到2.8T总参数并全量开源的MoE模型。大家可能注意到,这集技术报告的领读距离模型发布已经过去一段时间。期间,我们试图寻找一位合适的嘉宾,我们希望这位嘉宾的学术和工作背景非常适合讲K3。最后找到孙宇涛。宇涛目前是清华大学计算机系博士候选人,上海创智学院璞锐学者。他从博士开始的研究方向是LLM架构、预训练,架构创新一直是他的兴趣点,这正好是K3亮点之一。宇涛透过领读K3技术报告也串联讲解了十多篇相关论文。他的语速非常快——前方语速预警。OUTLINE:02:00 宇涛的自我介绍和研究经历,为什么对架构创新最感兴趣?16:05 从Kimi K3论文出发,论文讲解的框架与脉络19:02 宇涛开始领读论文:1 导读Kimi K3将设计概括为沿sequence、depth和width三个维度扩展信息流;其中depth与width本质上仍是模型容量扩展的两种组织⽅式,这⼀“三维 scaling”更多是贯穿论⽂的叙事框架。2 Model Architecture2.1 线性注意⼒的前世今⽣[Microsoft Research] RetNet:最初的 data-independent decay 与 chunk-recurrent 递归形式[NVIDIA] Gated DeltaNet:引⼊ gated delta rule[Moonshot AI] Kimi Linear:fine-grained decay及其带来的infra变化2.2 Gated MLA[Alibaba Qwen] Gated Attention for Large Language Models:attention gating与训练稳定性2.3 Attention Residuals[Microsoft Research] On Layer Normalization in the Transformer Architecture:Pre-LN、Post-LN与训练稳定性[ByteDance Seed] Hyper-Connections:残差连接的新维度[Moonshot AI] Attention Residuals:以跨深度attention取代固定残差累积,使各层选择性聚合此前表⽰2.4 Stable LatentMoE与SiTU-GLU[NVIDIA] LatentMoE:通过latent expert space降低MoE通信与权重访问成本[OpenAI] GPT-OSS:以clamped SwiGLU控制激活值2.5 Muon[Moonshot AI] Moonlight / Muon is Scalable for LLM Training:更好的优化表现及随之突出的 activation outlier问题2.6 Quantile Balancing[科学空间] 《MoE环游记:6、最优分配促均衡》:从最优分配推导Quantile Balancing2.7 Native Vision3 Pre-Training3.1 Scaling Law[OpenBMB] MiniCPM:WSD的训练预算扩展、稳定阶段复⽤、继续训练与提前停⽌3.2 Long-Context Extension[Cohere] RNoPE:交替使⽤RoPE与NoPE,兼顾位置建模与长上下⽂检索4 Post-Training4.1 Post-Training Pipeline4.2 Reinforcement Learning[Moonshot AI] Kimi K1.5:提出partial rollout,通过复⽤未完成轨迹降低长CoT rollout开销[Moonshot AI] Kimi K2.5:reasoning-effort budget control 与 Agentic Generative Reward Model4.3 Multi-Teacher On-Policy Distillation[Microsoft Research] MiniLLM:基于student-generated samples的on-policy distillation4.4 Deployment-Aware Post-Training4.5 Draft Model Fine-Tuning5 Infrastructure5.1 Pre-Training5.1.1 KDA Kernel[FLA] Flash Linear Attention:KDA kernel与线性注意⼒⾼效实现5.1.2 Distributed Training[DeepSeek-AI] DeepSeek-V3 Technical Report:DualPipe与细粒度MoE communication–computation overlapMoE overlap的资源trade-off与cross-PP activation transfer5.1.3 Perfectly Balanced Expert-Parallel MoE Training5.1.4 Memory-Efficient Training[Moonshot AI] Mooncake:Mooncake Transfer Engine与cross-PP activation remote offload5.1.5 Multimodal Encoder Optimization5.2 RL5.2.1 Long-Context RL Infrastructure5.2.2 Sandbox Infrastructure5.3 Inference
Podzilla Summary coming soon
Sign up to get notified when the full AI-powered summary is ready.
Free forever for up to 3 podcasts. No credit card required.

153. 和曾鸣聊产业史观:残酷的真相、会消亡的公司、优秀≠卓越、“OAI、Anth大概率不是原生时代大赢家”

151. 17岁被2026年ICML收录论文的小少年:我bet开心!开心!开心!

150. 对英伟达研究副总裁刘洺堉的4小时访谈:Cosmos 3、世界模型、武术、黄仁勋影响我的,和你不需要击败所有对手

149. 亲历中美neo labs资本狂潮,和清华刘子鸣聊:AI for AI、机制可解释性和Max Tegmark
Free AI-powered recaps of 张小珺Jùn|商业访谈录 and your other favorite podcasts, delivered to your inbox.
Free forever for up to 3 podcasts. No credit card required.