
Neel and his team are trying to do something phenomenally difficult: understand an intelligence that didn't come with a manual. Together, they explore the cutting-edge "neuroscience" of artificial intelligence—revealing the surprising, elegant structures being discovered inside these networks (like spare autoencoders), the inherent limits of looking under the hood, and why interpretability is absolutely essential if we are to build safe, aligned and trustworthy AI as we move towards AGI. Learn more about this area of research via https://deepmind.google/ Timecodes 00:00 Introduction 02:41 Motivation for interpretability research 04:01 Mechanistic interpretability 08:14 Chain of thought monitoring 18:14 Interpretability techniques 35:00 Auditing models for safety 48:53 What comes next for interpretability Please leave us a review on Spotify or Apple Podcasts if you enjoyed this episode. We always want to hear from our audience whether that's in the form of feedback, new idea or a guest recommendation! Hosted by Simplecast, an AdsWizz company. See pcm.adswizz.com for information about our collection and use of personal data for advertising.
Podzilla Summary coming soon
Sign up to get notified when the full AI-powered summary is ready.
Free forever for up to 3 podcasts. No credit card required.

When millions of AI agents meet

10 Years of AlphaGo: The Turning Point for AI | Thore Graepel & Pushmeet Kohli

The Future of Intelligence with Demis Hassabis (Co-founder and CEO of DeepMind)

The Arrival of AGI with Shane Legg (co-founder of DeepMind)
Free AI-powered recaps of Google DeepMind: The Podcast and your other favorite podcasts, delivered to your inbox.
Free forever for up to 3 podcasts. No credit card required.