
David Manheim is head of methodology at AI Evaluation Consensus. He joins the podcast to discuss how AI evaluations can become more reliable, transparent, and useful for decisions. We cover common failures such as unclear reporting, training to the test, benchmark saturation, and models changing behavior when they know they are being tested. The conversation also examines real-world tests, biosecurity, persuasion, forecasting, human oversight, and why even “normal” AI progress could be disruptive.LINKS:David Manheim websiteCHAPTERS: (00:00) Episode Preview (01:04) Evaluation consensus project (07:01) Evaluation awareness challenges (12:28) Reporting capabilities clearly (19:38) Benchmarks beyond humans (29:52) Proxies and biosecurity (42:01) Persuasion and democracy (53:59) Forecasting with AI (01:08:44) Oversight and disruption (01:16:42) Supporting better evals PRODUCED BY: https://aipodcast.ing SOCIAL LINKS: Website: https://podcast.futureoflife.org Twitter (FLI): https://x.com/FLI_org Twitter (Gus): https://x.com/gusdocker LinkedIn: https://www.linkedin.com/company/future-of-life-institute/ YouTube: https://www.youtube.com/channel/UC-rCCy3FQ-GItDimSR9lhzw/ Apple: https://geo.itunes.apple.com/us/podcast/id1170991978 Spotify: https://open.spotify.com/show/2Op1WO3gwVwCrYHg4eoGyP
Podzilla Summary coming soon
Sign up to get notified when the full AI-powered summary is ready.
Free forever for up to 3 podcasts. No credit card required.

What Would a US-China AI Deal Look Like? (with Beatrice Fihn)

Why AI Hacking Is Becoming Hard to Control (with Benjamin Weinstein-Raun)

First 24 Hours of a Bioweapon Attack - Annie Jacobsen

How AI Shifts the Offense-Defense Balance in Biosecurity (with Paul-Enguerrand Fady)
Free AI-powered recaps of Future of Life Institute Podcast and your other favorite podcasts, delivered to your inbox.
Free forever for up to 3 podcasts. No credit card required.