
What happens when AI passes our tests faster than we create new ones? Adam Khoja and Richard Ren helped build Humanity’s Last Exam. They return to explain what the disappearing benchmarks tell us about the future, and why smarter AI doesn’t automatically mean safer AI.Adam Khoja and Richard Ren are research engineers at the Center for AI Safety (CAIS).We go through the benchmarks they’ve built — Humanity’s Last Exam, the Remote Labor Index, MASK — and the safety-washing problem, where AI companies pass off raw capability gains as safety progress. Then we get to their forecasts: when AI outperforms research mathematicians, whether there will be a billion general-purpose robots by 2035, and why they put 80% odds that historians will judge we faced at least a 33% chance of catastrophe.Watch on YouTube: https://www.youtube.com/watch?v=gBPzgJkT9e8Timestamps00:00:00 — Cold Open00:00:42 — Adam Khoja and Richard Ren Return00:04:57 — Humanity’s Last Exam00:14:47 — AI Outrunning Its Benchmarks00:20:24 — The Remote Labor Index00:27:59 — AI and the Next Human Job00:30:08 — MASK: Catching AI in a Lie00:38:06 — Safety Washing: Capabilities Passed Off as Safety00:44:30 — Ethical Knowledge vs. Ethical Behavior00:48:58 — Forecasting AI from 2026 to 206000:53:53 — AI vs. Research Mathematicians00:55:53 — Putting a Number on P(Doom)01:04:30 — A Billion Robots by 203501:08:09 — Dyson Swarms and the Physical Singularity01:09:56 — Risk, Precision, and Taking ActionLinksAdam Khoja's first Doom Debates appearance — https://www.youtube.com/watch?v=QqESBXuo6EIRichard Ren's first Doom Debates appearance — https://www.youtube.com/watch?v=1glFImnyp6oCollision — Richard Ren's Substack — https://richardren.substack.com/"Predictions on AI (2026–2060)" — Adam and Richard's forecasts, written December 2025, with resolution status — https://richardren.substack.com/p/predictions-on-ai-20262060Manifold Markets — where Adam built his forecasting track record — https://manifold.markets/Center for AI Safety — https://safe.ai/Center for AI Safety — careers — https://safe.ai/careersStatement on AI Risk (Center for AI Safety, May 2023) — https://safe.ai/work/statement-on-ai-riskHumanity's Last Exam — the 2,500-question closed-book exam at the frontier of human knowledge — https://agi.safe.ai/"Humanity's Last Exam" (paper) — https://arxiv.org/abs/2501.14249Remote Labor Index — real Upwork projects, measuring what fraction of remote work AI can actually finish — https://www.remotelabor.ai/"Remote Labor Index: Measuring AI Automation of Remote Work" (paper) — https://arxiv.org/abs/2510.26787MASK — the honesty benchmark: does a model contradict its own stated beliefs under pressure? — https://www.mask-benchmark.ai/"The MASK Benchmark: Disentangling Honesty From Accuracy in AI Systems" (paper) — https://arxiv.org/abs/2503.03750"Safetywashing: Do AI Safety Benchmarks Actually Measure Safety Progress?" — Richard Ren et al. (NeurIPS 2024) — the paper behind the safety-washing segment — https://arxiv.org/abs/2407.21792WMDP — the weaponization benchmark for bio, cyber, and chem — https://www.wmdp.ai/"The WMDP Benchmark: Measuring and Reducing Malicious Use With Unlearning" (paper) — https://arxiv.org/abs/2403.03218"Measuring Massive Multitask Language Understanding" (MMLU) — Dan Hendrycks et al., 2020 — https://arxiv.org/abs/2009.03300METR, "Measuring AI Ability to Complete Long Tasks" — the time-horizon graph — https://metr.org/blog/2025-03-19-measuring-ai-ability-to-complete-long-tasks/</p
Podzilla Summary coming soon
Sign up to get notified when the full AI-powered summary is ready.
Free forever for up to 3 podcasts. No credit card required.

The Week in AI that Changed the World, With Robert Wright | Nonzero × Doom Debates

He May Have Found AI's FEELINGS — Richard Ren, Center for AI Safety Researcher

USA and China Will Each Be BETRAYED By Their Own AIs — Adam Khoja, Center for AI Safety

Sam Altman Is Gaslighting About AI Risk After His Own AI Just Went Rogue
Free AI-powered recaps of Doom Debates! and your other favorite podcasts, delivered to your inbox.
Free forever for up to 3 podcasts. No credit card required.