
Tl;dr I am currently worried about current alignment techniques + how they are applied to frontier models. This decomposes into two hypotheses: Alignment techniques are not working to address misalignment from RL. Alignment techniques are actively obscuring evidence about misalignment. I think we do not currently have enough (public) evidence to conclude whether either of these claims are true. However, if both of these were true that would imply that alignment techniques are net bad and...
Podzilla Summary coming soon
Sign up to get notified when the full AI-powered summary is ready.
Free forever for up to 3 podcasts. No credit card required.

"For Love of the Lightcone, Don’t Partisanize AI Safety" by DanB

"AI as orderly evacuation vs stampede" by Richard_Ngo

"Cooperation with AIs seems to be a low-hanging fruit for better evals" by Clément Dumas

"If Anyone Builds It, Everyone Dies: One Year Closer" by Eliezer Yudkowsky, So8res, Duncan Sabien (Inactive)
Free AI-powered recaps of LessWrong (Curated & Popular) and your other favorite podcasts, delivered to your inbox.
Free forever for up to 3 podcasts. No credit card required.