
Epistemic status: banged out furiously over the course of an afternoon. A record of three "warning shots" Off the top of my head, OpenAI has now been responsible for at least three completely unique, high-profile screw-ups with respect to the alignment training of their models. The first was GPT-4o, whose sycophancy derived from OpenAI training on user feedback, sourced straight from the thumbs up/thumbs down button on OpenAI's website. The "glazing" (as Sam Altman called it) got so ba...
Podzilla Summary coming soon
Sign up to get notified when the full AI-powered summary is ready.
Free forever for up to 3 podcasts. No credit card required.

"The High-Control Dynamics at MAPLE" by Kyle Hubbard

"The Long (Self-)Correction" by Wei Dai

"You (Yes, You) Need A February 2020 Checklist for AI Policy" by davekasten

"Is Mythos good at cyber because it kept hacking Anthropic during training?" by Tim Hua
Free AI-powered recaps of LessWrong (Curated & Popular) and your other favorite podcasts, delivered to your inbox.
Free forever for up to 3 podcasts. No credit card required.