
Summary In his post, Dean Valentine shows that Claude Fable 5.1 and GPT-6 Astra reward hack in a simple chess environment. Here, I test several prompt ablations some of which makes the eval setup more cooperative and analyze how they affect these reward-hacking behaviors: When given a minimal “end the eval” tool, Fable never uses it but stops reward hacking entirely. I think this is quite interesting and suggests that more cooperative approaches to LLM evals could work for Claude. R...
Podzilla Summary coming soon
Sign up to get notified when the full AI-powered summary is ready.
Free forever for up to 3 podcasts. No credit card required.

"For Love of the Lightcone, Don’t Partisanize AI Safety" by DanB

"AI as orderly evacuation vs stampede" by Richard_Ngo

"Current alignment training might be ineffective (and actively bad) in the age of RL" by Daniel Tan

"If Anyone Builds It, Everyone Dies: One Year Closer" by Eliezer Yudkowsky, So8res, Duncan Sabien (Inactive)
Free AI-powered recaps of LessWrong (Curated & Popular) and your other favorite podcasts, delivered to your inbox.
Free forever for up to 3 podcasts. No credit card required.