Paper Review: Open Problems and Fundamental Limitations of Reinforcement Learning from Human Feedback
10 August 2023
A systematic survey of what's broken in RLHF — from reward hacking to evaluation gaps — and what techniques can fix, supplement, or replace it for safer AI alignment.