Catastrophic Forgetting in Continual RLHF: A Measurement Framework for Round-Over-Round Capability Degradation
Reinforcement learning from human feedback (RLHF) is widely understood to incur an alignment tax: aligning a language model with human preferences can degrade capabilities the base model possessed. This phenomenon is well documented in…
DOI