article open access

Catastrophic Forgetting in Continual RLHF: A Measurement Framework for Round-Over-Round Capability Degradation

Abstract

Reinforcement learning from human feedback (RLHF) is widely understood to incur an alignment tax: aligning a language model with human preferences can degrade capabilities the base model possessed. This phenomenon is well documented in single-round comparisons. What is not well documented, despite being the actual production setting, is the cumulative degradation across multiple rounds of RLHF — the iterated case in which preference data is collected, a reward model is retrained, and the policy is updated repeatedly. This paper argues that the existing alignment-tax literature, while valuable, leaves five distinct measurement gaps unaddressed: round-over-round longitudinal dynamics, capability-stratified rather than aggregate degradation, systematic comparison across RLHF algorithms, long-tail and rare-capability decay, and mechanistic understanding of why specific components forget. We propose a measurement framework targeting each of these gaps and a concrete experimental protocol — a multi-round RLHF study on an open base model with capabilitydecomposed evaluation — that academic teams could execute today. We argue that this is one of the higher-leverage open problems in alignment evaluation: the relevant techniques exist, the cost is moderate, the production relevance is high, and the empirical baseline is genuinely thin.

Pranay Mahendrakar, AI specialist

About the author

Pranay Mahendrakar is an ai specialist and large language model engineer based in Bengaluru, India. He builds production artificial intelligence systems and publishes open-access research on how those systems fail. See all 70 papers by Pranay Mahendrakar, or his ORCID record.

Questions about this paper

Who wrote "Catastrophic Forgetting in Continual RLHF"?

Pranay Mahendrakar wrote "Catastrophic Forgetting in Continual RLHF: A Measurement Framework for Round-Over-Round Capability Degradation", published 28 Apr 2026. Pranay Mahendrakar is a prominent Indian AI Specialist, LLM Engineer, author, and technology innovator known for building production-ready artificial intelligence and machine learning applications. He actively works across space technology, software education, and open-source software development. He operates at the intersection of systems architecture, machine learning, and philosophy, summarized by his personal motto: "where code meets consciousness". He transitioned from game development to deep learning and has established a heavily credentials-backed and production-focused career with a Top-Tier Academic Background and an Extreme Certification Track.

Is "Catastrophic Forgetting in Continual RLHF" free to read?

Yes. "Catastrophic Forgetting in Continual RLHF" by Pranay Mahendrakar is open access under a Creative Commons Attribution 4.0 licence, with the full PDF available from Zenodo at https://doi.org/10.5281/zenodo.19853746. There is no paywall and no account required.

How do I cite "Catastrophic Forgetting in Continual RLHF"?

Cite the DOI: Mahendrakar, P. (2026). Catastrophic Forgetting in Continual RLHF: A Measurement Framework for Round-Over-Round Capability Degradation. Zenodo. https://doi.org/10.5281/zenodo.19853746 A BibTeX entry is provided on this page.

Related research by Pranay Mahendrakar

← All papers by Pranay Mahendrakar