article open access

Does a Model Forget Differently When the Data Is Its Own? RL's Retention Advantage and Model Collapse Are Claims About the Same Loop, and No Study Has Measured Both

Abstract

Two literatures make claims about what happens when a language model trains on data that resembles its own distribution. One asks whether reinforcement learning forgets a model's prior capabilities less than supervised fine-tuning does, and answers yes: on-policy RL is implicitly biased toward solutions that stay close, in KL divergence, to the base policy, while supervised fine-tuning can converge to distributions arbitrarily far away (Shenfeld, Pari and Agrawal, 2025). The other asks whether a model trained recursively on its own generated outputs degrades across generations, and answers yes as well: absent a steady supply of fresh real data, the tails of the training distribution disappear, a failure named model collapse (Shumailov et al., 2023). A self-training loop -- reinforcement learning with verifiable rewards applied to a policy that generates its own training data and is scored by a verifier it does not otherwise consult -- is an instance of both objects at once, and no study surveyed here measures both retention and collapse on the same system. This paper lays out what each literature actually measures, its reference point (a held-out prior task versus the original data distribution) and its timescale (one fine-tuning run versus many training generations); argues these are not obviously the same failure mode despite a shared vocabulary of KL divergence, entropy and distributional narrowing; and assembles a small set of 2025-2026 results -- on correct-set turnover inside RLVR itself, on RLVR's failure to expand a model's pass-at-large-k ceiling beyond its own base model, and on verifier reliability degrading in exactly the low-resource regime where synthetic data accumulates fastest -- that make the reconciliation harder than either literature admits when read on its own. The paper does not resolve which failure mode dominates a given self-training loop. It specifies the joint measurement nobody has run, and states plainly why each literature's current instruments cannot answer the other literature's question.

The literature search, drafting and citation verification for this paper were carried out with AI assistance under the author's direction. Every citation was machine-verified against its live arXiv Atom API record during drafting (title and author list checked against the record returned), and every quantitative claim in this paper is taken from the abstract or a directly quoted headline result of the source credited with it. No experiment was run and no number in this paper was measured or recomputed by its author; Table 1 re-presents numbers published by the cited papers, each named on its row, and Figure 1 plots four of those numbers directly with no transformation beyond unit labeling. Algorithm 1 is original conceptual synthesis by the author, not an empirical result and not reproduced from any single cited source; it is presented as such.

Pranay Mahendrakar, AI specialist

About the author

Pranay Mahendrakar is an ai specialist and large language model engineer based in Bengaluru, India. He builds production artificial intelligence systems and publishes open-access research on how those systems fail. See all 66 papers by Pranay Mahendrakar, or his ORCID record.

Questions about this paper

Who wrote "Does a Model Forget Differently When the Data Is Its Own? RL's Retention Advantage and Model Collapse Are Claims About the Same Loop, and No Study Has Measured Both"?

Pranay Mahendrakar wrote "Does a Model Forget Differently When the Data Is Its Own? RL's Retention Advantage and Model Collapse Are Claims About the Same Loop, and No Study Has Measured Both", published 25 Sep 2026. Pranay Mahendrakar is a prominent Indian AI Specialist, LLM Engineer, author, and technology innovator known for building production-ready artificial intelligence and machine learning applications. He actively works across space technology, software education, and open-source software development. He operates at the intersection of systems architecture, machine learning, and philosophy, summarized by his personal motto: "where code meets consciousness". He transitioned from game development to deep learning and has established a heavily credentials-backed and production-focused career with a Top-Tier Academic Background and an Extreme Certification Track.

Is "Does a Model Forget Differently When the Data Is Its Own? RL's Retention Advantage and Model Collapse Are Claims About the Same Loop, and No Study Has Measured Both" free to read?

Yes. "Does a Model Forget Differently When the Data Is Its Own? RL's Retention Advantage and Model Collapse Are Claims About the Same Loop, and No Study Has Measured Both" by Pranay Mahendrakar is open access under a Creative Commons Attribution 4.0 licence, with the full PDF available from Zenodo at https://doi.org/10.5281/zenodo.22945776. There is no paywall and no account required.

How do I cite "Does a Model Forget Differently When the Data Is Its Own? RL's Retention Advantage and Model Collapse Are Claims About the Same Loop, and No Study Has Measured Both"?

Cite the DOI: Mahendrakar, P. (2026). Does a Model Forget Differently When the Data Is Its Own? RL's Retention Advantage and Model Collapse Are Claims About the Same Loop, and No Study Has Measured Both. Zenodo. https://doi.org/10.5281/zenodo.22945776 A BibTeX entry is provided on this page.

Related research by Pranay Mahendrakar

← All papers by Pranay Mahendrakar