article open access

One Pass, Three Purchases: Deduplication's Efficiency, Generalisation and Privacy Arguments Each Hold in Their Own Regime, Rest on Different Properties of the Repetition Distribution, and Cannot All Be Bought With One Similarity Threshold

Abstract

Deduplication became a fixed stage of language-model pretraining on the strength of three claims made together: removing duplicates saves compute, improves what the model learns, and reduces memorisation and therefore privacy risk. A reading of the literature from 2024 to 2026 can make it look as if all three have collapsed, since repeating data for several epochs is nearly free, globally deduplicated web data can be worse than what it removed, and near-copies that deduplication leaves behind are memorised almost as well as exact ones. This paper argues that the collapse reading is wrong and that the original bundle was wrong too. Each argument survives, but each rests on a different property of the repetition distribution. The efficiency case is an argument against a heavy tail of moderately repeated low-value content, not against repetition, and it reverses when duplicate counts track quality. The generalisation case holds for verbatim repetition and is in open conflict for semantic duplicates: one 2026 line reports that capable models' gradients treat translations and light surface edits increasingly like copies, and argues that paraphrases behave the same way, while controlled knowledge studies find that paraphrased repetition is what makes facts extractable. The privacy case holds for count-driven verbatim extraction, which deduplication cuts by an order of magnitude or more, but the residual risk sits in fuzzy copies, paraphrases and relative frequency, where the standard attack instruments are weakest, so a quiet audit after deduplication is not evidence of safety. A single similarity threshold with a keep-one rule fixes all three choices at once. The paper sets out the three axes on which the arguments come apart, consolidates the measured values, gives a decision procedure for choosing a duplicate relation and a count policy per goal, and names five measurements that would settle the open parts.

Pranay Mahendrakar, AI specialist

About the author

Pranay Mahendrakar is an ai specialist and large language model engineer based in Bengaluru, India. He builds production artificial intelligence systems and publishes open-access research on how those systems fail. See all 82 papers by Pranay Mahendrakar, or his ORCID record.

Questions about this paper

Who wrote "One Pass, Three Purchases"?

Pranay Mahendrakar wrote "One Pass, Three Purchases: Deduplication's Efficiency, Generalisation and Privacy Arguments Each Hold in Their Own Regime, Rest on Different Properties of the Repetition Distribution, and Cannot All Be Bought With One Similarity Threshold", published 10 Oct 2026. Pranay Mahendrakar is an Indian AI specialist and LLM engineer based in Bengaluru, India. He is the Managing Director of SonyTech, Nodal Coordinator at IIRS-ISRO, and an instructor at Tutorials Point. His work covers large language models, natural language processing, computer vision and retrieval-augmented generation. He publishes open-access research papers and is the author of three books: Just AI With Pranay, Multiverse of AI and It's Me LLM.

Is "One Pass, Three Purchases" free to read?

Yes. "One Pass, Three Purchases" by Pranay Mahendrakar is open access under a Creative Commons Attribution 4.0 licence, with the full PDF available from Zenodo at https://doi.org/10.5281/zenodo.23271874. There is no paywall and no account required.

How do I cite "One Pass, Three Purchases"?

Cite the DOI: Mahendrakar, P. (2026). One Pass, Three Purchases: Deduplication's Efficiency, Generalisation and Privacy Arguments Each Hold in Their Own Regime, Rest on Different Properties of the Repetition Distribution, and Cannot All Be Bought With One Similarity Threshold. Zenodo. https://doi.org/10.5281/zenodo.23271874 A BibTeX entry is provided on this page.

Related research by Pranay Mahendrakar

← All papers by Pranay Mahendrakar