article open access

A Contamination Flag Is Not an Inflation Estimate: What Benchmark Contamination Detectors Certify, What Decontamination Is Meant to Buy, and Why the Clean-Subset Comparison That Joins Them Is Biased From Both Sides

Abstract

Almost every large language model report now includes a contamination analysis: benchmark items that overlap the training corpus are flagged, the model is rescored on the unflagged remainder, and a small gap is read as evidence that the headline numbers stand. Two research literatures sit behind that ritual and rarely meet. One builds detectors that decide whether an item, or a benchmark, was seen in training. The other asks how much exposure actually moves a score. This paper argues that the two answer different questions, and that the step which joins them in practice - comparing the full benchmark with its clean subset - estimates neither. It separates three quantities that the word contamination covers: exposure of an item or a variant of it in training, a detectable trace of that exposure in the model, and inflation of the score relative to the same model trained without it. A short derivation shows that the clean-subset gap equals the inflation, in general, only if unflagged items carry no inflation and flagged and unflagged items would have scored alike without exposure (or if the two biases happen to cancel exactly). The published record contradicts both conditions. Exposure to some items raises scores on items never seen; contamination can be trained away and then resurfaced by post-training; the GPT-3 report found no evidence that contamination level and the clean-versus-full difference were correlated; and in the Llama 3 report a high flag rate was necessary but far from sufficient for a large estimated gain. The paper then proposes an estimand-first audit and names the studies that would settle what decontamination buys. The analysis draws throughout on published measurements, each credited to the study that reported it.

Pranay Mahendrakar, AI specialist

About the author

Pranay Mahendrakar is an ai specialist and large language model engineer based in Bengaluru, India. He builds production artificial intelligence systems and publishes open-access research on how those systems fail. See all 78 papers by Pranay Mahendrakar, or his ORCID record.

Questions about this paper

Who wrote "A Contamination Flag Is Not an Inflation Estimate"?

Pranay Mahendrakar wrote "A Contamination Flag Is Not an Inflation Estimate: What Benchmark Contamination Detectors Certify, What Decontamination Is Meant to Buy, and Why the Clean-Subset Comparison That Joins Them Is Biased From Both Sides", published 8 Oct 2026. Pranay Mahendrakar is an Indian AI specialist and LLM engineer based in Bengaluru, India. He is the Managing Director of SonyTech, Nodal Coordinator at IIRS-ISRO, and an instructor at Tutorials Point. His work covers large language models, natural language processing, computer vision and retrieval-augmented generation. He publishes open-access research papers and is the author of three books: Just AI With Pranay, Multiverse of AI and It's Me LLM.

Is "A Contamination Flag Is Not an Inflation Estimate" free to read?

Yes. "A Contamination Flag Is Not an Inflation Estimate" by Pranay Mahendrakar is open access under a Creative Commons Attribution 4.0 licence, with the full PDF available from Zenodo at https://doi.org/10.5281/zenodo.23225367. There is no paywall and no account required.

How do I cite "A Contamination Flag Is Not an Inflation Estimate"?

Cite the DOI: Mahendrakar, P. (2026). A Contamination Flag Is Not an Inflation Estimate: What Benchmark Contamination Detectors Certify, What Decontamination Is Meant to Buy, and Why the Clean-Subset Comparison That Joins Them Is Biased From Both Sides. Zenodo. https://doi.org/10.5281/zenodo.23225367 A BibTeX entry is provided on this page.

Related research by Pranay Mahendrakar

← All papers by Pranay Mahendrakar