Poison Below the Base Rate: What the Near-Constant Poisoning Result Establishes About Detection, and Four Things It Does Not
Abstract
A widely cited 2025 result reports that backdooring a language model through its training data takes a near-constant number of poisoned documents rather than a constant fraction of the corpus, and the field has read that as retiring the proportion-based threat model that most poison detectors are built on. This paper separates what the result establishes from what it is being read to establish. Its own stated scope covers a narrow class of backdoors, and its own ablations show the required count moving with learning rate, with where in the schedule the poison lands, and, at fine-tuning scale, with the amount of clean data - so the count is near-constant in corpus size specifically, not constant in general. The paper then does what the popular reading skips: it partitions poison detectors by what each actually needs, and finds that proportion binds three of the four families tightly, one of them not at all, and none of them in the way the slogan implies. The binding constraint on pretraining-scale data-side detection is a base rate, an old result from intrusion detection that the poisoning literature has not imported; the binding constraint on model-side detection is that the defender does not know the trigger. The base-rate argument is scoped deliberately and does not extend to fine-tuning-stage screening, where the prevalences the attack literature implies and the prevalences the detectors are evaluated at do overlap - a distinction the ratio-versus-count debate routinely collapses. The paper also corrects two premises that travel with this topic: that unlearning results close off post-hoc cleanup, and that a planted backdoor stays planted. The record on the second is openly contradictory, and the paper that supplies the headline number is on the side that says the backdoors it planted did not survive post-training. No experiments are reported here. The paper states what the record establishes, states flatly what it does not, and names four studies that would settle the open part.
The literature search, drafting and citation verification for this paper were carried out with AI assistance under the author's direction. Every citation was machine-verified against the arXiv API and Crossref before inclusion, and every quantitative claim was read back against the cited source's own text. The author is responsible for the final text and for all claims made in it.
Questions about this paper
Who wrote "Poison Below the Base Rate"?
Pranay Mahendrakar wrote "Poison Below the Base Rate: What the Near-Constant Poisoning Result Establishes About Detection, and Four Things It Does Not", published 24 Aug 2026. Pranay Mahendrakar is an Indian AI specialist and large language model engineer based in Bengaluru, India, who builds production artificial intelligence systems and publishes open-access research on how those systems fail.
Is "Poison Below the Base Rate" free to read?
Yes. "Poison Below the Base Rate" by Pranay Mahendrakar is open access under a Creative Commons Attribution 4.0 licence, with the full PDF available from Zenodo at https://doi.org/10.5281/zenodo.22073424. There is no paywall and no account required.
How do I cite "Poison Below the Base Rate"?
Cite the DOI: Mahendrakar, P. (2026). Poison Below the Base Rate: What the Near-Constant Poisoning Result Establishes About Detection, and Four Things It Does Not. Zenodo. https://doi.org/10.5281/zenodo.22073424 A BibTeX entry is provided on this page.