article open access

Three Guarantees Under One Word: What Importance-Based Selective Forgetting in Language Models Can and Cannot Promise

Abstract

Unlearning methods are asked to deliver a guarantee, and the word is used for three different ones: that a model's outputs no longer reveal the target under some stated class of queries, that the target is absent from the weights, and that no adversary within a stated budget can restore it. This paper separates the three, reads the published record as establishing that they come apart in practice, and asks which one each of the field's motivating use cases actually requires. Read through that separation, a literature that appears to disagree about whether unlearning works turns out to be reporting different guarantees under the same heading: benchmark forget-quality numbers are the first guarantee measured under the weakest access model, and the recovery results that appear to refute them are the third guarantee measured at budgets the benchmarks never applied. The paper then audits the specific premise that licenses importance-based and weight-attribution methods - that the target is concentrated in identifiable parameters, that those parameters can be found, and that changing them removes rather than reroutes - and finds published counter-evidence against each link, with the third link the weakest and the least addressed. On the demand side, only one of the three motivating use cases is well served by the guarantee the field is optimising, and for the right-to-erasure case an unprovability result suggests that the auditable deliverable is a documented procedure rather than a property of the weights. This paper reports no experiments and no measurements of its own. It states what the published record establishes, states flatly what it does not, proposes a three-part disclosure that would let a reader tell the guarantees apart, and names five studies that would settle the open part - including one matched-compute ablation that has been run for image classifiers and, as far as the author can determine, never for language models.

The literature search, drafting and citation verification for this paper were carried out with AI assistance under the author's direction. Every citation was machine-verified against the arXiv API and Crossref before inclusion, and every quantitative claim was read back against the cited source's own text. The author is responsible for the final text and for all claims made in it.

Pranay Mahendrakar, AI specialist

About the author

Pranay Mahendrakar is an ai specialist and large language model engineer based in Bengaluru, India. He builds production artificial intelligence systems and publishes open-access research on how those systems fail. See all 19 papers by Pranay Mahendrakar, or his ORCID record.

Questions about this paper

Who wrote "Three Guarantees Under One Word"?

Pranay Mahendrakar wrote "Three Guarantees Under One Word: What Importance-Based Selective Forgetting in Language Models Can and Cannot Promise", published 22 Aug 2026. Pranay Mahendrakar is an Indian AI specialist and large language model engineer based in Bengaluru, India, who builds production artificial intelligence systems and publishes open-access research on how those systems fail.

Is "Three Guarantees Under One Word" free to read?

Yes. "Three Guarantees Under One Word" by Pranay Mahendrakar is open access under a Creative Commons Attribution 4.0 licence, with the full PDF available from Zenodo at https://doi.org/10.5281/zenodo.22054075. There is no paywall and no account required.

How do I cite "Three Guarantees Under One Word"?

Cite the DOI: Mahendrakar, P. (2026). Three Guarantees Under One Word: What Importance-Based Selective Forgetting in Language Models Can and Cannot Promise. Zenodo. https://doi.org/10.5281/zenodo.22054075 A BibTeX entry is provided on this page.

Related research by Pranay Mahendrakar

← All papers by Pranay Mahendrakar