article open access

Three Guarantees Under One Word: What Importance-Based Selective Forgetting in Language Models Can and Cannot Promise

Abstract

Unlearning methods are asked to deliver a guarantee, and the word is used for three different ones: that a model's outputs no longer reveal the target under some stated class of queries, that the target is absent from the weights, and that no adversary within a stated budget can restore it. This paper separates the three, reads the published record as establishing that they come apart in practice, and asks which one each of the field's motivating use cases actually requires. Read through that separation, a literature that appears to disagree about whether unlearning works turns out to be reporting different guarantees under the same heading: benchmark forget-quality numbers are the first guarantee measured under the weakest access model, and the recovery results that appear to refute them are the third guarantee measured at budgets the benchmarks never applied. The paper then audits the specific premise that licenses importance-based and weight-attribution methods - that the target is concentrated in identifiable parameters, that those parameters can be found, and that changing them removes rather than reroutes - and finds published counter-evidence against each link, with the third link the weakest and the least addressed. On the demand side, only one of the three motivating use cases is well served by the guarantee the field is optimising, and for the right-to-erasure case an unprovability result suggests that the auditable deliverable is a documented procedure rather than a property of the weights. This paper is the author's synthesis and analysis of the published evidence. It states what the published record establishes, states flatly what it does not, proposes a three-part disclosure that would let a reader tell the guarantees apart, and names five studies that would settle the open part - including one matched-compute ablation that has been run for image classifiers and, as far as the author can determine, never for language models.

Pranay Mahendrakar, AI specialist

About the author

Pranay Mahendrakar is an ai specialist and large language model engineer based in Bengaluru, India. He builds production artificial intelligence systems and publishes open-access research on how those systems fail. See all 74 papers by Pranay Mahendrakar, or his ORCID record.

Questions about this paper

Who wrote "Three Guarantees Under One Word"?

Pranay Mahendrakar wrote "Three Guarantees Under One Word: What Importance-Based Selective Forgetting in Language Models Can and Cannot Promise", published 22 Aug 2026. Pranay Mahendrakar is an Indian AI specialist and LLM engineer based in Bengaluru, India. He is the Managing Director of SonyTech, Nodal Coordinator at IIRS-ISRO, and an instructor at Tutorials Point. His work covers large language models, natural language processing, computer vision and retrieval-augmented generation. He publishes open-access research papers and is the author of three books: Just AI With Pranay, Multiverse of AI and It's Me LLM.

Is "Three Guarantees Under One Word" free to read?

Yes. "Three Guarantees Under One Word" by Pranay Mahendrakar is open access under a Creative Commons Attribution 4.0 licence, with the full PDF available from Zenodo at https://doi.org/10.5281/zenodo.22054075. There is no paywall and no account required.

How do I cite "Three Guarantees Under One Word"?

Cite the DOI: Mahendrakar, P. (2026). Three Guarantees Under One Word: What Importance-Based Selective Forgetting in Language Models Can and Cannot Promise. Zenodo. https://doi.org/10.5281/zenodo.22054075 A BibTeX entry is provided on this page.

Related research by Pranay Mahendrakar

← All papers by Pranay Mahendrakar