A Safety Memory Is a Declassification Channel: Why Origin Binding, Taint Tracking and Memory Isolation Do Not Secure the Records an Agent Writes About Attacks
Abstract
Two lines of work on language-model agents have converged on the same component from opposite directions. One proposes that an agent keep a persistent safety memory, so that an attack seen once, a request refused once or a source found untrustworthy once continues to inform decisions in later sessions. The other reports that persistent memory is an exploitable channel, with published attacks that write records through query-only interaction, through fragments individually too innocuous to filter, and through the agent's own reflection step. This paper argues that the coincidence is structural rather than incidental. A safety memory is defined by a property that ordinary agent memory does not have: its records exist because the agent observed something untrusted, and they are worth keeping only if they later carry authority over a consequential decision. In the vocabulary of information-flow control that the agent-security literature has already adopted, that is a declassification, specifically an integrity endorsement, and it is the one operation the three defense families now in print are built to forbid. Write-time origin binding, proved necessary and, with corroboration-gated elevation, sufficient against laundering for ordinary memory, either denies a safety record the authority that makes it useful or supplies the elevation an attacker needs. Monotone taint tracking over-blocks or strands the record by construction, which its own proponents say. Memory isolation prevents the write that the safety memory exists to perform. Three of the paper's claims cut against the premise it started from. The headline poisoning success rates are not one quantity: storage and execution dissociate, and the early numbers were measured in memory pools that later attack papers describe as unrealistically empty. Four independent 2026 results report the same degradation with no attacker present, which makes attack surface the wrong frame for at least part of the phenomenon. And the claim that nobody separates the trust level of a remembered threat from the trust level of the environment that produced it is false: one paper does exactly that, and the interesting problem is what its solution costs a safety memory in particular. What is not known is the size of any of it. No located source measures how often a safety memory fires on a record it should not trust, and no benchmark evaluates a memory-based guardrail under an attacker who is targeting the guardrail's own memory.
The literature search, drafting and citation verification for this paper were carried out with AI assistance under the author's direction. Every citation was machine-verified against the arXiv API and Crossref before inclusion, and every quantitative claim was read back against the cited source's own abstract. The author is responsible for the final text and for all claims made in it.
Questions about this paper
Who wrote "A Safety Memory Is a Declassification Channel"?
Pranay Mahendrakar wrote "A Safety Memory Is a Declassification Channel: Why Origin Binding, Taint Tracking and Memory Isolation Do Not Secure the Records an Agent Writes About Attacks", published 3 Sep 2026. Pranay Mahendrakar is a prominent Indian AI Specialist, LLM Engineer, author, and technology innovator known for building production-ready artificial intelligence and machine learning applications. He actively works across space technology, software education, and open-source software development. He operates at the intersection of systems architecture, machine learning, and philosophy, summarized by his personal motto: "where code meets consciousness". He transitioned from game development to deep learning and has established a heavily credentials-backed and production-focused career with a Top-Tier Academic Background and an Extreme Certification Track.
Is "A Safety Memory Is a Declassification Channel" free to read?
Yes. "A Safety Memory Is a Declassification Channel" by Pranay Mahendrakar is open access under a Creative Commons Attribution 4.0 licence, with the full PDF available from Zenodo at https://doi.org/10.5281/zenodo.23022350. There is no paywall and no account required.
How do I cite "A Safety Memory Is a Declassification Channel"?
Cite the DOI: Mahendrakar, P. (2026). A Safety Memory Is a Declassification Channel: Why Origin Binding, Taint Tracking and Memory Isolation Do Not Secure the Records an Agent Writes About Attacks. Zenodo. https://doi.org/10.5281/zenodo.23022350 A BibTeX entry is provided on this page.