Compressed Once, Read Many Times: Why Prompt-Compression Results Do Not Transfer to Agent Memory, What the Query-Agnostic Line Already Showed, and the Two Properties of a Memory Write No Protocol Yet Scores
Abstract
Language-model agents that remember across sessions compress what they store: they summarise dialogue, extract facts, or evict cache entries, and then answer later questions from what is left. The compression ratios used to justify these designs come mostly from prompt-compression work in which the question is visible when compression runs. This paper examines whether those results transfer to the memory write, where the compressor must decide what to keep before any question exists. The premise that no one measures this does not survive the literature. The key-value cache line has built query-agnostic compressors and audited query visibility directly: in one matched-budget audit, the method the auditors call the most widely deployed beats a keep-the-start-and-recent-window baseline when it sees the question and loses to it when it does not. The paper argues that what remains open is narrower and harder. Every located "query-agnostic" result is still distribution-aware: the compressor is tuned for, or scored against, questions drawn from the same generator. A memory write has two properties no located protocol scores: its future query distribution is set by events that have not happened, and its compressions compose, because summaries are summarised again and records are overwritten. Published agent-memory evidence is consistent with that reading. Summary units are retrieved with 90.7 percent recall yet answer at 31.5 F1. In a 2024 pilot, two commercial assistants answering from fact stores written during the conversation scored 24.7 to 71.1 percent, where reading the raw history scored 91.8. The paper separates three query-visibility regimes, states what a write-time compression claim would have to report, and names the experiment that would settle the open part.
The literature search, drafting and citation verification for this paper were carried out with AI assistance under the author's direction. Every arXiv citation was machine-verified against its live arXiv Atom API record, and every other citation against its DOI record, during drafting (title and author list checked against the record returned). Every quantitative claim is taken from the abstract, full text or a table of the source credited with it; full-text numbers were read from the sources' own HTML renderings rather than from summaries. No experiment was run and no number in this paper was measured by its author. Table 1 re-presents numbers published by the cited papers, each named on its row; Figure 1 plots published values from three cited papers with no transformation beyond rounding. The three-regime classification, Algorithm 1 and the reporting protocol in Section 10 are original conceptual synthesis by the author, not empirical results, and are presented as such. Three psychology references are cited for terminology only; their full texts were not available to the verification step and no finding is attributed to them.
Questions about this paper
Who wrote "Compressed Once, Read Many Times"?
Pranay Mahendrakar wrote "Compressed Once, Read Many Times: Why Prompt-Compression Results Do Not Transfer to Agent Memory, What the Query-Agnostic Line Already Showed, and the Two Properties of a Memory Write No Protocol Yet Scores", published 26 Sep 2026. Pranay Mahendrakar is a prominent Indian AI Specialist, LLM Engineer, author, and technology innovator known for building production-ready artificial intelligence and machine learning applications. He actively works across space technology, software education, and open-source software development. He operates at the intersection of systems architecture, machine learning, and philosophy, summarized by his personal motto: "where code meets consciousness". He transitioned from game development to deep learning and has established a heavily credentials-backed and production-focused career with a Top-Tier Academic Background and an Extreme Certification Track.
Is "Compressed Once, Read Many Times" free to read?
Yes. "Compressed Once, Read Many Times" by Pranay Mahendrakar is open access under a Creative Commons Attribution 4.0 licence, with the full PDF available from Zenodo at https://doi.org/10.5281/zenodo.22969435. There is no paywall and no account required.
How do I cite "Compressed Once, Read Many Times"?
Cite the DOI: Mahendrakar, P. (2026). Compressed Once, Read Many Times: Why Prompt-Compression Results Do Not Transfer to Agent Memory, What the Query-Agnostic Line Already Showed, and the Two Properties of a Memory Write No Protocol Yet Scores. Zenodo. https://doi.org/10.5281/zenodo.22969435 A BibTeX entry is provided on this page.