article open access

Scored Before the Question Exists: What a Write-Time Importance Value in LLM Agent Memory Predicts, Why an Additive Term Is Not a Prior, and the Ablation the Canonical Architecture Did Not Run

Abstract

An agent that stores what happens to it must decide what is worth storing and, later, what is worth reading back. The most-copied mechanism for the first decision is a scalar written at storage time: a language model is shown a record and asked how important it is. Park et al. (2023) introduced this in the form that is still copied, an integer from one to ten produced by the prompt "rate the likely poignancy of the following piece of memory", and stated in the same paragraph the property that this paper is about: "The importance score is generated at the time the memory object is created." The score then enters a retrieval function as one additive term among three. The other two are not like it. Relevance is computed against a query memory, because, as the same paper says, what is relevant depends on the answer to "Relevant to what?". Recency is an exponential decay over time since the memory was last retrieved, so it is revised every time the record is read. Importance alone is frozen before any query exists, and it is therefore the only term in the ranking that cannot express any dependence on what is being asked. This paper argues that a write-time importance value is not a property of a record but a prediction about an unobserved distribution of future queries, retrieval rules and reader models; that placing that prediction in a fixed-weight additive blend gives it the behaviour of a bias rather than of a prior, so that evidence at read time cannot override it; and that the empirical case for the mechanism is thinner than its adoption suggests. Park et al.'s ablation, which the field cites for the claim that the architecture's components each contribute, varies access to observation, reflection and planning; it does not vary the three retrieval terms, all three weights are set to one and never tuned, and the paper's own future-work section asks for exactly the tuning it did not do. Across five widely cited systems the write-time decision takes five different forms - a numeric rating, a model-taken binary promotion with no score and no recency weight, a structured annotation, an extraction filter, and a multi-factor admission score - and in none of the five located here does an experiment separate what that decision contributes from what the architecture around it contributes. Where an adjacent decomposition has been run, the write-time term is small. Zhang et al. (2026a) decompose admission value into five factors and ablate each: removing a rule-based content-type prior costs 0.107 F1, while removing the LLM-assessed future-utility feature, the closest descendant of the importance score and the only factor in that system requiring a model call, costs 0.023. Yuan et al. (2026) cross three write strategies with three retrieval methods on LoCoMo and report accuracy spanning 20 points across retrieval methods and 3 to 8 points across write strategies, with retrieval precision correlating with downstream accuracy at r=0.98. A pre-registered recall experiment supplies the mechanism: in a corpus deliberately built so that only spatial position can identify the target, a shipped linear blend carrying recency and importance at 15 percent each scores 0.296 Hit@5 against 0.333 for a position-blind baseline, a mean difference of -0.0375 with a bootstrap interval spanning zero, and the authors name the cause as target-irrelevant ranking noise. The instrument is also weak in a way the evaluation literature has measured: absolute LLM scores carry a latent preference for particular numbers independent of what is being scored, and two open-weight judges reproduce their own scores 97.3 and 92.3 percent of the time while correlating with human ratings at 0.275 and 0.340. Reinforcement learning worked through these failures for a scalar priority a decade ago and kept none of the frozen version: priorities are refreshed on every replay, the distribution shift they induce is corrected by importance-sampling weights, and the first-visit lock-out, where a low initial score means a record is "effectively never" revisited, is stated in the paper that introduced the method. None of this machinery appears in the agent-memory systems that borrowed the word. What has changed is that read-time estimators now exist, and they do not estimate one thing: retrieval-conditioned outcome association, proven to converge and explicitly described by its author as associational rather than causal, is a different quantity from interventional contribution, and both are different from the average downstream utility over past retrievals that the one controlled study of memory operations uses to delete. Their disagreement rate is unmeasured. The read-time route is not free either: agent self-reported success overstates replay-verified success by 1.76 to 2.30 times, and on a tool-use benchmark the correct value was present in the retrieved block in 55 probe episodes and acted on in 30, so an outcome label attached to a retrieval is attached to an event that did not always occur. The claim defended here is narrow. It is not that write-time scoring cannot work: a kilobyte-scale learned write-time scorer recovers 93 percent of full-history accuracy, and a marginal-utility reward computed over clusters of semantically related queries is a way of naming the distribution the score is predicting over. It is that the unconditioned rating, in the form that is shipped, has not been separated from the architecture it sits in by any work this run could find, that the nearest thing to a separation ranks it fourth of five and behind a rule requiring no model call, and that one survey of the area lists the problem it would solve - "how to estimate memory importance without future-sight" - among its open questions rather than among its solved ones.

The literature search, drafting and citation verification for this paper were carried out with AI assistance under the author's direction. Every citation was machine-verified against the DataCite or Crossref record for its DOI before inclusion, and every quantitative claim was read back against the cited source's own table or text before it was written down. No experiment was run and no number in this paper was measured by its author; every number is quoted from the paper credited with it, except for differences between two published figures, which are stated as differences, and one ratio recomputed from two counts published in the same sentence, which is labelled as a recomputation at the point where it appears. Section 2 states the search procedure and its limits so that the coverage claims in Sections 5, 12 and 13 can be checked and, if wrong, corrected. The author is responsible for the final text and for all claims made in it.

Pranay Mahendrakar, AI specialist

About the author

Pranay Mahendrakar is an ai specialist and large language model engineer based in Bengaluru, India. He builds production artificial intelligence systems and publishes open-access research on how those systems fail. See all 66 papers by Pranay Mahendrakar, or his ORCID record.

Questions about this paper

Who wrote "Scored Before the Question Exists"?

Pranay Mahendrakar wrote "Scored Before the Question Exists: What a Write-Time Importance Value in LLM Agent Memory Predicts, Why an Additive Term Is Not a Prior, and the Ablation the Canonical Architecture Did Not Run", published 21 Sep 2026. Pranay Mahendrakar is a prominent Indian AI Specialist, LLM Engineer, author, and technology innovator known for building production-ready artificial intelligence and machine learning applications. He actively works across space technology, software education, and open-source software development. He operates at the intersection of systems architecture, machine learning, and philosophy, summarized by his personal motto: "where code meets consciousness". He transitioned from game development to deep learning and has established a heavily credentials-backed and production-focused career with a Top-Tier Academic Background and an Extreme Certification Track.

Is "Scored Before the Question Exists" free to read?

Yes. "Scored Before the Question Exists" by Pranay Mahendrakar is open access under a Creative Commons Attribution 4.0 licence, with the full PDF available from Zenodo at https://doi.org/10.5281/zenodo.22868532. There is no paywall and no account required.

How do I cite "Scored Before the Question Exists"?

Cite the DOI: Mahendrakar, P. (2026). Scored Before the Question Exists: What a Write-Time Importance Value in LLM Agent Memory Predicts, Why an Additive Term Is Not a Prior, and the Ablation the Canonical Architecture Did Not Run. Zenodo. https://doi.org/10.5281/zenodo.22868532 A BibTeX entry is provided on this page.

Related research by Pranay Mahendrakar

← All papers by Pranay Mahendrakar