26 / papers

Reasoning & Memory

How models hold context, carry state, and compose steps beyond next-token prediction. 26 open-access papers by Pranay Mahendrakar, each with a permanent DOI.

article

Pooling Experience Spends Independence Twice: Shared Memory in LLM Agent Teams Correlates Errors and Opens a Common-Mode Channel, Why Little of the First Was Left to Spend, and the Order, Gate and Readout That May Decide Whether Pooling Pays

Teams of language-model agents increasingly read and write a common memory: a message pool, a blackboard, a bank of distilled experience, a knowledge base served to several frameworks at once. The case for pooling is efficiency, since no…

article

Errors That Should Compound, and When They Do: The Product Rule Assumes Independent Steps, Conformal Guarantees Assume Exchangeable Instances, and the Unit of Independence Decides What Each Can Claim

Two bodies of work quantify uncertainty in multi-step language-model reasoning, and both rest on an independence assumption placed at some unit. Step-level error models multiply per-step reliabilities, which, at an illustrative 95%…

article

Perspectives Without Independence: Multi-Agent and Multi-Persona Reasoning Under Compute-Normalised Comparison, Why the Gains That Survive Are Not the Ones Diversity Predicts, and the Controls That Would Tell Them Apart

Multi-agent debate, multi-persona prompting and related schemes are usually justified by diversity of viewpoint: several perspectives err differently, so their combination is more reliable than any one of them. That justification is a…

article

Self-Model or Self-Simulation? A Machine Self-Awareness Index Averages Sub-Scores With No Common Referent and No Fixed Sign, Why Persistent Identity, Goal Stability and Memory Continuity Are Not Evidence of Self-Access, and the Validity Tests Any Composite Would Have to Pass

Some proposals to quantify machine self-awareness combine several sub-scores - persistent identity, goal stability, cross-session memory continuity, contradiction detection, uncertainty awareness, self-prediction and introspective access…

article

Compressed Once, Read Many Times: Why Prompt-Compression Results Do Not Transfer to Agent Memory, What the Query-Agnostic Line Already Showed, and the Two Properties of a Memory Write No Protocol Yet Scores

Language-model agents that remember across sessions compress what they store: they summarise dialogue, extract facts, or evict cache entries, and then answer later questions from what is left. The compression ratios used to justify these…

article

Three Things Called Budget Awareness: Observability, Forecasting and Allocation in LLM Agents, Why Every Published Allocation Gain Is Keyed to a Signal Measured After the Fact, and the Run-to-Run Variance No Forecast Is Scored Against

Two 2024-2026 literatures make claims about resource use in language-model agents that look incompatible. One reports that allocating test-time compute according to problem difficulty beats spending it uniformly, by margins up to a factor…

article

Scored Before the Question Exists: What a Write-Time Importance Value in LLM Agent Memory Predicts, Why an Additive Term Is Not a Prior, and the Ablation the Canonical Architecture Did Not Run

An agent that stores what happens to it must decide what is worth storing and, later, what is worth reading back. The most-copied mechanism for the first decision is a scalar written at storage time: a language model is shown a record and…

About this research line

What has Pranay Mahendrakar published on Reasoning & Memory?

Pranay Mahendrakar has published 26 open-access papers on Reasoning & Memory: "Pooling Experience Spends Independence Twice", "Errors That Should Compound, and When They Do", "Perspectives Without Independence", "Self-Model or Self-Simulation? A Machine Self-Awareness Index Averages Sub-Scores With No Common Referent and No Fixed Sign, Why Persistent Identity, Goal Stability and Memory Continuity Are Not Evidence of Self-Access, and the Validity Tests Any Composite Would Have to Pass", "Compressed Once, Read Many Times", "Curriculum Is Three Claims, Not One", "Faithful to What? Four Instruments for Chain-of-Thought Faithfulness Disagree With Each Other, and the One Ground-Truth Check Run So Far Found Most of Them Near Chance", "When Deliberation Hurts", "Two Kinds of Missing", "Three Things Called Budget Awareness", "Scored Before the Question Exists", "A Lesson Is an Untested Counterfactual", "Consolidation Without Weights", "Delete Names Five Operations", "Transfer Is a Directed Relation", "A Safety Memory Is a Declassification Channel", "A Conflict Is Constructed Before It Is Measured", "Newer Is Not Truer", "A Step Score Is Not a Step Verdict", "Out-of-Distribution With Respect to What? Four Referents Behind One Predicate, and Why OOD Detection Results Do Not Transfer to Foundation-Model Reasoning", "Intent Is Not a Property of the Record", "The Crossover Does Not Carry the Claim", "Memory Architectures Beyond Attention", "Causal Representation Learning from Observational Video", "Energy-Based Models for Reasoning" and "Mechanistic Interpretability of In-Context Learning". Each is deposited on Zenodo with a permanent DOI and is free to read.

← All research topics