A Trust Score Needs a Consumer: Four Places a Per-Source Trust Value Can Act on an LLM Agent, the Missing Test of a Graded Discount Against an Adaptive Attacker, and Why the Defences That Report Guarantees Use a Gate Instead of a Score
Abstract
Work on trust between language-model agents produces two kinds of object. Protocol work produces identity, attestation, stake and constraint, all bound at the transport layer before any content reaches a model. Behavioural work produces a per-source score: a reputation, a credibility, a reliability weight. Statements of motivation across the security literature suggest that the score has nowhere to go, because verified and unverified content arrive in the same undifferentiated context window and the model has no way to discount the low-trust part. This paper argues that this suggestion, taken literally, is false, and that it is true in a narrower form that matters more. It is false because at least four consumers of a trust value exist and have published results: admission and routing before the model, annotation written into the prompt, modulation inside the forward pass, and gates on actions after the model. Credibility annotations and attention scaling do move model outputs, and models already weigh source labels they were never asked to weigh. It is true in a narrower form because every graded discount the model reads that was located for this paper, whether written into the prompt or applied inside the forward pass, was evaluated against sources that err or against attacks fixed in advance, never against an attacker who adapts to the discount; the one graded router located that was attacked by a source writing its own evidence was captured. The label, the count of copies and the evidence behind a score are each writable by an attacker, and published attacks write all three: forged role tags, repeated low-credibility text, fabricated episodes that launder reputation. The trust defences that report guarantees against attacking sources consume trust as a discrete label or capability enforced by code outside the model, never as a graded weight the model reads, although their guarantees rest mostly on proofs and on fixed attacks rather than on adaptive evaluation. Against sources that err, by contrast, a graded discount may do better than exclusion when scores are noisy. The paper sets out the four-consumer partition, analyses which inputs of a graded consumer an attacker can write, argues that trust scored per source cannot follow influence that arrives per token, consolidates the measurements in one table, states the routing decision as an algorithm, and names eight studies that would settle the open part. No experiments are reported here.
The literature search, drafting and citation verification for this paper were carried out with AI assistance under the author's direction. Every arXiv citation was machine-verified against its live arXiv Atom API record, and every other citation against its Crossref or DOI record, during drafting (title and author list checked against the record returned). Every quantitative claim is taken from the abstract, full text or a table of the source credited with it; full-text numbers were read from the sources' own arXiv HTML renderings rather than from summaries. No experiment was run and no number in this paper was measured by its author. Table 1 re-presents numbers published by the cited papers, each named on its row; Figure 1 plots published values from four cited papers with no transformation. The consumer partition, the erring/attacking distinction as applied here, Algorithm 1 and the reporting protocol in Section 11 are original conceptual synthesis by the author, not empirical results, and are presented as such. Seven classic references (Douceur 2002; Josang, Ismail and Boyd 2007; Hardy 1988; Resnick et al. 2000; Saltzer and Schroeder 1975; Kamvar et al. 2003; Lamport et al. 1982) were verified as records; where no deposited abstract was available they are cited only for what their titles state or for the concept they named.
Questions about this paper
Who wrote "A Trust Score Needs a Consumer"?
Pranay Mahendrakar wrote "A Trust Score Needs a Consumer: Four Places a Per-Source Trust Value Can Act on an LLM Agent, the Missing Test of a Graded Discount Against an Adaptive Attacker, and Why the Defences That Report Guarantees Use a Gate Instead of a Score", published 29 Sep 2026. Pranay Mahendrakar is a prominent Indian AI Specialist, LLM Engineer, author, and technology innovator known for building production-ready artificial intelligence and machine learning applications. He actively works across space technology, software education, and open-source software development. He operates at the intersection of systems architecture, machine learning, and philosophy, summarized by his personal motto: "where code meets consciousness". He transitioned from game development to deep learning and has established a heavily credentials-backed and production-focused career with a Top-Tier Academic Background and an Extreme Certification Track.
Is "A Trust Score Needs a Consumer" free to read?
Yes. "A Trust Score Needs a Consumer" by Pranay Mahendrakar is open access under a Creative Commons Attribution 4.0 licence, with the full PDF available from Zenodo at https://doi.org/10.5281/zenodo.23026299. There is no paywall and no account required.
How do I cite "A Trust Score Needs a Consumer"?
Cite the DOI: Mahendrakar, P. (2026). A Trust Score Needs a Consumer: Four Places a Per-Source Trust Value Can Act on an LLM Agent, the Missing Test of a Graded Discount Against an Adaptive Attacker, and Why the Defences That Report Guarantees Use a Gate Instead of a Score. Zenodo. https://doi.org/10.5281/zenodo.23026299 A BibTeX entry is provided on this page.