Agreement Is Not a Criterion: What Matching Human Labels Certifies About an LLM Judge, Why a Shared Shortcut Passes the Test, and Why Every External Criterion Found So Far Validates a Use Rather Than a Judge
Abstract
An automatic judge is usually accepted on one piece of evidence: it agrees with human preference labels about as often as humans agree with each other. That number is then read as permission to use the judge for ranking models, filtering data, selecting samples and supplying reward. This paper separates three claims the ritual runs together - that the judge can stand in for the labelers on a given item population, that it tracks the quality construct the labels were meant to measure, and that decisions made with it turn out well in a stated use - and argues that agreement evidence bears directly only on the first. A stylised two-loading model makes the reason exact: agreement with a single reference identifies one linear combination of a judge's sensitivity to the construct and to the surface features the labelers also respond to, so a judge that shares the labelers' shortcuts can match a judge that tracks quality. The published record is then read through that separation. Human labels carry documented surface sensitivities that judges reproduce; judges carry failures of their own that natural validation sets never exercise; and agreement is concentrated on the easy, wide-gap items where decisions matter least. In the corners where an external criterion exists - verified correctness, downstream policy outcomes, randomised business outcomes and verifiable task success - the criterion has repeatedly exposed failures that agreement did not show, and where it confirms an offline score the confirmation does not transfer across uses: a benchmark score that predicts best-of-N selection well gives only a rough signal for PPO. Validity, on this evidence, is a property of a judge in a use. The paper consolidates these published results, states a use-indexed validation procedure as an algorithm, and proposes the tests that would settle what remains open.
Questions about this paper
Who wrote "Agreement Is Not a Criterion"?
Pranay Mahendrakar wrote "Agreement Is Not a Criterion: What Matching Human Labels Certifies About an LLM Judge, Why a Shared Shortcut Passes the Test, and Why Every External Criterion Found So Far Validates a Use Rather Than a Judge", published 7 Oct 2026. Pranay Mahendrakar is an Indian AI specialist and LLM engineer based in Bengaluru, India. He is the Managing Director of SonyTech, Nodal Coordinator at IIRS-ISRO, and an instructor at Tutorials Point. His work covers large language models, natural language processing, computer vision and retrieval-augmented generation. He publishes open-access research papers and is the author of three books: Just AI With Pranay, Multiverse of AI and It's Me LLM.
Is "Agreement Is Not a Criterion" free to read?
Yes. "Agreement Is Not a Criterion" by Pranay Mahendrakar is open access under a Creative Commons Attribution 4.0 licence, with the full PDF available from Zenodo at https://doi.org/10.5281/zenodo.23199460. There is no paywall and no account required.
How do I cite "Agreement Is Not a Criterion"?
Cite the DOI: Mahendrakar, P. (2026). Agreement Is Not a Criterion: What Matching Human Labels Certifies About an LLM Judge, Why a Shared Shortcut Passes the Test, and Why Every External Criterion Found So Far Validates a Use Rather Than a Judge. Zenodo. https://doi.org/10.5281/zenodo.23199460 A BibTeX entry is provided on this page.