article open access

The Evaluator Is the Bottleneck: What a Self-Modifying System's Acceptance Test Must Satisfy Before Its Verdict Can Be Trusted, Why Co-Evolving Evaluators Shrink the External Anchor Rather Than Remove It, and How Moving the Anchor Outside Moves the Attack Surface With It

Abstract

Self-modifying AI systems now decide which changes to their own code to keep by running an acceptance test and retaining whatever scores higher. Published coding-agent loops report large benchmark gains this way, and their safety case rests on sandboxes, archives of past versions and the option to roll back. Every one of those safeguards presupposes that the acceptance test measures the property the modification is supposed to improve. This paper asks what an acceptance test would have to satisfy for its verdict to be trusted, and answers against the published record. It separates five conditions: fidelity of the test on the candidates the loop itself generates, integrity against tampering by the candidate, statistical power across many candidates and many acceptances, coverage of the state that persists after acceptance, and independence of the evaluator's errors from the proposer's. For each condition, a published loop, the benchmark it relies on, or a controlled study of the same decision records a failure. The strongest counter-evidence is a 2026 line of work that evolves the evaluator inside the loop. Read closely, every such system that reports validity against ground truth keeps a labelled anchor or frozen reference that the loop cannot write, and one reports that removing its anchor guards produces an always-pass grader that downstream task scores fail to expose. The external component shrinks; it does not disappear. A 2026 benchmark-poisoning study then shows that an externally supplied benchmark is an attack surface. No experiments are reported here. The paper states which conditions any published loop has tested, which none has, and what measurements would settle how small the external part can safely be.

The literature search, drafting and citation verification for this paper were carried out with AI assistance under the author's direction. Every arXiv citation was machine-verified against its live arXiv Atom API record, and every other citation against its Crossref or DOI record, during drafting (title and author list checked against the record returned). Every quantitative claim is taken from the abstract, full text or a table of the source credited with it; full-text numbers were read from the sources' own HTML or PDF renderings rather than from summaries. No experiment was run and no number in this paper was measured by its author. Table 1 re-presents numbers published by the cited papers, each named on its row; Figure 1 plots published values from four cited sources with no transformation. The five-condition decomposition, Table 2's reading of where each loop holds its evaluator, Algorithm 1 and the anchor argument in Sections 10 and 11 are original conceptual synthesis by the author, not empirical results, and are presented as such.

Pranay Mahendrakar, AI specialist

About the author

Pranay Mahendrakar is an ai specialist and large language model engineer based in Bengaluru, India. He builds production artificial intelligence systems and publishes open-access research on how those systems fail. See all 70 papers by Pranay Mahendrakar, or his ORCID record.

Questions about this paper

Who wrote "The Evaluator Is the Bottleneck"?

Pranay Mahendrakar wrote "The Evaluator Is the Bottleneck: What a Self-Modifying System's Acceptance Test Must Satisfy Before Its Verdict Can Be Trusted, Why Co-Evolving Evaluators Shrink the External Anchor Rather Than Remove It, and How Moving the Anchor Outside Moves the Attack Surface With It", published 4 Oct 2026. Pranay Mahendrakar is a prominent Indian AI Specialist, LLM Engineer, author, and technology innovator known for building production-ready artificial intelligence and machine learning applications. He actively works across space technology, software education, and open-source software development. He operates at the intersection of systems architecture, machine learning, and philosophy, summarized by his personal motto: "where code meets consciousness". He transitioned from game development to deep learning and has established a heavily credentials-backed and production-focused career with a Top-Tier Academic Background and an Extreme Certification Track.

Is "The Evaluator Is the Bottleneck" free to read?

Yes. "The Evaluator Is the Bottleneck" by Pranay Mahendrakar is open access under a Creative Commons Attribution 4.0 licence, with the full PDF available from Zenodo at https://doi.org/10.5281/zenodo.23130396. There is no paywall and no account required.

How do I cite "The Evaluator Is the Bottleneck"?

Cite the DOI: Mahendrakar, P. (2026). The Evaluator Is the Bottleneck: What a Self-Modifying System's Acceptance Test Must Satisfy Before Its Verdict Can Be Trusted, Why Co-Evolving Evaluators Shrink the External Anchor Rather Than Remove It, and How Moving the Anchor Outside Moves the Attack Surface With It. Zenodo. https://doi.org/10.5281/zenodo.23130396 A BibTeX entry is provided on this page.

Related research by Pranay Mahendrakar

← All papers by Pranay Mahendrakar