article open access

Predicting Your Own Failure Is Not Predicting Another Model's Advantage: What Router Results Establish About Agent Self-Knowledge, and the Incremental-Validity Test Nobody Has Run

Abstract

A common reading of the LLM routing literature holds that delegation works without self-knowledge: a router allocates queries using external features of the task and observed outcome statistics over a model pool, and never consults the candidate model's opinion of its own competence. This paper argues that the reading does not survive the evidence. Signals derived from the candidate model's own generation are competitive with trained external routers and, in the one located comparison that varies distribution shift, substantially better out of distribution; a self-confidence gate has been reported to beat a frozen external router at lower token cost; and the strongest current cascade and escalation systems are built on exactly such signals. What the evidence does support is a different partition, which the internal-versus-external frame obscures. Predicting whether you will fail is a self-directed prediction and is comparatively tractable. Predicting which alternative would do better is an other-directed prediction, and the same score that ranks own-failure well has been reported to fall sharply when asked which collaboration protocol pays off. Delegation is the second kind of judgement, not the first, and the learning-to-defer literature has known for years that a deferral rule must model the expert rather than only the deferrer. A further problem is that no located routing experiment isolates introspection at all. Every deployed self-signal is a supervised probe fitted to labelled outcomes, which is an external estimator that happens to read internal features, so the axis the delegation story rests on is not the axis the experiments vary. The control that would settle it already exists in the introspection literature, where a model's self-prediction is scored against what a second model with matched knowledge predicts about it. That control has not been run on a router. Four measurements are set out that would settle the open part, and what is not known is stated flatly: nobody has reported whether self-assessed capability adds anything to a delegation decision once an external estimator with the same information is already in the system.

The literature search, drafting and citation verification for this paper were carried out with AI assistance under the author's direction. Every citation was machine-verified against the arXiv API and Crossref before inclusion, and every quantitative claim was read back against the cited source's own abstract. The author is responsible for the final text and for all claims made in it.

Pranay Mahendrakar, AI specialist

About the author

Pranay Mahendrakar is an ai specialist and large language model engineer based in Bengaluru, India. He builds production artificial intelligence systems and publishes open-access research on how those systems fail. See all 66 papers by Pranay Mahendrakar, or his ORCID record.

Questions about this paper

Who wrote "Predicting Your Own Failure Is Not Predicting Another Model's Advantage"?

Pranay Mahendrakar wrote "Predicting Your Own Failure Is Not Predicting Another Model's Advantage: What Router Results Establish About Agent Self-Knowledge, and the Incremental-Validity Test Nobody Has Run", published 4 Sep 2026. Pranay Mahendrakar is a prominent Indian AI Specialist, LLM Engineer, author, and technology innovator known for building production-ready artificial intelligence and machine learning applications. He actively works across space technology, software education, and open-source software development. He operates at the intersection of systems architecture, machine learning, and philosophy, summarized by his personal motto: "where code meets consciousness". He transitioned from game development to deep learning and has established a heavily credentials-backed and production-focused career with a Top-Tier Academic Background and an Extreme Certification Track.

Is "Predicting Your Own Failure Is Not Predicting Another Model's Advantage" free to read?

Yes. "Predicting Your Own Failure Is Not Predicting Another Model's Advantage" by Pranay Mahendrakar is open access under a Creative Commons Attribution 4.0 licence, with the full PDF available from Zenodo at https://doi.org/10.5281/zenodo.23022340. There is no paywall and no account required.

How do I cite "Predicting Your Own Failure Is Not Predicting Another Model's Advantage"?

Cite the DOI: Mahendrakar, P. (2026). Predicting Your Own Failure Is Not Predicting Another Model's Advantage: What Router Results Establish About Agent Self-Knowledge, and the Incremental-Validity Test Nobody Has Run. Zenodo. https://doi.org/10.5281/zenodo.23022340 A BibTeX entry is provided on this page.

Related research by Pranay Mahendrakar

← All papers by Pranay Mahendrakar