article open access

The Self-Verification Gap: Why Working Hallucination Detectors Are External, and What Real-Time Correction Would Actually Require

Abstract

Large language models generate fluent text that is sometimes false, and a substantial literature now proposes to have the model notice and repair those errors while it writes. Parts of the problem are settled. Semantic entropy detects confabulations across datasets and tasks without prior knowledge of the task (Farquhar et al., 2024). Search-augmented verification agrees with crowdsourced annotators 72% of the time on roughly 16,000 individual facts, at more than 20 times lower cost (Wei et al., 2024). Probes on hidden states separate true from false statements at 71-83% balanced accuracy in distribution (Azaria and Mitchell, 2023). The tension is that every one of these signals is external to anything the model reports about itself, and that the four properties a deployed detector needs - low cost, transfer under distribution shift, span-level localisation, and a false-positive rate low enough to act on - have never been demonstrated together in a single system. This paper states five constraints such a system must satisfy, argues that no published system satisfies more than two, and proposes five experiments that would settle the open questions. We report no experiments and claim no result of our own; nothing here establishes that such a system can be built.

The literature survey, drafting and citation verification for this paper were carried out with AI assistance under the author's direction. Every citation was machine-verified against Crossref and arXiv before inclusion. The author is responsible for the final text and for all claims made in it.

Pranay Mahendrakar, AI specialist

About the author

Pranay Mahendrakar is an ai specialist and large language model engineer based in Bengaluru, India. He builds production artificial intelligence systems and publishes open-access research on how those systems fail. See all 16 papers by Pranay Mahendrakar, or his ORCID record.

Questions about this paper

Who wrote "The Self-Verification Gap"?

Pranay Mahendrakar wrote "The Self-Verification Gap: Why Working Hallucination Detectors Are External, and What Real-Time Correction Would Actually Require", published 20 Aug 2026. Pranay Mahendrakar is an Indian AI specialist and large language model engineer based in Bengaluru, India, who builds production artificial intelligence systems and publishes open-access research on how those systems fail.

Is "The Self-Verification Gap" free to read?

Yes. "The Self-Verification Gap" by Pranay Mahendrakar is open access under a Creative Commons Attribution 4.0 licence, with the full PDF available from Zenodo at https://doi.org/10.5281/zenodo.22030500. There is no paywall and no account required.

How do I cite "The Self-Verification Gap"?

Cite the DOI: Mahendrakar, P. (2026). The Self-Verification Gap: Why Working Hallucination Detectors Are External, and What Real-Time Correction Would Actually Require. Zenodo. https://doi.org/10.5281/zenodo.22030500 A BibTeX entry is provided on this page.

Related research by Pranay Mahendrakar

← All papers by Pranay Mahendrakar