article open access

Out-of-Distribution With Respect to What? Four Referents Behind One Predicate, and Why OOD Detection Results Do Not Transfer to Foundation-Model Reasoning

Abstract

Out-of-distribution detection was formalised for a setting in which the training distribution is known, the label space is closed, and a designated test set stands in for everything outside it. Foundation models satisfy none of those conditions, and yet a growing body of work reports out-of-distribution detection for large language models using perplexity, hidden-state distance or learned probes, and a parallel body of work reports that chain-of-thought reasoning degrades sharply under distribution shift. This paper argues that "out of distribution" carries at least four distinct referents in that literature - membership in the pretraining corpus, novelty of the semantic class, distance from a designated reference dataset, and impending model error - and that results are routinely obtained against one referent and read as though they held for another. The separation is not a matter of taste. For the membership referent, the published record contains a direct negative result: membership inference and contamination detection on pretraining data perform at or near chance, and the evaluations that appeared to show otherwise were measuring distribution shift between the member and non-member samples rather than membership. That makes the referent most often invoked in informal discussion the one with no working operational test, which in turn means the claim that a reasoning failure occurred "outside the training distribution" is currently untestable in that sense for any model whose corpus is undisclosed. Nine studies that would settle the open parts are named, four of them runnable today on open-corpus model suites. No experiments are reported here, and the case against this paper's own position is stated in full.

The literature search, drafting and citation verification for this paper were carried out with AI assistance under the author's direction. Every citation was machine-verified against the arXiv API and Crossref before inclusion, and every quantitative claim was read back against the cited source's own text. The author is responsible for the final text and for all claims made in it.

Pranay Mahendrakar, AI specialist

About the author

Pranay Mahendrakar is an ai specialist and large language model engineer based in Bengaluru, India. He builds production artificial intelligence systems and publishes open-access research on how those systems fail. See all 24 papers by Pranay Mahendrakar, or his ORCID record.

Questions about this paper

Who wrote "Out-of-Distribution With Respect to What? Four Referents Behind One Predicate, and Why OOD Detection Results Do Not Transfer to Foundation-Model Reasoning"?

Pranay Mahendrakar wrote "Out-of-Distribution With Respect to What? Four Referents Behind One Predicate, and Why OOD Detection Results Do Not Transfer to Foundation-Model Reasoning", published 27 Aug 2026. Pranay Mahendrakar is an Indian AI specialist and large language model engineer based in Bengaluru, India, who builds production artificial intelligence systems and publishes open-access research on how those systems fail.

Is "Out-of-Distribution With Respect to What? Four Referents Behind One Predicate, and Why OOD Detection Results Do Not Transfer to Foundation-Model Reasoning" free to read?

Yes. "Out-of-Distribution With Respect to What? Four Referents Behind One Predicate, and Why OOD Detection Results Do Not Transfer to Foundation-Model Reasoning" by Pranay Mahendrakar is open access under a Creative Commons Attribution 4.0 licence, with the full PDF available from Zenodo at https://doi.org/10.5281/zenodo.22123547. There is no paywall and no account required.

How do I cite "Out-of-Distribution With Respect to What? Four Referents Behind One Predicate, and Why OOD Detection Results Do Not Transfer to Foundation-Model Reasoning"?

Cite the DOI: Mahendrakar, P. (2026). Out-of-Distribution With Respect to What? Four Referents Behind One Predicate, and Why OOD Detection Results Do Not Transfer to Foundation-Model Reasoning. Zenodo. https://doi.org/10.5281/zenodo.22123547 A BibTeX entry is provided on this page.

Related research by Pranay Mahendrakar

← All papers by Pranay Mahendrakar