article open access

Faithful to What? Four Instruments for Chain-of-Thought Faithfulness Disagree With Each Other, and the One Ground-Truth Check Run So Far Found Most of Them Near Chance

Abstract

A chain-of-thought (CoT) trace is called "faithful" when it accurately represents the computation that produced the model's answer, as opposed to a plausible-sounding story invented after the fact (Jacovi and Goldberg, 2020). Whether published CoT traces meet this bar is disputed, and the dispute has practical stakes: chain-of-thought monitoring is proposed as an AI safety tool precisely because a faithful trace would let a human read off a model's intent (Korbak et al., 2025). This paper does not take a side in that dispute directly. It surveys what four structurally different families of published instrument -- hint-insertion tests, causal mediation and activation-level analysis, ablation of the visible reasoning text, and meta-evaluation against constructed ground truth -- actually measure, and finds that they frequently disagree with each other on the same models and the same data, not only across different research groups' setups. Switching the counterfactual operator used to test one model on one task crosses the threshold between "faithful" and "not faithful" in close to a fifth of tested configurations, with disagreements up to 44 percentage points (Basu and Chakraborty, 2026). Three classifiers scoring identical hint-acknowledgment transcripts report acknowledgment rates of 74.4%, 82.6% and 69.7% and can reverse which model ranks as more faithful (Young, 2026a). A widely cited scaling result -- that larger, more capable models produce less faithful reasoning (Lanham et al., 2023) -- correlates strongly with a confound its own instrument does not control for: raw task accuracy (R-squared 0.74; Bentham, Stringham and Marasovic, 2024). The most direct test available, a 2026 benchmark built from tasks with actual ground-truth internal causes, reports that most existing faithfulness metrics perform near chance and that the best of them reaches only 0.70 AUROC (Gur-Arieh, Marasovic and Geva, 2026). None of this settles whether chain-of-thought carries genuine causal signal or whether it is mostly post-hoc rationalization; it complicates the question of how anyone would currently tell. The paper's contribution is not a resolution but a map: it separates what each instrument family can and cannot detect, states where they have been run on the same object and disagreed, and names what a study that could actually adjudicate the dispute would need to do that no published study yet does.

The literature search, drafting and citation verification for this paper were carried out with AI assistance under the author's direction. Every citation was machine-verified against its live arXiv Atom API record before inclusion, with title and author list checked against the record returned, and every quantitative claim in this paper is taken from the abstract or stated headline result of the source credited with it. No experiment was run and no number in this paper was measured or recomputed by its author; the table re-presents numbers published by the cited papers, each named on its row. This draft was produced in a session whose execution environment did not permit running this project's own citation-check script (cite_check.py), its Zenodo publication script (publish_paper.py), or any matplotlib code to render the figure this project's own house style requires; the reference list below was instead verified by hand against live arXiv API responses fetched during drafting, and the draft is staged rather than published for exactly that reason. See the Limitations section for the full account.

Pranay Mahendrakar, AI specialist

About the author

Pranay Mahendrakar is an ai specialist and large language model engineer based in Bengaluru, India. He builds production artificial intelligence systems and publishes open-access research on how those systems fail. See all 66 papers by Pranay Mahendrakar, or his ORCID record.

Questions about this paper

Who wrote "Faithful to What? Four Instruments for Chain-of-Thought Faithfulness Disagree With Each Other, and the One Ground-Truth Check Run So Far Found Most of Them Near Chance"?

Pranay Mahendrakar wrote "Faithful to What? Four Instruments for Chain-of-Thought Faithfulness Disagree With Each Other, and the One Ground-Truth Check Run So Far Found Most of Them Near Chance", published 24 Sep 2026. Pranay Mahendrakar is a prominent Indian AI Specialist, LLM Engineer, author, and technology innovator known for building production-ready artificial intelligence and machine learning applications. He actively works across space technology, software education, and open-source software development. He operates at the intersection of systems architecture, machine learning, and philosophy, summarized by his personal motto: "where code meets consciousness". He transitioned from game development to deep learning and has established a heavily credentials-backed and production-focused career with a Top-Tier Academic Background and an Extreme Certification Track.

Is "Faithful to What? Four Instruments for Chain-of-Thought Faithfulness Disagree With Each Other, and the One Ground-Truth Check Run So Far Found Most of Them Near Chance" free to read?

Yes. "Faithful to What? Four Instruments for Chain-of-Thought Faithfulness Disagree With Each Other, and the One Ground-Truth Check Run So Far Found Most of Them Near Chance" by Pranay Mahendrakar is open access under a Creative Commons Attribution 4.0 licence, with the full PDF available from Zenodo at https://doi.org/10.5281/zenodo.22925257. There is no paywall and no account required.

How do I cite "Faithful to What? Four Instruments for Chain-of-Thought Faithfulness Disagree With Each Other, and the One Ground-Truth Check Run So Far Found Most of Them Near Chance"?

Cite the DOI: Mahendrakar, P. (2026). Faithful to What? Four Instruments for Chain-of-Thought Faithfulness Disagree With Each Other, and the One Ground-Truth Check Run So Far Found Most of Them Near Chance. Zenodo. https://doi.org/10.5281/zenodo.22925257 A BibTeX entry is provided on this page.

Related research by Pranay Mahendrakar

← All papers by Pranay Mahendrakar