Errors That Should Compound, and When They Do: The Product Rule Assumes Independent Steps, Conformal Guarantees Assume Exchangeable Instances, and the Unit of Independence Decides What Each Can Claim
Abstract
Two bodies of work quantify uncertainty in multi-step language-model reasoning, and both rest on an independence assumption placed at some unit. Step-level error models multiply per-step reliabilities, which, at an illustrative 95% per-step accuracy, predicts that a fifty-step chain succeeds about 8% of the time; yet longer thinking often helps, and correct answers are frequently reached through erroneous steps. Conformal methods are imported for distribution-free guarantees, yet exchangeability, their one assumption, fails between the steps of an autoregressive chain by construction. This paper argues that the two problems share a structure: in each, what can be claimed depends on the unit at which independence is assumed. The product rule holds only if every step is load-bearing, step errors are independent, and failure is absorbing. Published measurements break each premise. Fed the true per-step error rates, with errors positively associated as the evidence finds, the product would then understate success; fed the clean-context rates that are usually measured, it can err in either direction, because errors in context raise later error rates. The breakage tracks model-relative task difficulty: in one study a single corrupted step propagated to the answer in 3.9% of continuations on an easy benchmark and 64.5% on a hard arithmetic task. Within that evidence, errors compound sharply where the task is hard for the model and rarely where it is easy. The conformal literature, where it is careful, names its unit and does not assume exchangeable steps: it lifts the unit to the whole reasoning instance, and its guarantees then hold, but only marginally over instances drawn from the calibration distribution, only about the annotation used, and not under the deployment shift that agent loops create. Composing per-step guarantees repeats the product rule's independence premise; a union bound avoids the premise at the price of conservatism, and the methods that avoid both calibrate the chain as one unit. What is offered is less a new finding than a side-by-side account: a synthesis of published measurements, a table of what each unit licenses, a procedure for reading uncertainty claims, and six measurements that would decide what is still open. No new data were collected.
The literature search, drafting and citation verification for this paper were carried out with AI assistance under the author's direction. Every arXiv citation was machine-verified against its live arXiv Atom API record, and every other citation against its Crossref record, during drafting; title and author list were checked against the record returned. The full texts (arXiv HTML renderings) of the load-bearing sources were read for the passages and numbers attributed to them, including Jia and Mu (2026), Zheng et al. (2025, ProcessBench), Kang et al. (2025), Zeng et al. (2025a), Ren et al. (2023, KnowNo), Cheung et al. (2026, CROP), Kotte (2026b, PASC) and Li et al. (2024, TRAQ). The remaining sources were read at abstract level, and the claims attached to them are limited to what their abstracts state. No experiment was run and no number in this paper was measured by its author. Table 1 and Figure 1 re-present published values, each named with its source. The three-premise decomposition of the product rule, the three units of exchangeability, Table 2, Algorithm 1 and the proposed tests in Section 11 are conceptual synthesis by the author, not empirical results, and are presented as such.
Questions about this paper
Who wrote "Errors That Should Compound, and When They Do"?
Pranay Mahendrakar wrote "Errors That Should Compound, and When They Do: The Product Rule Assumes Independent Steps, Conformal Guarantees Assume Exchangeable Instances, and the Unit of Independence Decides What Each Can Claim", published 1 Oct 2026. Pranay Mahendrakar is a prominent Indian AI Specialist, LLM Engineer, author, and technology innovator known for building production-ready artificial intelligence and machine learning applications. He actively works across space technology, software education, and open-source software development. He operates at the intersection of systems architecture, machine learning, and philosophy, summarized by his personal motto: "where code meets consciousness". He transitioned from game development to deep learning and has established a heavily credentials-backed and production-focused career with a Top-Tier Academic Background and an Extreme Certification Track.
Is "Errors That Should Compound, and When They Do" free to read?
Yes. "Errors That Should Compound, and When They Do" by Pranay Mahendrakar is open access under a Creative Commons Attribution 4.0 licence, with the full PDF available from Zenodo at https://doi.org/10.5281/zenodo.23086063. There is no paywall and no account required.
How do I cite "Errors That Should Compound, and When They Do"?
Cite the DOI: Mahendrakar, P. (2026). Errors That Should Compound, and When They Do: The Product Rule Assumes Independent Steps, Conformal Guarantees Assume Exchangeable Instances, and the Unit of Independence Decides What Each Can Claim. Zenodo. https://doi.org/10.5281/zenodo.23086063 A BibTeX entry is provided on this page.