article open access

The Allowance Is Doing the Work: Constructed Premises in Chain-of-Thought Step Verification, What Formal Validity Still Certifies Once They Are Permitted, and the Dependence Test Argumentation Runs and Step Verification Does Not

Abstract

A verifier that checks the steps of a chain of thought cannot demand that each step state everything it relies on, because no real step does. Every published step verifier therefore permits a class of premises that the text does not contain: the 2023 system that introduced step-wise deductive verification says in as many words that its format permits the use of commonsense knowledge not listed in the premises, and the 2025 neuro-symbolic system that extended the idea to law and biomedicine constructs such premises automatically when a step does not follow. This paper argues that the allowance is not a concession at the margin of those systems but the mechanism that produces almost all of their output, and that the field has not measured what the allowance costs. The size is on record and has not been read this way. Turning premise construction off in the 2025 system drops its verification pass rate from 45.2 to 3.3 percent on a synthetic rulebase, from 25.3 to 2.9 on biomedical question answering, and from 15.2 to 0.6 on statutory tax reasoning; a 2026 argumentation system reports that direct translation without construction passes the solver on zero instances, under four language models, on both of its datasets. What a solver certifies afterwards is therefore conditional on premises the verifier wrote, and a 2026 result makes the hazard concrete rather than hypothetical: a system refined against solver feedback alone reaches proofs that succeed with the original premise deleted on 25.03 and 22.36 percent of its verified cases, and the authors state that refining toward provability inflates validity faster than faithfulness. Three routes could fix the target. Answer correctness cannot, because chains that reach the right answer without supporting it are a named category with their own label. Human annotation cannot in the general case: pooled three-way agreement on detecting that something has been left unstated is Krippendorff's alpha 0.516, falling to 0.453 once the implicit element must also be typed; no warrant resource in that literature's own survey reports agreement on the reconstruction at all; the step-level benchmarks quarantine or discard the items their annotators cannot agree on; and the largest complication category behind that disagreement, for attribution steps, is world knowledge. Closing the world synthetically works and does not transfer. The instrument that would measure the exposure exists: a dependence probe that re-runs the proof with the premise removed, validated in 2026 on argumentative inference, where it reduces premise-bypassing from about a quarter of verified cases to 4.06 and 2.72 percent. No chain-of-thought step verifier located here reports it. The unsupported leap resists definition less than it resists measurement, and the measurement is available.

The literature search, drafting and citation verification for this paper were carried out with AI assistance under the author's direction. Every citation was machine-verified against the DataCite or Crossref record for its DOI before inclusion, and every quantitative claim was read back against the cited source's own table or text before it was written down. No experiment was run and no number in this paper was measured by its author; every number is quoted from the paper credited with it. Section 2 states the search procedure and its limits so that the coverage claims in Sections 11 and 13 can be checked and, if wrong, corrected. The author is responsible for the final text and for all claims made in it.

Pranay Mahendrakar, AI specialist

About the author

Pranay Mahendrakar is an ai specialist and large language model engineer based in Bengaluru, India. He builds production artificial intelligence systems and publishes open-access research on how those systems fail. See all 66 papers by Pranay Mahendrakar, or his ORCID record.

Questions about this paper

Who wrote "The Allowance Is Doing the Work"?

Pranay Mahendrakar wrote "The Allowance Is Doing the Work: Constructed Premises in Chain-of-Thought Step Verification, What Formal Validity Still Certifies Once They Are Permitted, and the Dependence Test Argumentation Runs and Step Verification Does Not", published 18 Sep 2026. Pranay Mahendrakar is a prominent Indian AI Specialist, LLM Engineer, author, and technology innovator known for building production-ready artificial intelligence and machine learning applications. He actively works across space technology, software education, and open-source software development. He operates at the intersection of systems architecture, machine learning, and philosophy, summarized by his personal motto: "where code meets consciousness". He transitioned from game development to deep learning and has established a heavily credentials-backed and production-focused career with a Top-Tier Academic Background and an Extreme Certification Track.

Is "The Allowance Is Doing the Work" free to read?

Yes. "The Allowance Is Doing the Work" by Pranay Mahendrakar is open access under a Creative Commons Attribution 4.0 licence, with the full PDF available from Zenodo at https://doi.org/10.5281/zenodo.22821403. There is no paywall and no account required.

How do I cite "The Allowance Is Doing the Work"?

Cite the DOI: Mahendrakar, P. (2026). The Allowance Is Doing the Work: Constructed Premises in Chain-of-Thought Step Verification, What Formal Validity Still Certifies Once They Are Permitted, and the Dependence Test Argumentation Runs and Step Verification Does Not. Zenodo. https://doi.org/10.5281/zenodo.22821403 A BibTeX entry is provided on this page.

Related research by Pranay Mahendrakar

← All papers by Pranay Mahendrakar