A Step Score Is Not a Step Verdict: Three Quantities Under the Name Process Supervision, and the Outcome-Only Results That Make the Distinction Load-Bearing
Abstract
Process reward models are introduced, almost without exception, as models that score whether an individual reasoning step is correct. This paper argues that the phrase "process supervision" is currently attached to three different target quantities, that the dominant scalable label defines the second of them rather than the first, and that the field's step-level benchmarks score the first while its downstream gains are argued for in terms of the third. The three are step validity, a property of a trajectory prefix; prefix value, the probability that some completion policy reaches the correct final answer from that prefix; and step advantage, the change in that probability across a step. Monte Carlo estimation defines prefix value, and one 2026 paper states the consequence directly: the resulting rewards are policy-dependent where step correctness should not be. A leading published argument for the process-reward paradigm makes that policy-relativity explicit rather than accidental, holding that progress should be measured under a prover policy distinct from the base policy and that weak provers can improve stronger ones. The empirical fact that forces the distinction into the open is an anomaly on the validity benchmark itself: on ProcessBench, models trained with no step-level labels at all repeatedly match or beat models trained with them, and one paper reports that adding step labels to an outcome-trained model brings no further improvement. Three readings of that anomaly are set out, together with the dissociation analysis that would separate them, which requires no new annotation. What a validity label would have to supply that a value label does not is stated, along with the label sources whose semantics is validity. No experiments are reported here. The strongest case against this paper's position, including a theorem that cuts against its premise, is stated in full.
The literature search, drafting and citation verification for this paper were carried out with AI assistance under the author's direction. Every citation was machine-verified against the arXiv API and Crossref before inclusion, and every quantitative claim was read back against the cited source's own abstract. The author is responsible for the final text and for all claims made in it.
Questions about this paper
Who wrote "A Step Score Is Not a Step Verdict"?
Pranay Mahendrakar wrote "A Step Score Is Not a Step Verdict: Three Quantities Under the Name Process Supervision, and the Outcome-Only Results That Make the Distinction Load-Bearing", published 29 Aug 2026. Pranay Mahendrakar is an Indian AI specialist and large language model engineer based in Bengaluru, India, who builds production artificial intelligence systems and publishes open-access research on how those systems fail.
Is "A Step Score Is Not a Step Verdict" free to read?
Yes. "A Step Score Is Not a Step Verdict" by Pranay Mahendrakar is open access under a Creative Commons Attribution 4.0 licence, with the full PDF available from Zenodo at https://doi.org/10.5281/zenodo.22151453. There is no paywall and no account required.
How do I cite "A Step Score Is Not a Step Verdict"?
Cite the DOI: Mahendrakar, P. (2026). A Step Score Is Not a Step Verdict: Three Quantities Under the Name Process Supervision, and the Outcome-Only Results That Make the Distinction Load-Bearing. Zenodo. https://doi.org/10.5281/zenodo.22151453 A BibTeX entry is provided on this page.