Transfer Is a Directed Relation: Four Quantities Under One Word in the Cross-Domain RLVR Debate, and Why the Structural-Similarity Reading Does Not Survive Its Own Evidence
Abstract
Reinforcement learning with verifiable rewards is the standard route to reasoning-tuned language models, and the field has split over whether its gains leave the training domain. One body of work reports that most models succeeding at mathematics fail to transfer and that single-domain post-training yields no statistically significant out-of-domain improvement; another reports that post-training on constraint-satisfaction puzzles alone raises hard-mathematics accuracy substantially. This paper argues that the two are not in direct contradiction, because the word transfer carries at least four logically independent quantities: whether training on a domain improves others, whether a domain improves when others are trained, whether a domain is preserved rather than degraded, and whether joint training beats separate training. Papers measuring one are routinely cited as evidence about another. The paper then argues that the reconciliation the field has settled on - that transfer follows structural similarity between source and target - is contradicted by the directional findings of the negative result usually cited for it, which reports unstructured domains transferring to structured ones while failing to transfer to each other. Similarity is symmetric; the reported relation is not, and multi-task learning has treated directed, sign-bearing task-affinity matrices as its normal object of study since Taskonomy. Two further problems are set out: the source-side gain may be substantially elicitation of a pretraining-frequent behaviour rather than acquired skill, which would relocate transfer to the base model, and the reported effect sizes sit near a documented seed-to-seed noise floor. What is not known is stated flatly: no published experiment reports a full directed transfer matrix over a fixed domain set, one protocol and more than one model family. Five measurements that would settle the open part are specified.
The literature search, drafting and citation verification for this paper were carried out with AI assistance under the author's direction. Every citation was machine-verified against the arXiv API and Crossref before inclusion, and every quantitative claim was read back against the cited source's own abstract or, where a claim is drawn from a paper's body, against the located passage. The author is responsible for the final text and for all claims made in it.
Questions about this paper
Who wrote "Transfer Is a Directed Relation"?
Pranay Mahendrakar wrote "Transfer Is a Directed Relation: Four Quantities Under One Word in the Cross-Domain RLVR Debate, and Why the Structural-Similarity Reading Does Not Survive Its Own Evidence", published 5 Sep 2026. Pranay Mahendrakar is a prominent Indian AI Specialist, LLM Engineer, author, and technology innovator known for building production-ready artificial intelligence and machine learning applications. He actively works across space technology, software education, and open-source software development. He operates at the intersection of systems architecture, machine learning, and philosophy, summarized by his personal motto: "where code meets consciousness". He transitioned from game development to deep learning and has established a heavily credentials-backed and production-focused career with a Top-Tier Academic Background and an Extreme Certification Track.
Is "Transfer Is a Directed Relation" free to read?
Yes. "Transfer Is a Directed Relation" by Pranay Mahendrakar is open access under a Creative Commons Attribution 4.0 licence, with the full PDF available from Zenodo at https://doi.org/10.5281/zenodo.23022357. There is no paywall and no account required.
How do I cite "Transfer Is a Directed Relation"?
Cite the DOI: Mahendrakar, P. (2026). Transfer Is a Directed Relation: Four Quantities Under One Word in the Cross-Domain RLVR Debate, and Why the Structural-Similarity Reading Does Not Survive Its Own Evidence. Zenodo. https://doi.org/10.5281/zenodo.23022357 A BibTeX entry is provided on this page.