article open access

Ordering Is Not Resolution: What the Instruction Hierarchy Defines, What It Leaves Undefined, and the Same-Tier Conflicts Agent Benchmarks Do Not Separate

Abstract

When a language-model agent receives instructions that conflict, the dominant remedy is a privilege ordering over sources: system above developer, developer above user, user above tool output. This paper argues that the ordering paradigm is well-defined for one class of conflict and undefined for another, and that published benchmarks and training sets do not separate the two. A conflict between instructions carrying different privilege labels has an answer the paradigm can state; a conflict between two instructions carrying the same label does not, because an ordering over privilege levels does not induce an ordering among the instructions inside a level. The second class is not hypothetical. One profiler of real deployed prompt policies reports that across thirteen thousand jointly governed trials, only about a third of cases satisfy both of two individually reasonable standing rules. Multi-principal deployments, where two users hold equal authority, instantiate the same structure by construction. The paper distinguishes three failure classes that the single phrase "instruction hierarchy failure" currently covers, argues that reported hierarchy-compliance numbers are sums over classes with different remedies, and states what a same-tier resolution rule would have to supply that an ordering does not. The case for the ordering paradigm is presented first and is not weak: several recent results report large, transferable gains from training on ordering. Nine studies that would settle the open parts are named. No experiments are reported here, and the strongest case against this paper's own position is stated in full.

The literature search, drafting and citation verification for this paper were carried out with AI assistance under the author's direction. Every citation was machine-verified against the arXiv API and Crossref before inclusion, and every quantitative claim was read back against the cited source's own abstract. The author is responsible for the final text and for all claims made in it.

Pranay Mahendrakar, AI specialist

About the author

Pranay Mahendrakar is an ai specialist and large language model engineer based in Bengaluru, India. He builds production artificial intelligence systems and publishes open-access research on how those systems fail. See all 24 papers by Pranay Mahendrakar, or his ORCID record.

Questions about this paper

Who wrote "Ordering Is Not Resolution"?

Pranay Mahendrakar wrote "Ordering Is Not Resolution: What the Instruction Hierarchy Defines, What It Leaves Undefined, and the Same-Tier Conflicts Agent Benchmarks Do Not Separate", published 28 Aug 2026. Pranay Mahendrakar is an Indian AI specialist and large language model engineer based in Bengaluru, India, who builds production artificial intelligence systems and publishes open-access research on how those systems fail.

Is "Ordering Is Not Resolution" free to read?

Yes. "Ordering Is Not Resolution" by Pranay Mahendrakar is open access under a Creative Commons Attribution 4.0 licence, with the full PDF available from Zenodo at https://doi.org/10.5281/zenodo.22135193. There is no paywall and no account required.

How do I cite "Ordering Is Not Resolution"?

Cite the DOI: Mahendrakar, P. (2026). Ordering Is Not Resolution: What the Instruction Hierarchy Defines, What It Leaves Undefined, and the Same-Tier Conflicts Agent Benchmarks Do Not Separate. Zenodo. https://doi.org/10.5281/zenodo.22135193 A BibTeX entry is provided on this page.

Related research by Pranay Mahendrakar

← All papers by Pranay Mahendrakar