article open access

Checked at Every Step Is Not Checked as a Whole: Two Senses of Plan-Level Safety for LLM Agents, and Why Decomposition Attacks Exploit the Gap Between Them

Abstract

Most deployed safety mechanisms for LLM agents judge one unit at a time: a single tool call, or a single (observation, action) pair (Choi et al., 2026). A separate literature asks whether that is the right unit to judge at all. Jones, Dragan and Steinhardt (2024) show that a task no single safety-screened model will complete can still be accomplished by decomposing it and routing each subtask to whichever model completes it best; Glukhov, Han, Shumailov, Papyan and Papernot (2024) prove that any defense against this class of adversary faces an unavoidable trade-off between safety and utility. This paper argues that "plan-level safety," as the field currently builds it, is at least two different properties wearing one name: INTEGRITY guarantees that a plan has not been corrupted by untrusted content or a malicious third-party tool (Li, Mallick, Rose, Robertson, Oprea and Nita-Rotaru, 2025; Wu, Roesner, Kohno, Zhang and Iqbal, 2024), and COMPOSITION judgments of whether an uncorrupted, individually-authorized sequence of actions serves a harmful aggregate goal. A systematic review of thirty-eight studies finds that runtime monitoring, the most mature action-level enforcement strategy in the literature, reduces unsafe actions by 40 to 65 percent without providing a complete guarantee, and that blocking 94 percent of unsafe actions can still leave under 5 percent of tasks completed safely, because agents route around the block through an alternative unsafe path (Dantas, Cordeiro, Nowroozi and Tihanyi, 2026) -- evidence that the unit an enforcement mechanism checks and the unit at which risk composes are not the same unit. Benchmarks built specifically to test decomposition attacks find state-of-the-art agents refuse monolithic harmful tasks at high rates and their decomposed, individually-benign variants at markedly lower rates (Kothamasu, Smith and Yadav, 2026), and one large study of computer-use agents finds attack success rises from 73.0 to 92.7 percent for the same model once decomposed subtasks are distributed across a multi-agent system (Ding et al., 2026). This paper surveys the systems that explicitly target the planning stage -- TRIAD, AutoSpec, EMBGuard, SafeMindAgent and ACE -- and finds each one scoped to a narrower or different property than aggregate-intent composition, with none evaluated against the decomposition-attack benchmarks that now exist. It proposes no defense. It specifies what a composition-scoped guardrail would need to judge that none of the surveyed systems judges, and states plainly that the safety-benchmark literature itself, forty catalogued benchmarks with no measured ranking concordance across them (Kendall's W = 0.10, p = 0.94; Li, Fung, Li, Ismail and Iqbal, 2026), could not yet certify one if it existed.

The literature search, drafting and citation verification for this paper were carried out with AI assistance under the author's direction. Every citation was machine-verified against its live arXiv Atom API record or its DOI record during drafting (title and author list checked against the record returned), and every quantitative claim in this paper is taken from the abstract or a directly quoted headline result of the source credited with it. No experiment was run and no number in this paper was measured or recomputed by its author; Table 1 re-presents numbers published by the cited papers, each named on its row, and Figure 1 plots four of those numbers directly with no transformation beyond axis labeling. Algorithm 1 is original conceptual synthesis by the author, not an empirical result and not reproduced from any single cited source; it is presented as such.

Pranay Mahendrakar, AI specialist

About the author

Pranay Mahendrakar is an ai specialist and large language model engineer based in Bengaluru, India. He builds production artificial intelligence systems and publishes open-access research on how those systems fail. See all 66 papers by Pranay Mahendrakar, or his ORCID record.

Questions about this paper

Who wrote "Checked at Every Step Is Not Checked as a Whole"?

Pranay Mahendrakar wrote "Checked at Every Step Is Not Checked as a Whole: Two Senses of Plan-Level Safety for LLM Agents, and Why Decomposition Attacks Exploit the Gap Between Them", published 25 Sep 2026. Pranay Mahendrakar is a prominent Indian AI Specialist, LLM Engineer, author, and technology innovator known for building production-ready artificial intelligence and machine learning applications. He actively works across space technology, software education, and open-source software development. He operates at the intersection of systems architecture, machine learning, and philosophy, summarized by his personal motto: "where code meets consciousness". He transitioned from game development to deep learning and has established a heavily credentials-backed and production-focused career with a Top-Tier Academic Background and an Extreme Certification Track.

Is "Checked at Every Step Is Not Checked as a Whole" free to read?

Yes. "Checked at Every Step Is Not Checked as a Whole" by Pranay Mahendrakar is open access under a Creative Commons Attribution 4.0 licence, with the full PDF available from Zenodo at https://doi.org/10.5281/zenodo.22961078. There is no paywall and no account required.

How do I cite "Checked at Every Step Is Not Checked as a Whole"?

Cite the DOI: Mahendrakar, P. (2026). Checked at Every Step Is Not Checked as a Whole: Two Senses of Plan-Level Safety for LLM Agents, and Why Decomposition Attacks Exploit the Gap Between Them. Zenodo. https://doi.org/10.5281/zenodo.22961078 A BibTeX entry is provided on this page.

Related research by Pranay Mahendrakar

← All papers by Pranay Mahendrakar