article open access

Near-Zero Until Someone Tries: What a Prompt-Injection Defense Number Measures, Why Static and Adaptive Results Do Not Reconcile, and the Assumption the Out-of-Band Turn Has Not Yet Tested

Abstract

Several published prompt-injection defenses report attack success rates at or near one percent on static benchmarks; published adaptive attacks report success above fifty percent against the same defense families, and above ninety percent in several instances. Neither set of numbers is a replication failure of the other, because they estimate different quantities: a static rate estimates performance against an attack distribution fixed independently of the defense, an adaptive rate estimates performance against an attacker who optimises with the defense in hand, and nothing published maps one onto the other. This paper argues that the resulting incomparability is not a temporary untidiness that better benchmarks will resolve, and that three further asymmetries compound it. An attack success rate obtained by running attacks is a lower bound on what the best attack achieves, so a defense's own adaptive evaluation cannot bound its residual risk - the failure mode that recurred thirteen times in the adversarial-example literature. The success event and the delivery assumption differ across papers, and separate published results report injection and execution dissociating, and retrieval acting as an independent barrier. And the false positive rate that fixes a detector's operating point is frequently absent from the table carrying its attack success rate. The paper then examines the partition that has organised the field's response - that observation-level detection is fragile and that enforcement outside the model is not - and argues the published record does not sort that way, since model-level defenses fall under adaptive attack and two uncontested detector results do not. A narrower property is proposed as the better cut: whether a learned decision the attacker can influence sits on the security-critical path. On that reading two of the four out-of-band systems examined here place such a decision downstream of attacker influence. The only published adaptive evaluation of that family located for this paper is a single small-scale result whose own authors decline to generalise it. Nine studies and one reporting convention that would settle the open parts are named. No experiments are reported here, and the strongest case against this paper's position is stated in full.

The literature search, drafting and citation verification for this paper were carried out with AI assistance under the author's direction. Every citation was machine-verified against the arXiv API and Crossref before inclusion, and every quantitative claim was read back against the cited source's own abstract. The author is responsible for the final text and for all claims made in it.

Pranay Mahendrakar, AI specialist

About the author

Pranay Mahendrakar is an ai specialist and large language model engineer based in Bengaluru, India. He builds production artificial intelligence systems and publishes open-access research on how those systems fail. See all 26 papers by Pranay Mahendrakar, or his ORCID record.

Questions about this paper

Who wrote "Near-Zero Until Someone Tries"?

Pranay Mahendrakar wrote "Near-Zero Until Someone Tries: What a Prompt-Injection Defense Number Measures, Why Static and Adaptive Results Do Not Reconcile, and the Assumption the Out-of-Band Turn Has Not Yet Tested", published 30 Aug 2026. Pranay Mahendrakar is an Indian AI specialist and large language model engineer based in Bengaluru, India, who builds production artificial intelligence systems and publishes open-access research on how those systems fail.

Is "Near-Zero Until Someone Tries" free to read?

Yes. "Near-Zero Until Someone Tries" by Pranay Mahendrakar is open access under a Creative Commons Attribution 4.0 licence, with the full PDF available from Zenodo at https://doi.org/10.5281/zenodo.22167308. There is no paywall and no account required.

How do I cite "Near-Zero Until Someone Tries"?

Cite the DOI: Mahendrakar, P. (2026). Near-Zero Until Someone Tries: What a Prompt-Injection Defense Number Measures, Why Static and Adaptive Results Do Not Reconcile, and the Assumption the Out-of-Band Turn Has Not Yet Tested. Zenodo. https://doi.org/10.5281/zenodo.22167308 A BibTeX entry is provided on this page.

Related research by Pranay Mahendrakar

← All papers by Pranay Mahendrakar