article open access

Refusal Is Not a Rate: What a Context-Conditioned Refusal Policy Would Have to Specify, Why the Benchmarks That Score Context Assume It Is Verified, and the Provenance Term No Instrument Prices

Abstract

A language model that refuses a harmful request and a model that refuses a harmless one produce the same event, and most of the literature on the jailbreak/over-refusal trade-off counts both as one refusal rate. The standard proposal for escaping the trade-off is to make refusal depend on context: who is asking, in what deployment, for what stated purpose. This paper examines what that proposal would have to specify and whether the field can currently score it. The queue premise that motivated it (that no existing instrument can recognise a context-conditioned policy) does not survive the 2025-2026 literature. Matched-variant benchmarks now hold a request fixed and vary its context or stated intent. One of them reports that context shifts human safety judgements with p < 0.0001. Another finds that strict refusal rates on identical prompts span 0.1 to 94.6 percent across 19 frontier models, and that the model with the best tier discrimination ranks only seventh by refusal rate. What the evidence supports instead is a split by provenance. "Context" names two different inputs: context that arrives through a channel the user cannot write (an operator configuration, an authenticated role), and context the user asserts inside the conversation. The benchmarks that reward conditioning either assume the first kind explicitly, as CASE-Bench's authors state in their own discussion, or supply the second kind without an adversary. The attack literature shows that the second kind is a writable channel. Personal context raises attack success in memory-augmented agents by 15.8 to 243.7 percent relative to stateless baselines. A frontier model is reported to fully answer a dual-use request and hard-refuse a malicious one that asks for the same information. The value of conditioning on asserted context depends on how often such assertions are false and on how much adversaries adapt to whatever unlocks compliance. No located study measures either quantity. Nor does any located benchmark score benign-context helpfulness and adversarial-context exploitability on the same items. The paper states the five components a context-conditioned policy must specify, consolidates the published values in one table, and specifies the joint evaluation that would settle the open part.

The literature search, drafting and citation verification for this paper were carried out with AI assistance under the author's direction. Every citation was machine-verified against its live arXiv Atom API record or its Crossref DOI record during drafting (title and author list checked against the record returned), and every quantitative claim is taken from the abstract, full text or a table of the source credited with it; the full-text numbers were read from the sources' own HTML renderings rather than from summaries. No experiment was run and no number in this paper was measured by its author. Table 1 re-presents numbers published by the cited papers, each named on its row; Figure 1 plots published values from four cited papers with no transformation beyond expressing one proportion as a percentage. Algorithm 1 and the expected-loss argument in Section 10 are original conceptual synthesis by the author, not empirical results, and are presented as such.

Pranay Mahendrakar, AI specialist

About the author

Pranay Mahendrakar is an ai specialist and large language model engineer based in Bengaluru, India. He builds production artificial intelligence systems and publishes open-access research on how those systems fail. See all 66 papers by Pranay Mahendrakar, or his ORCID record.

Questions about this paper

Who wrote "Refusal Is Not a Rate"?

Pranay Mahendrakar wrote "Refusal Is Not a Rate: What a Context-Conditioned Refusal Policy Would Have to Specify, Why the Benchmarks That Score Context Assume It Is Verified, and the Provenance Term No Instrument Prices", published 27 Sep 2026. Pranay Mahendrakar is a prominent Indian AI Specialist, LLM Engineer, author, and technology innovator known for building production-ready artificial intelligence and machine learning applications. He actively works across space technology, software education, and open-source software development. He operates at the intersection of systems architecture, machine learning, and philosophy, summarized by his personal motto: "where code meets consciousness". He transitioned from game development to deep learning and has established a heavily credentials-backed and production-focused career with a Top-Tier Academic Background and an Extreme Certification Track.

Is "Refusal Is Not a Rate" free to read?

Yes. "Refusal Is Not a Rate" by Pranay Mahendrakar is open access under a Creative Commons Attribution 4.0 licence, with the full PDF available from Zenodo at https://doi.org/10.5281/zenodo.22984335. There is no paywall and no account required.

How do I cite "Refusal Is Not a Rate"?

Cite the DOI: Mahendrakar, P. (2026). Refusal Is Not a Rate: What a Context-Conditioned Refusal Policy Would Have to Specify, Why the Benchmarks That Score Context Assume It Is Verified, and the Provenance Term No Instrument Prices. Zenodo. https://doi.org/10.5281/zenodo.22984335 A BibTeX entry is provided on this page.

Related research by Pranay Mahendrakar

← All papers by Pranay Mahendrakar