8 / papers

Safety & Verification

Guarantees, specifications, and the gap between a proof and a deployed system. 8 open-access papers by Pranay Mahendrakar, each with a permanent DOI.

article

The Precursor Assumption: What an Early-Warning Threshold Must Deliver in a Frontier Safety Framework, Why the Continuity Evidence Covers Aggregates Under Fixed Elicitation Rather Than the Thresholded Task, and the Lead Time No Published Record Measures

Frontier safety frameworks decide when a model needs stronger safeguards by testing it against capability thresholds at a fixed cadence, and they place an early-warning threshold below each capability threshold so that mitigations can be…

article

Fingerprints That Survive What? Model Lineage Names Three Relations, Robustness Is Indexed by Which Party Is Adversarial, and a Benchmark's Distillation Column Scores the Parent That Was Not Distilled From

Language-model fingerprinting is asked to answer questions of the form "is this model derived from that one?", and its methods are routinely reported as robust to fine-tuning, quantization, pruning, merging and distillation. This paper…

article

Refusal Is Not a Rate: What a Context-Conditioned Refusal Policy Would Have to Specify, Why the Benchmarks That Score Context Assume It Is Verified, and the Provenance Term No Instrument Prices

A language model that refuses a harmful request and a model that refuses a harmless one produce the same event, and most of the literature on the jailbreak/over-refusal trade-off counts both as one refusal rate. The standard proposal for…

About this research line

What has Pranay Mahendrakar published on Safety & Verification?

Pranay Mahendrakar has published 8 open-access papers on Safety & Verification: "The Precursor Assumption", "Fingerprints That Survive What? Model Lineage Names Three Relations, Robustness Is Indexed by Which Party Is Adversarial, and a Benchmark's Distillation Column Scores the Parent That Was Not Distilled From", "Refusal Is Not a Rate", "Checked at Every Step Is Not Checked as a Whole", "Who Pulls the Plug? Self-Report, External Authority, and the Limits of Self-Isolation in AI Control", "Formal Verification of Neural Network Safety Beyond Toy Examples", "AI-Generated Text Detection Under Paraphrasing" and "Emergent Covert Signaling in Multi-Agent LLM Negotiation". Each is deposited on Zenodo with a permanent DOI and is free to read.

← All research topics