45 / papers

Evaluation & Detection

Measuring what models do rather than what benchmarks say they do. 45 open-access papers by Pranay Mahendrakar, each with a permanent DOI.

article

Spend the Budget on the Known or the Unknown? Open-Set Active Learning Scores Two Opposed Objectives Under One Name, Its Filtering Side's Own Results Do Not Treat Unknowns as Waste, and the Winner Is Set by a Relevance Label and a Query Price the Benchmarks Fix in Advance

Active learning chooses which unlabelled examples a human should label. When the unlabelled pool contains classes the model has never been shown, two lines of work give opposite instructions under the same name. One, usually called…

article

The Evaluator Is the Bottleneck: What a Self-Modifying System's Acceptance Test Must Satisfy Before Its Verdict Can Be Trusted, Why Co-Evolving Evaluators Shrink the External Anchor Rather Than Remove It, and How Moving the Anchor Outside Moves the Attack Surface With It

Self-modifying AI systems now decide which changes to their own code to keep by running an acceptance test and retaining whatever scores higher. Published coding-agent loops report large benchmark gains this way, and their safety case…

article

Whose Goals? Autotelic Agents Generate Goals Within Spaces They Are Given: The Goal Space and the Referee Stay Outside the Agent, and a Foundation Model in the Loop Relocates Them Rather Than Removing Them

Autotelic agents are described as learning to represent, generate, select and solve their own goals, and a new generation of systems built on foundation models is reported to do so without hand-coded goal representations, without human…

article

Problem Choice Without a Referee: An Automated Novelty Check Certifies an Empty Search, a Rediscovery Benchmark Credits the Hypothesis Humans Found Next, and Every Located Referee of Worth Independent of Proposer and Field Needs the Objective Given

Pipelines that claim to automate scientific discovery gate their search on a judgement that a proposed idea or problem is novel and worth pursuing. The best-controlled evidence on that judgement points two ways at once: in a blind study…

article

Errors That Should Compound, and When They Do: The Product Rule Assumes Independent Steps, Conformal Guarantees Assume Exchangeable Instances, and the Unit of Independence Decides What Each Can Claim

Two bodies of work quantify uncertainty in multi-step language-model reasoning, and both rest on an independence assumption placed at some unit. Step-level error models multiply per-step reliabilities, which, at an illustrative 95%…

article

Fingerprints That Survive What? Model Lineage Names Three Relations, Robustness Is Indexed by Which Party Is Adversarial, and a Benchmark's Distillation Column Scores the Parent That Was Not Distilled From

Language-model fingerprinting is asked to answer questions of the form "is this model derived from that one?", and its methods are routinely reported as robust to fine-tuning, quantization, pruning, merging and distillation. This paper…

article

Drift or New Class? Without Labels a Drifted Class and a New One Can Produce the Same Stream, the Two Lines of Work With the Most Explicit Assumptions Each Get an Answer by Freezing the Variable the Other Lets Move, and No Located Benchmark Scores the Attribution

A classifier deployed on a stream eventually sees inputs its model does not explain. Two different events can produce them: a known class can have drifted, or a class that did not exist in training can have appeared. The stream-mining…

article

A Trust Score Needs a Consumer: Four Places a Per-Source Trust Value Can Act on an LLM Agent, the Missing Test of a Graded Discount Against an Adaptive Attacker, and Why the Defences That Report Guarantees Use a Gate Instead of a Score

Work on trust between language-model agents produces two kinds of object. Protocol work produces identity, attestation, stake and constraint, all bound at the transport layer before any content reaches a model. Behavioural work produces a…

article

Self-Model or Self-Simulation? A Machine Self-Awareness Index Averages Sub-Scores With No Common Referent and No Fixed Sign, Why Persistent Identity, Goal Stability and Memory Continuity Are Not Evidence of Self-Access, and the Validity Tests Any Composite Would Have to Pass

Some proposals to quantify machine self-awareness combine several sub-scores - persistent identity, goal stability, cross-session memory continuity, contradiction detection, uncertainty awareness, self-prediction and introspective access…

article

Four Things Called Forgetting: Interference, Transience, Reset and Unlearning Remove Different Things, Why Forgetting Looks Easy by Accident and Hard on Purpose Only Under Different Instruments, and the Savings Measurement That Would Tell Them Apart

Machine learning uses one word for four operations. Catastrophic forgetting is damage that fine-tuning does to earlier capabilities. Transience is the fading of individual training examples during ordinary training. Resets deliberately…

article

Refusal Is Not a Rate: What a Context-Conditioned Refusal Policy Would Have to Specify, Why the Benchmarks That Score Context Assume It Is Verified, and the Provenance Term No Instrument Prices

A language model that refuses a harmful request and a model that refuses a harmless one produce the same event, and most of the literature on the jailbreak/over-refusal trade-off counts both as one refusal rate. The standard proposal for…

article

How Many Classes Are Out There? In Category Discovery the Class Count Is Granularity Carried Over From the Labelled Set, Supplying It Can Supply the Taxonomy, and Without It the Count Is Set by a Hyperparameter

Category discovery asks a model to sort an unlabelled image collection into classes, some of which it has never been shown, using a labelled subset of other classes as its guide. Almost every method needs one number before it can produce…

article

The Allowance Is Doing the Work: Constructed Premises in Chain-of-Thought Step Verification, What Formal Validity Still Certifies Once They Are Permitted, and the Dependence Test Argumentation Runs and Step Verification Does Not

A verifier that checks the steps of a chain of thought cannot demand that each step state everything it relies on, because no real step does. Every published step verifier therefore permits a class of premises that the text does not…

article

Near-Zero Until Someone Tries: What a Prompt-Injection Defense Number Measures, Why Static and Adaptive Results Do Not Reconcile, and the Assumption the Out-of-Band Turn Has Not Yet Tested

Several published prompt-injection defenses report attack success rates at or near one percent on static benchmarks; published adaptive attacks report success above fifty percent against the same defense families, and above ninety percent…

About this research line

What has Pranay Mahendrakar published on Evaluation & Detection?

Pranay Mahendrakar has published 45 open-access papers on Evaluation & Detection: "Spend the Budget on the Known or the Unknown? Open-Set Active Learning Scores Two Opposed Objectives Under One Name, Its Filtering Side's Own Results Do Not Treat Unknowns as Waste, and the Winner Is Set by a Relevance Label and a Query Price the Benchmarks Fix in Advance", "The Evaluator Is the Bottleneck", "Whose Goals? Autotelic Agents Generate Goals Within Spaces They Are Given", "Problem Choice Without a Referee", "Errors That Should Compound, and When They Do", "Fingerprints That Survive What? Model Lineage Names Three Relations, Robustness Is Indexed by Which Party Is Adversarial, and a Benchmark's Distillation Column Scores the Parent That Was Not Distilled From", "Drift or New Class? Without Labels a Drifted Class and a New One Can Produce the Same Stream, the Two Lines of Work With the Most Explicit Assumptions Each Get an Answer by Freezing the Variable the Other Lets Move, and No Located Benchmark Scores the Attribution", "A Trust Score Needs a Consumer", "Self-Model or Self-Simulation? A Machine Self-Awareness Index Averages Sub-Scores With No Common Referent and No Fixed Sign, Why Persistent Identity, Goal Stability and Memory Continuity Are Not Evidence of Self-Access, and the Validity Tests Any Composite Would Have to Pass", "Four Things Called Forgetting", "Refusal Is Not a Rate", "How Many Classes Are Out There? In Category Discovery the Class Count Is Granularity Carried Over From the Labelled Set, Supplying It Can Supply the Taxonomy, and Without It the Count Is Set by a Hyperparameter", "Does a Model Forget Differently When the Data Is Its Own? RL's Retention Advantage and Model Collapse Are Claims About the Same Loop, and No Study Has Measured Both", "Curriculum Is Three Claims, Not One", "Faithful to What? Four Instruments for Chain-of-Thought Faithfulness Disagree With Each Other, and the One Ground-Truth Check Run So Far Found Most of Them Near Chance", "When Deliberation Hurts", "Who Pulls the Plug? Self-Report, External Authority, and the Limits of Self-Isolation in AI Control", "Two Kinds of Missing", "Sufficient for Whose Answer? Three Predicates Under One Word in Selective Retrieval-Augmented Generation, and Why the Only One Tied to Correctness Cannot Be Computed Where the Gate Runs", "The Allowance Is Doing the Work", "Not Acting Is Not One Decision", "Measured Against an Incomplete Answer Key", "The Reassessment That Did Not Travel", "Attribution Is Scored on a Finished Trace", "A Failure to Reject Is Not a Finding", "Two Supply Chains, One Artifact", "Consistency Is Not Correctness", "Transfer Is a Directed Relation", "Predicting Your Own Failure Is Not Predicting Another Model's Advantage", "A Conflict Is Constructed Before It Is Measured", "An Unknown Unknown Is Defined Relative to an Oracle", "Near-Zero Until Someone Tries", "A Step Score Is Not a Step Verdict", "Ordering Is Not Resolution", "Out-of-Distribution With Respect to What? Four Referents Behind One Predicate, and Why OOD Detection Results Do Not Transfer to Foundation-Model Reasoning", "Automatic Benchmarks Measure an Asymmetry", "Poison Below the Base Rate", "The Crossover Does Not Carry the Claim", "Three Guarantees Under One Word", "A Second Model Is Not a Second Opinion", "The Self-Verification Gap", "Cross-Lingual Hallucination Patterns in Indic Languages", "AI-Generated Text Detection Under Paraphrasing", "Catastrophic Forgetting in Continual RLHF" and "The Emotional Intelligence Paradox in Large Language Models". Each is deposited on Zenodo with a permanent DOI and is free to read.

← All research topics