What has Pranay Mahendrakar published on Evaluation & Detection?
Pranay Mahendrakar has published 45 open-access papers on Evaluation & Detection: "Spend the Budget on the Known or the Unknown? Open-Set Active Learning Scores Two Opposed Objectives Under One Name, Its Filtering Side's Own Results Do Not Treat Unknowns as Waste, and the Winner Is Set by a Relevance Label and a Query Price the Benchmarks Fix in Advance", "The Evaluator Is the Bottleneck", "Whose Goals? Autotelic Agents Generate Goals Within Spaces They Are Given", "Problem Choice Without a Referee", "Errors That Should Compound, and When They Do", "Fingerprints That Survive What? Model Lineage Names Three Relations, Robustness Is Indexed by Which Party Is Adversarial, and a Benchmark's Distillation Column Scores the Parent That Was Not Distilled From", "Drift or New Class? Without Labels a Drifted Class and a New One Can Produce the Same Stream, the Two Lines of Work With the Most Explicit Assumptions Each Get an Answer by Freezing the Variable the Other Lets Move, and No Located Benchmark Scores the Attribution", "A Trust Score Needs a Consumer", "Self-Model or Self-Simulation? A Machine Self-Awareness Index Averages Sub-Scores With No Common Referent and No Fixed Sign, Why Persistent Identity, Goal Stability and Memory Continuity Are Not Evidence of Self-Access, and the Validity Tests Any Composite Would Have to Pass", "Four Things Called Forgetting", "Refusal Is Not a Rate", "How Many Classes Are Out There? In Category Discovery the Class Count Is Granularity Carried Over From the Labelled Set, Supplying It Can Supply the Taxonomy, and Without It the Count Is Set by a Hyperparameter", "Does a Model Forget Differently When the Data Is Its Own? RL's Retention Advantage and Model Collapse Are Claims About the Same Loop, and No Study Has Measured Both", "Curriculum Is Three Claims, Not One", "Faithful to What? Four Instruments for Chain-of-Thought Faithfulness Disagree With Each Other, and the One Ground-Truth Check Run So Far Found Most of Them Near Chance", "When Deliberation Hurts", "Who Pulls the Plug? Self-Report, External Authority, and the Limits of Self-Isolation in AI Control", "Two Kinds of Missing", "Sufficient for Whose Answer? Three Predicates Under One Word in Selective Retrieval-Augmented Generation, and Why the Only One Tied to Correctness Cannot Be Computed Where the Gate Runs", "The Allowance Is Doing the Work", "Not Acting Is Not One Decision", "Measured Against an Incomplete Answer Key", "The Reassessment That Did Not Travel", "Attribution Is Scored on a Finished Trace", "A Failure to Reject Is Not a Finding", "Two Supply Chains, One Artifact", "Consistency Is Not Correctness", "Transfer Is a Directed Relation", "Predicting Your Own Failure Is Not Predicting Another Model's Advantage", "A Conflict Is Constructed Before It Is Measured", "An Unknown Unknown Is Defined Relative to an Oracle", "Near-Zero Until Someone Tries", "A Step Score Is Not a Step Verdict", "Ordering Is Not Resolution", "Out-of-Distribution With Respect to What? Four Referents Behind One Predicate, and Why OOD Detection Results Do Not Transfer to Foundation-Model Reasoning", "Automatic Benchmarks Measure an Asymmetry", "Poison Below the Base Rate", "The Crossover Does Not Carry the Claim", "Three Guarantees Under One Word", "A Second Model Is Not a Second Opinion", "The Self-Verification Gap", "Cross-Lingual Hallucination Patterns in Indic Languages", "AI-Generated Text Detection Under Paraphrasing", "Catastrophic Forgetting in Continual RLHF" and "The Emotional Intelligence Paradox in Large Language Models". Each is deposited on Zenodo with a permanent DOI and is free to read.