publication open access

AI-Generated Text Detection Under Paraphrasing: What Has Been Solved, What Has Not, and Why Bias Matters

Pranay Mahendrakar 0009-0003-7224-029X

Abstract

The detection of AI-generated text under adversarial transformation — particularly paraphrasing — is widely characterised as an unsolved problem. The reality is more structured. Two technically distinct approaches, post-hoc detection and watermarking, are routinely conflated in policy and popular discussion despite having very different robustness properties, deployment requirements, and failure modes. This paper makes three claims. First, post-hoc detectors are demonstrably brittle to paraphrasing and exhibit serious bias against non-native English writers; their use in high-stakes contexts is currently indefensible. Second, watermarking has made substantial recent progress — semantic-invariant schemes, distortion-free constructions, and the production-scale deployment of SynthID-Text in Google's Gemini at twenty-million-response scale — but remains vulnerable to determined paraphrase attacks and faces a deployment-coordination problem the technical literature largely ignores. Third, the contested theoretical question of whether robust detection is fundamentally possible (Sadasivan et al., 2023, versus subsequent watermarking work) is genuinely unresolved and matters for what policy should expect of the technology. We propose a research agenda focused on semantic-level robustness measurement, the multivendor cooperative-detection problem, fairness as a first-class evaluation criterion, and a clearer separation of what detection can and cannot deliver in practice.

Related work

← All papers