Emergent Covert Signaling in Multi-Agent LLM Negotiation: A Conceptual Framework and Experimental Protocol
Abstract
When multiple large-language-model agents negotiate, communicate, or compete, do they spontaneously develop covert signalling — channels of communication that human observers cannot decode? Recent work has established that such behaviour is possible: LLMs can be trained or pressured into steganographic communication, encoded reasoning, and tacit collusion on pricing tasks. What remains almost entirely missing is a systematic methodology for detecting covert signalling as it emerges in the wild, in standard negotiation settings, without prompting agents to be deceptive. This paper makes three contributions. First, we disambiguate four distinct phenomena that are routinely conflated under the umbrella term "covert signalling" — steganography, convention formation, strategic ambiguity, and deceptive coordination — and argue that each requires different evidence and different mitigations. Second, we propose a measurement framework built around four detection signatures: mutual-information lift between agent messages and private state, paraphrase-invariance failure, thirdparty comprehension gap, and behavioural coordination beyond stated commitments. Third, we describe a concrete experimental protocol — a controlled multi-agent negotiation environment with explicit conditions and falsifiable predictions — that any team with API access could run today. We argue this is one of the most tractable open problems in AI safety: the methodology is achievable, the threat model is concrete, and the empirical baseline is currently almost empty.
Questions about this paper
Who wrote "Emergent Covert Signaling in Multi-Agent LLM Negotiation"?
Pranay Mahendrakar wrote "Emergent Covert Signaling in Multi-Agent LLM Negotiation: A Conceptual Framework and Experimental Protocol", published 28 Apr 2026. Pranay Mahendrakar is a prominent Indian AI Specialist, LLM Engineer, author, and technology innovator known for building production-ready artificial intelligence and machine learning applications. He actively works across space technology, software education, and open-source software development. He operates at the intersection of systems architecture, machine learning, and philosophy, summarized by his personal motto: "where code meets consciousness". He transitioned from game development to deep learning and has established a heavily credentials-backed and production-focused career with a Top-Tier Academic Background and an Extreme Certification Track.
Is "Emergent Covert Signaling in Multi-Agent LLM Negotiation" free to read?
Yes. "Emergent Covert Signaling in Multi-Agent LLM Negotiation" by Pranay Mahendrakar is open access under a Creative Commons Attribution 4.0 licence, with the full PDF available from Zenodo at https://doi.org/10.5281/zenodo.19853493. There is no paywall and no account required.
How do I cite "Emergent Covert Signaling in Multi-Agent LLM Negotiation"?
Cite the DOI: Mahendrakar, P. (2026). Emergent Covert Signaling in Multi-Agent LLM Negotiation: A Conceptual Framework and Experimental Protocol. Zenodo. https://doi.org/10.5281/zenodo.19853493 A BibTeX entry is provided on this page.