Cross-Lingual Hallucination Patterns in Indic Languages: A Typology-Aware Research Agenda for Dravidian and Indo-Aryan Comparison
Abstract
Multilingual hallucination in large language models has begun to receive systematic attention. Recent work has measured aggregate hallucination rates across 19 to 30 languages (FACTSCORE-multilingual, MFAVA, Poly-FEVER), introduced multilingual benchmarks (HalluVerseM3, BHRAM-IL for some Indo-Aryan languages), and demonstrated that LLMs hallucinate more in low-resource languages than in English. What this growing literature does not yet provide is a typology-aware comparison of how hallucination patterns — not just rates — differ across closely-related languages. The Dravidian languages (Tamil, Telugu, Kannada, Malayalam) and the Indo-Aryan languages (Hindi, Marathi, Bengali, Gujarati) of India offer a natural experimental setting: substantial variation in training data quantity, sharply different morphological typology (Dravidian agglutination versus Indo-Aryan fusion, presence or absence of grammatical gender), shared script families in some cases, and pervasive code-mixing with English in real use. This paper makes three claims. First, the question "how often does the model hallucinate in language X" is the wrong primary question; the more informative question is "which categories of hallucination dominate in language X, and why." Second, four hypotheses about Indic-language hallucination patterns can be distinguished empirically and have meaningful policy and research implications. Third, the experimental work needed to test these hypotheses is small enough that a single research team in India could execute it within a year. We propose a measurement framework, a concrete experimental protocol, and a case for why this question is particularly tractable in the Indian research context.