article open access

Cross-Lingual Hallucination Patterns in Indic Languages: A Typology-Aware Research Agenda for Dravidian and Indo-Aryan Comparison

Abstract

Multilingual hallucination in large language models has begun to receive systematic attention. Recent work has measured aggregate hallucination rates across 19 to 30 languages (FACTSCORE-multilingual, MFAVA, Poly-FEVER), introduced multilingual benchmarks (HalluVerseM3, BHRAM-IL for some Indo-Aryan languages), and demonstrated that LLMs hallucinate more in low-resource languages than in English. What this growing literature does not yet provide is a typology-aware comparison of how hallucination patterns — not just rates — differ across closely-related languages. The Dravidian languages (Tamil, Telugu, Kannada, Malayalam) and the Indo-Aryan languages (Hindi, Marathi, Bengali, Gujarati) of India offer a natural experimental setting: substantial variation in training data quantity, sharply different morphological typology (Dravidian agglutination versus Indo-Aryan fusion, presence or absence of grammatical gender), shared script families in some cases, and pervasive code-mixing with English in real use. This paper makes three claims. First, the question "how often does the model hallucinate in language X" is the wrong primary question; the more informative question is "which categories of hallucination dominate in language X, and why." Second, four hypotheses about Indic-language hallucination patterns can be distinguished empirically and have meaningful policy and research implications. Third, the experimental work needed to test these hypotheses is small enough that a single research team in India could execute it within a year. We propose a measurement framework, a concrete experimental protocol, and a case for why this question is particularly tractable in the Indian research context.

Pranay Mahendrakar, AI specialist

About the author

Pranay Mahendrakar is an ai specialist and large language model engineer based in Bengaluru, India. He builds production artificial intelligence systems and publishes open-access research on how those systems fail. See all 70 papers by Pranay Mahendrakar, or his ORCID record.

Questions about this paper

Who wrote "Cross-Lingual Hallucination Patterns in Indic Languages"?

Pranay Mahendrakar wrote "Cross-Lingual Hallucination Patterns in Indic Languages: A Typology-Aware Research Agenda for Dravidian and Indo-Aryan Comparison", published 28 Apr 2026. Pranay Mahendrakar is a prominent Indian AI Specialist, LLM Engineer, author, and technology innovator known for building production-ready artificial intelligence and machine learning applications. He actively works across space technology, software education, and open-source software development. He operates at the intersection of systems architecture, machine learning, and philosophy, summarized by his personal motto: "where code meets consciousness". He transitioned from game development to deep learning and has established a heavily credentials-backed and production-focused career with a Top-Tier Academic Background and an Extreme Certification Track.

Is "Cross-Lingual Hallucination Patterns in Indic Languages" free to read?

Yes. "Cross-Lingual Hallucination Patterns in Indic Languages" by Pranay Mahendrakar is open access under a Creative Commons Attribution 4.0 licence, with the full PDF available from Zenodo at https://doi.org/10.5281/zenodo.19854165. There is no paywall and no account required.

How do I cite "Cross-Lingual Hallucination Patterns in Indic Languages"?

Cite the DOI: Mahendrakar, P. (2026). Cross-Lingual Hallucination Patterns in Indic Languages: A Typology-Aware Research Agenda for Dravidian and Indo-Aryan Comparison. Zenodo. https://doi.org/10.5281/zenodo.19854165 A BibTeX entry is provided on this page.

Related research by Pranay Mahendrakar

← All papers by Pranay Mahendrakar