MEGA Hub

Investigating Hallucination in Conversations for Low Resource Languages

Authors

Do you know Amit Das?You can claim authorship or link another user.Do you know Md. Najib Hasan?You can claim authorship or link another user.Do you know Souvika Sarkar?You can claim authorship or link another user.Do you know Zheng Zhang?You can claim authorship or link another user.Do you know Fatemeh Jamshidi?You can claim authorship or link another user.Do you know Tathagata Bhattacharya?You can claim authorship or link another user.Do you know Nilanjana Raychawdhury?You can claim authorship or link another user.Do you know Dongji Feng?You can claim authorship or link another user.Do you know Vinija Jain?You can claim authorship or link another user.Do you know Aman Chadha?You can claim authorship or link another user.

Abstract

Large Language Models (LLMs) have demonstrated remarkable proficiency in generating text that closely resemble human writing. However, they often generate factually incorrect statements, a problem typically referred to as 'hallucination'. Addressing hallucination is crucial for enhancing the reliability and effectiveness of LLMs. While much research has focused on hallucinations in English, our study extends this investigation to conversational data in three languages: Hindi, Farsi, and Mandarin. We offer a comprehensive analysis of a dataset to examine both factual and linguistic errors in these languages for GPT-3.5, GPT-4o, Llama-3.1, Gemma-2.0, DeepSeek-R1 and Qwen-3. We found that LLMs produce very few hallucinated responses in Mandarin but generate a significantly higher number of hallucinations in Hindi and Farsi.

Community

00