MEGA Hub

The Illusion of Cross-Lingual Safety in Low-Resource Languages

Authors

Do you know Abigail Oppong?You can claim authorship or link another user.Do you know P Sam Sahil?You can claim authorship or link another user.Do you know Tadesse Destaw Belay?You can claim authorship or link another user.Do you know Maryam Ibrahim Mukhtar?You can claim authorship or link another user.Do you know Esmael Ahmed Abdu?You can claim authorship or link another user.Do you know Tassallah Abdullahi?You can claim authorship or link another user.Do you know Jessica Oparebea?You can claim authorship or link another user.Do you know Saminu Mohammad Aliyu?You can claim authorship or link another user.Do you know Idris Abdulmumin?You can claim authorship or link another user.Do you know Abubakar Juma Chilala?You can claim authorship or link another user.Do you know Nicholaus Dismas Ladislaus?You can claim authorship or link another user.Do you know Alfred Malengo Kondoro?You can claim authorship or link another user.Do you know Lemofouet Valdini Douglace?You can claim authorship or link another user.Do you know Shamsuddeen Hassan Muhammad?You can claim authorship or link another user.Do you know Seid Muhie Yimam?You can claim authorship or link another user.

Abstract

Safety alignment in large language models (LLMs) is largely developed in English, assuming these safeguards generalize across multilingual settings. However, this assumption remains underexplored and exposes a vulnerability in low-resource languages. We investigate cross-lingual safety transfer in four African languages, Twi, Hausa, Amharic, and Swahili, using LoDNA, a new safety dataset that pairs literal translations with culturally localized prompts. To move beyond generation-based evaluation, we propose a latent geometric framework that probes hidden-state refusal representations in LLMs. Our experimental results show that cross-lingual safety transfer is severely limited; harmful prompts retain less than 10% of the English refusal signal across most language-model pairs. Literal and localized prompts are semantically aligned (cosine 0.95-0.996) but drift across layers, suggesting models encode the concepts without routing them to safety mechanisms. These findings demonstrate that current multilingual safety alignment is superficial, providing strong evidence against the assumption of a universal, language-agnostic harm manifold within the specific low-resource languages studied. Warning: This paper contains example data that may be offensive or harmful.

Community

00