MEGA Hub

Cultivar: A Contrastive and Locale-Oriented Translation Benchmark for Investigating Contamination and Localisation Robustness

Authors

Do you know Pinzhen Chen?You can claim authorship or link another user.Do you know Koel Dutta Chowdhury?You can claim authorship or link another user.Do you know Xiaoya Xu?You can claim authorship or link another user.Do you know David Tan?You can claim authorship or link another user.Do you know Doreen Osmelak?You can claim authorship or link another user.Do you know Ona de Gibert?You can claim authorship or link another user.Do you know Ariun-Erdene Tumurchuluun?You can claim authorship or link another user.Do you know Ashok Urlana?You can claim authorship or link another user.Do you know Fedor Sizov?You can claim authorship or link another user.Do you know Hale Sirin?You can claim authorship or link another user.Do you know Jesujoba Alabi?You can claim authorship or link another user.Do you know Karrar Talib Abed?You can claim authorship or link another user.Do you know Mateusz Klimaszewski?You can claim authorship or link another user.Do you know Nikolay Bogoychev?You can claim authorship or link another user.Do you know Niyati Bafna?You can claim authorship or link another user.Do you know Patricia Schmidtova?You can claim authorship or link another user.Do you know Preksha Manjunath Shanbhag?You can claim authorship or link another user.Do you know Sherrie Shen?You can claim authorship or link another user.Do you know Vilem Zouhar?You can claim authorship or link another user.Do you know Vivek Iyer?You can claim authorship or link another user.Do you know Yasser Hamidullah?You can claim authorship or link another user.Do you know Yusser Al Ghussin?You can claim authorship or link another user.Do you know Zheng Zhao?You can claim authorship or link another user.

Abstract

Multilingual translation benchmarks are typically sourced in English and translated into other languages, treating language pairs as the unit of evaluation---a design that is prone to contamination over time and overlooks locale and cultural considerations. We therefore advocate for source-contrastive evaluation and instantiate it with Cultivar, a localised subset of FLORES, which enables locale-specific translation evaluation. When paired with unlocalised counterparts, performance discrepancy allows the probing of data contamination and localisation robustness. We benchmark 32 open-weight models and find that MT-specialised models are less robust, a few models potentially overfit FLORES, and models tend to translate US content better than that of other locales, regardless of language.

Community

00