MEGA Hub

Seeds Before Objectives: Rethinking Evaluation for Low-Resource Garhwali ASR

Authors

Do you know Karamvir Singh Batra?You can claim authorship or link another user.Do you know Prathamjyot Singh?You can claim authorship or link another user.Do you know Ashima Sood?You can claim authorship or link another user.Do you know Jasmeet Singh?You can claim authorship or link another user.Do you know Sahil Sharma?You can claim authorship or link another user.

Abstract

At corpus sizes typical of low-resource dialects, single-run comparisons can yield gains that do not replicate. We show this for Garhwali, an under-resourced Indo-Aryan language of the central Himalaya, building the first reproducible multi-seed ASR benchmark on the official VAANI splits, with per-seed outputs and significance testing. Re-examining plausible gains, we find them fragile: neither Focal CTC nor a matra-weighted objective beats standard CTC under seed-level testing, the matra objective fails to cut even its targeted errors, and Hindi-to-Garhwali transfer gives no gain over direct fine-tuning. What holds up is mundane: w2v-BERT 2.0 with standard CTC reaches 47.0% WER over five seeds, beating the larger MMS-1B and comparable models; pretraining design, not parameter count, drives performance, and speed augmentation gives a small, largely consistent gain. Multi-seed evaluation on official splits separates real gains from seed noise.

Community

00

Publication notes

Author note
19 pages, 3 figures. Accepted for oral presentation at ICNLSP 2026, Trento, Italy, September 2026