MEGA Hub

myMediWhisper: Construction of Burmese Medical Speech Corpus and Whisper Fine-Tuning for Clinical Dialogue ASR

Authors

Do you know Ye Kyaw Thu?You can claim authorship or link another user.Do you know Ye Bhone Lin?You can claim authorship or link another user.Do you know Thura Aung?You can claim authorship or link another user.Do you know Htet Arkar?You can claim authorship or link another user.Do you know Myat Oo Swe?You can claim authorship or link another user.Do you know Thet Htet San?You can claim authorship or link another user.Do you know Min Thiha Tun?You can claim authorship or link another user.Do you know Thazin Myint Oo?You can claim authorship or link another user.Do you know Thepchai Supnithi?You can claim authorship or link another user.

Abstract

Although Whisper models benefit from large-scale multilingual pre-training, their performance on Burmese medical speech remains limited. This work presents a Burmese medical speech recognition framework built on a high-quality 28-hour corpus recorded and validated by native speakers. We fine-tune Whisper models using full fine-tuning (FFT) and parameter-efficient fine-tuning (PEFT) with LoRA. To evaluate robustness, we apply waveform- and spectrogram-level data augmentation under controlled noise and simulated room acoustics. While augmentation reduces performance on clean speech, it significantly improves robustness in noisy and reverberant environments across FFT and PEFT settings. Our best-performing system, fully fine-tuned myMediWhisper-Medium without augmentation, achieves a state-of-the-art Word Error Rate (WER) of 23.44%, outperforming much larger general-domain fine-tuned models. Dataset and other resources can be found at the Huggingface repository: https://huggingface.co/datasets/LULab/mediTalk-mm-rdy.

Community

00