MEGA Hub

Navigating Speech Enhancement for Real-Time MRI: A Systematic Assessment of Signal Quality, Source Preservation, and Downstream Tasks

Authors

Do you know Huang-Cheng Chou?You can claim authorship or link another user.Do you know Sean Foley?You can claim authorship or link another user.Do you know Haley Hsu?You can claim authorship or link another user.Do you know Kevin Huang?You can claim authorship or link another user.Do you know Szu-Jui Chen?You can claim authorship or link another user.Do you know Rong Chao?You can claim authorship or link another user.Do you know Louis Goldstein?You can claim authorship or link another user.Do you know Khalil Iskarous?You can claim authorship or link another user.Do you know Dani Byrd?You can claim authorship or link another user.Do you know Yu Tsao?You can claim authorship or link another user.Do you know Sudarsana Reddy Kadiri?You can claim authorship or link another user.Do you know John H. L. Hansen?You can claim authorship or link another user.Do you know Shrikanth Narayanan?You can claim authorship or link another user.

Abstract

Audio recorded during real-time magnetic resonance imaging (rtMRI) is heavily contaminated by scanner noise, but it remains unclear whether general-purpose speech enhancement improves the signal for speech research and downstream processing. Three off-the-shelf systems---Denoiser, PASE, and RE-USE---are evaluated across five rtMRI corpora using naturally recorded inputs, a clean-input probe, and an archived paired additive-noise probe. The multi-task evaluation spans learned quality predictors, speaker and phone representations, reference-based intelligibility and quality measures, acoustic--phonetic probes, automatic speech recognition (ASR), and paralinguistic tasks. The central result is that enhancement effects are endpoint dependent: higher predicted-quality scores do not reliably imply better ASR performance or greater source fidelity. Across 15 corpus--recognizer comparisons using corpus-provided processed inputs, RE-USE yielded lower word-error-rate point estimates in 11, whereas Denoiser yielded higher estimates in 13. In the paired additive-noise probe, PASE and RE-USE improved recognized-phone agreement, intelligibility, and perceptual-quality point estimates. Denoiser improved recognized-phone agreement and short-time objective intelligibility (STOI) but reduced speaker-embedding similarity. No system was uniformly best across corpora, recognizers, and endpoints. Enhanced rtMRI audio should therefore be treated as a task-specific transformed derivative rather than a universally improved replacement for the original or DSP-processed waveform.

Community

00

Publication notes

Author note
Submitted to the Journal of the Acoustical Society of America (JASA). 18 pages, 3 figures, 11 tabels