MEGA Hub

Structured Phonological Representations for Audio-Articulatory rtMRI Speech Classification

Authors

Do you know Abner Hernandez?You can claim authorship or link another user.Do you know Tomás Arias Vergara?You can claim authorship or link another user.Do you know Daiqi Liu?You can claim authorship or link another user.Do you know Andreas Maier?You can claim authorship or link another user.Do you know Paula Andrea Pérez-Toro?You can claim authorship or link another user.

Abstract

Real-time MRI makes it possible to observe vocal-tract articulation during speech, but mapping these articulatory patterns to phonetic and phonological categories remains challenging. We investigate whether PhonoQ, an audio-based model trained to recognize structured phonological features, provides useful information for audio--articulatory modeling. Specifically, we extract representations from PhonoQ's Conformer module, whose training is shaped by supervision for manner, place, voicing, and vowel features. Using articulatory contours with synchronized audio-derived features, we compare WavLM-large and HuBERT-large baselines with models that incorporate PhonoQ-derived representations. Across unseen-speech and unseen-subject settings, these features improve macro-F1 for phonological targets including manner, place, voicing, vowel height, and vowel backness, and also improve fine-grained 39-phoneme classification. In a contour-only inference setting, audio-derived teacher supervision yields modest but consistent gains over contour-only training, indicating that phonological information from synchronized audio can be partially transferred to articulatory models. Finally, posterior analyses show interpretable surface-sensitive patterns consistent with flapping-like /t/ realizations, /t/-/r/ retraction or affrication, and nasal place assimilation.

Community

00

Publication notes

Author note
Submitted for review at SLT 2026