MEGA Hub

Robust Multi-Tier Infant-Centered Audio Understanding with Whisper via Structured Speaker Conditioning

Authors

Do you know Xulin Fan?You can claim authorship or link another user.Do you know Jialu Li?You can claim authorship or link another user.Do you know Mohammad Nur Hossain Khan?You can claim authorship or link another user.Do you know Kexin Hu?You can claim authorship or link another user.Do you know Bashima Islam?You can claim authorship or link another user.Do you know Mark Hasegawa-Johnson?You can claim authorship or link another user.Do you know Nancy L. McElwain?You can claim authorship or link another user.

Abstract

Recent advances in model design and self-supervised audio representations have improved speech and audio understanding, yet infant-centered naturalistic recordings remain challenging due to limited labeled data, low signal-to-noise ratio, and cross-family domain shifts. We present a family-conditioned, multi-tier audio tagger that combines a LoRA-finetuned Whisper encoder with a lightweight, target-speaker-aware Transformer for long-context inference and framewise prediction across tiers. To improve temporal coherence, we incorporate a simple sequence-level smoothing loss, and to enhance robustness across households, we introduce a factorized speaker-token design with a shared tier token and a learned family-specific offset, reducing family bias and promoting generalizable representations. Together, these choices enable efficient and effective infant-centered audio tagging of daylong audio recordings in home environments.

Community

00

Publication notes

Author note
Accepted to Interspeech 2026