MEGA Hub

Finetuning Strategies for Querying Sounds by Vocal Imitation

Authors

Do you know Aditya Bhattacharjee?You can claim authorship or link another user.Do you know Christos Plachouras?You can claim authorship or link another user.Do you know Sungkyun Chang?You can claim authorship or link another user.Do you know Emmanouil Benetos?You can claim authorship or link another user.

Abstract

This technical report describes our winning submission to the AES AIMLA 2025 Challenge on querying sound effects by vocal imitation. We investigate two complementary fine-tuning strategies: contrastive learning with a frozen, pretrained CED encoder, and joint contrastive-triplet learning with semi-hard negatives using a MobileNetV3 encoder. This report has been updated for posterity to include details released after the challenge.

Community

00