MEGA Hub

Rethinking Expert Training for Model Merging with Prompt Learning

Authors

Do you know Christos Georgakilas?You can claim authorship or link another user.Do you know Aniello Panariello?You can claim authorship or link another user.Do you know Samir El Karrat Moreno?You can claim authorship or link another user.Do you know Simone Calderara?You can claim authorship or link another user.Do you know Dimosthenis Karatzas?You can claim authorship or link another user.Do you know Joost van de Weijer?You can claim authorship or link another user.

Abstract

Model merging aims to combine multiple domain-specialized experts trained from a shared foundation model into a single multi-task model. Existing approaches largely focus on improving the merging procedure itself and typically assume experts obtained through full-parameter fine-tuning. In this work, we revisit expert training for model merging. We first show that prompt-based adaptation provides a strong baseline: independently learned prompts can be exploited across tasks while keeping the backbone fixed, avoiding the interference introduced by weight merging. Building on this observation, we introduce Dual-Tuned Experts (DTEs), a two-stage training strategy that first learns prompts and then fine-tunes the vision encoder. This reduces the magnitude of task-specific parameter updates and produces experts with higher merge compatibility. Experiments across multiple CLIP architectures, full fine-tuning, and LoRA experts show that DTEs consistently improve merged performance of standard merging approaches and remain effective even when combining heterogeneous sets of experts.

Community

00

Publication notes

Author note
14 pages