MEGA Hub

STAGE: Controlled Objective Admission for Multi-Preference LLM Alignment

Authors

Do you know Yongqi Tong?You can claim authorship or link another user.Do you know Zhenyu Zhang?You can claim authorship or link another user.Do you know Ruirui Wang?You can claim authorship or link another user.Do you know Kewei Fu?You can claim authorship or link another user.Do you know Shaoqing Lin?You can claim authorship or link another user.Do you know Sijie Dong?You can claim authorship or link another user.Do you know Jiang-Ming Yang?You can claim authorship or link another user.Do you know Xin Zhang?You can claim authorship or link another user.Do you know Jianshe Li?You can claim authorship or link another user.

Abstract

Multi-preference alignment is often framed as scalarization: combine reward dimensions, then optimize. This leaves a temporal decision underspecified: when should each preference dimension enter policy optimization? We propose \methodname, a stability-guided active-set controller for controlled objective admission. \methodname starts from a small active set, retains admitted objectives, and expands when reward-deviation gates indicate low recent deviation or a patience budget is exhausted. A probing phase estimates a hard-to-easy order, and adaptive weighting emphasizes underperforming active dimensions. Automatic evaluations with 15 training preferences and 16 held-out benchmark columns show that \methodname obtains higher averages than simultaneous scalarization and shared-budget adapted baselines. Component ablations and expansion dynamics further support cumulative retention, gated admission, and probing-derived ordering as useful design choices in this setting. These results position objective-entry timing as a concrete control variable in reward-vector RLHF.

Community

00