MEGA Hub

SmartGR: Hierarchy and Beam-Aware Knowledge Distillation for Generative Recommendation

Authors

Do you know Ziheng Zhang?You can claim authorship or link another user.Do you know Yu Cui?You can claim authorship or link another user.Do you know Bohao Wang?You can claim authorship or link another user.Do you know Yong He?You can claim authorship or link another user.Do you know Chao Yu?You can claim authorship or link another user.Do you know Chuan Yuan?You can claim authorship or link another user.Do you know Wujie Sun?You can claim authorship or link another user.Do you know Can Wang?You can claim authorship or link another user.Do you know Jiawei Chen?You can claim authorship or link another user.

Abstract

Generative recommendation (GR) has emerged as a promising paradigm for recommender systems. Scaling up GR models can improve recommendation performance, but it also substantially increases inference cost. Knowledge distillation provides a practical solution by transferring knowledge from a large GR model to a lightweight one. However, existing distillation methods do not account for two GR-specific challenges: imbalanced distillation difficulty across the semantic ID (SID) hierarchy and incorrect prefix pruning during beam search. To address these challenges, we propose SmartGR, a novel distillation framework that utilizes Hierarchy-Aware SID Distillation to transfer the teacher's modeling capability across the hierarchy and leverages Beam-Aware Ranking Distillation to distill the teacher's ranking preferences during beam search. Extensive experiments on four benchmark datasets demonstrate the effectiveness and efficiency of SmartGR, improving the performance by 8.6% while achieving a 2.39$\times$ inference speedup on average.

Community

00

Publication notes

Author note
14 pages, 4 figures, 13 tables; includes appendices