MEGA Hub

Knowledge-Geometry Decoupling: Refreshable Pretrained Transfer for Streaming Recommendation

Authors

Do you know Zixuan Wang?You can claim authorship or link another user.Do you know Yuhong Chen?You can claim authorship or link another user.Do you know Yuxuan Zhu?You can claim authorship or link another user.Do you know Guidong Lei?You can claim authorship or link another user.Do you know Zhiluohan Guo?You can claim authorship or link another user.Do you know Yu Zhao?You can claim authorship or link another user.Do you know Kun Wang?You can claim authorship or link another user.Do you know Bangyang Hong?You can claim authorship or link another user.Do you know Kangle Wu?You can claim authorship or link another user.Do you know Yabo Ni?You can claim authorship or link another user.Do you know Anxiang Zeng?You can claim authorship or link another user.Do you know Cong Fu?You can claim authorship or link another user.Do you know Hui Li?You can claim authorship or link another user.

Abstract

Industrial recommenders increasingly adopt the pretrain-then-transfer paradigm, yet behavioral distribution drift raises two questions: what to learn from behavior sequences, and how to transfer the learned knowledge while the pretrained model is continually refreshed. To resolve them, we propose Knowledge-Geometry Decoupling (KGD). For what to learn, conventional next-token prediction treats adjacency as dependency and may encode spurious transitions across unrelated sessions. We introduce Behavioral Multi-Token Prediction (BMTP) to retain only collaboratively or semantically related future items as supervision, yielding cleaner and more transferable behavioral knowledge. For how to transfer, pretrained knowledge and task-specific geometry impose conflicting optimization demands on shared parameters. To handle it, KGD assigns them to separate parameter sets: a refreshable encoder owns behavioral knowledge, while a task learner reads contextualized encoder states through read-only cross-attention and writes task-specific geometry through Anchored Calibration Residual (ACR) orthogonal to the pretrained embedding. The decoupled ownership enables continual knowledge refresh without task-gradient interference or invalidating downstream adaptation. KGD improves over strong pretrain-transfer baselines by 4-12% on eight public benchmarks and sustains its advantage over a 90-day production stream where baselines show no gains. KGD has been fully deployed in Shopee. In a live A/B test on Shopee Homepage Search, it increases GMV per user by 1.75% and advertising revenue by 1.53%, demonstrating its high practical value. We provide the core implementation of KGD at https://github.com/FuCongResearchSquad/KGD4REC.

Community

00