MEGA Hub

Rethinking Item Tokenization in Generative Recommenders: From Fixed Atoms to Semantic Subwords

Authors

Do you know Xinrui Miao?You can claim authorship or link another user.Do you know Mingjia Yin?You can claim authorship or link another user.Do you know Jiaqing Zhang?You can claim authorship or link another user.Do you know Wei Guo?You can claim authorship or link another user.Do you know Yong Liu?You can claim authorship or link another user.Do you know Yuyang Ye?You can claim authorship or link another user.Do you know Hao Wang?You can claim authorship or link another user.Do you know Enhong Chen?You can claim authorship or link another user.

Abstract

In generative recommender systems, items are typically tokenized into fixed-length semantic ID sequences for autoregressive next-item prediction. However, for user-context modeling, this fine-grained representation triggers Intra-item Attention Overload: excessive attention is spent on low-level intra-item dependencies rather than high-level inter-item behavioral transitions. To address this, we propose Semantic Subword Tokenization (SST), which represents historical items as variable-length semantic subwords while preserving fixed-length target decoding. SST first applies Item-level Subword Tokenization (IST) to merge stable adjacent atom tokens into compact semantic subword tokens, thereby reducing intra-item reassembly in the encoder. It then introduces Behavior-induced Co-occurrence Augmentation (BCA) to inject coarse-grained semantic prefix transition signals, guiding the freed modeling capacity toward inter-item behavioral regularities. Extensive experiments on three public datasets and three generative recommender backbones show empirical improvements of SST over fixed-length and transferable variable-length SID baselines. Code is available at https://github.com/mxrcandy/Semantic-Subword-Tokenization.

Community

00

Publication notes

Author note
13 pages, 6 figures, 8 tables. Accepted to CIKM 2026