MEGA Hub

Do Sequential Recommendation Benchmarks Really Require Higher-Order Sequence Modelling?

Authors

Do you know Aleksandr V. Petrov?You can claim authorship or link another user.Do you know Praveen Chandar?You can claim authorship or link another user.Do you know Paul N. Bennett?You can claim authorship or link another user.Do you know Hugues Bouchard?You can claim authorship or link another user.Do you know Mounia Lalmas?You can claim authorship or link another user.

Abstract

Sequential recommenders increasingly use language-model architectures designed to capture complex, context-dependent interactions. Yet it remains unclear whether widely used benchmarks actually require this modelling capacity. We investigate this question using two simple, recency-weighted pairwise probes that do not learn higher-order sequence representations: Sequential Rules (SeqRules) and our Probabilistic Collaborative Transition Model (PCTM). Using the evaluation protocol of eSASRec, at least one probe exceeds our eSASRec reproduction by 15-38% on three Amazon datasets and by 4.4% on MovieLens-1M, but trails it by 27.3% on MovieLens-20M. On the four remaining datasets, at least one probe also outperforms our sampled-softmax SASRec reproduction by 9-28%, suggesting that these widely used benchmarks are poorly suited to measuring gains from higher-order sequence modelling. More broadly, comparing Transformer-based models against strong recency-weighted pairwise probes provides a concrete test of whether a benchmark can meaningfully measure gains from higher-order sequence modelling.

Community

00

Publication notes

Author note
Accepted at the 20th ACM Conference on Recommender Systems (RecSys 2026)