MEGA Hub

Impression Share Prediction: An Offline Evaluation Task for Ranking Systems

Authors

Do you know Mohsen Malmir?You can claim authorship or link another user.Do you know Houssam Nassif?You can claim authorship or link another user.Do you know Danish Nasir Shaikh?You can claim authorship or link another user.Do you know Taher Rahgooy?You can claim authorship or link another user.Do you know Murat Ali Bayir?You can claim authorship or link another user.

Abstract

Offline evaluation is a major gateway before online evaluation of ranking models in A/B testing. Standard offline metrics measure predictive accuracy, but are only a surrogate for downstream utility: a model can improve them while redistributing impressions across objective buckets in ways that degrade downstream utility. No offline method surfaces these impression share shifts before online evaluation. We propose \emph{impression share prediction} as an offline evaluation task: given a candidate ranking model, predict the distribution of impressions it would produce across objective buckets - impressions grouped by optimization goal (e.g., click, video view). The task is inherently counterfactual, since the candidate has never served live traffic. We propose a structural causal model of how model predictions and delivery capacity jointly determine impression allocation, and show the counterfactual effect is identified from observational data. Building on this, we develop a statistical learning framework that predicts impression shares from a candidate's early-interaction confidence signals and current system state, trained on historical data. On data from multiple ranking model families, a Random Forest reduces L1 error by 49\% over a constant baseline for models seen during training. For held-out models, evaluated by time since first appearance, the first hour is the closest analog to true online evaluation and the hardest: the Random Forest falls below the baseline because the capacity state still reflects the prior model. An encoder-conditioned architecture that simulates a 2-hour rollout over recent auction dynamics recovers $+$22\% L1 in this regime.

Community

00

Publication notes

DOI
10.1145/3773078.3831809