MEGA Hub

Kernel weighted importance sampling for off-policy evaluation in contextual bandits

Authors

Do you know Joshua Spear?You can claim authorship or link another user.Do you know Matthieu Komorowski?You can claim authorship or link another user.Do you know Rebecca Pope?You can claim authorship or link another user.Do you know Neil J Sebire?You can claim authorship or link another user.Do you know Erica E. M. Moodie?You can claim authorship or link another user.

Abstract

This article presents a novel estimator for performing off-policy evaluation using only offline data for contextual bandits. The proposed estimator, Kernel-WIS is demonstrated to be asymptotically consistent and to empirically outperform strong baselines (including vanilla weighted importance sampling), particularly under complex conditions including behaviour policy miss-specification. The benefit of Kernel-WIS is derived from combining the bounded property of vanilla weighted importance sampling with the linearity of vanilla importance sampling.

Community

00