MEGA Hub

A Unified Algorithmic Framework for Hybrid Reinforcement Learning in Tabular MDPs with Shifted Transition Dynamics

Authors

Do you know Zheshun Wu?You can claim authorship or link another user.Do you know Renjie Zheng?You can claim authorship or link another user.Do you know Jinhang Zuo?You can claim authorship or link another user.Do you know Zenglin Xu?You can claim authorship or link another user.Do you know Fang Kong?You can claim authorship or link another user.

Abstract

This paper investigates a hybrid reinforcement learning setting in tabular Markov Decision Processes (MDPs), where an agent aims to learn an optimal policy by combining online interactions with a target environment and offline data from a source environment. A central challenge is that offline data may be collected from outdated environments with shifted transition dynamics, making naive integration of historical data ineffective. To address this, we propose a unified algorithmic framework featuring two algorithms: MIN-UCB-VI for regret minimization and MAX-LCB-VI for best policy identification. Both algorithms leverage fine-grained bias information to more effectively exploit offline data under general transition shifts. We provide theoretical guarantees for our framework, including both instance-dependent and independent upper bounds on regret and sub-optimality gap. Furthermore, we establish matching lower bounds to demonstrate the optimality of our approach and validate our theoretical findings through extensive experiments.

Community

00

Publication notes

Author note
59 pages, 3 figures, and 2 tables