MEGA Hub

Redistribution-based Cost Inference Improves Sparse Safe Offline RL

Authors

Do you know Ebenezer Gelo?You can claim authorship or link another user.Do you know Geraud Nangue Tasse?You can claim authorship or link another user.Do you know Steven James?You can claim authorship or link another user.Do you know Benjamin Rosman?You can claim authorship or link another user.

Abstract

Safe offline RL typically assumes access to dense per-step cost annotations, but in practice supervisors provide only trajectory-level stop-feedback: a binary signal at the first unsafe transition, with no per-step attribution. We frame this as a temporal credit assignment problem and propose the Redistribution-based Cost Inference (RCI) framework, which converts sparse stop-feedback into dense per-step costs via return decomposition, then trains a constrained offline policy on the augmented dataset. We show that return-equivalent redistribution preserves the feasible policy set and the optimal Lagrangian in a CMDP, establishing that the transformation is lossless in theory while yielding better-conditioned cost critic learning in practice. Experiments on highway driving and robotic manipulation demonstrate substantially lower violation rates than sparse and classifier-based baselines, with robustness to heterogeneous dataset compositions and label noise.

Community

00

Publication notes

Author note
Accepted at the 1st IJCAI Workshop on Safe Physical AI (SPAI 2026), affiliated with IJCAI/ECAI 2026