MEGA Hub

Generalised Bellman recurrence and three dualities in sequential decision-making

Authors

Do you know Fernando E. Rosas?You can claim authorship or link another user.Do you know David Hyland?You can claim authorship or link another user.Do you know Daniel Polani?You can claim authorship or link another user.

Abstract

What gives the Bellman equation its form? We show that the recursive properties of optimal value functions follow from three conditions: that the dynamics decomposes through sufficient statistics, that the return decomposes recursively, and that the aggregation of uncertainty is compatible with both. When all three conditions hold on a common state, the Bellman equation arises from their mutual consistency; when one fails, tractability can often be recovered by augmenting the state or by deforming return or dynamics. The same conditions are shown to give rise to three dualities: one between probability and return, one between return and aggregation, and one between aggregation and probability. Our framework reveals these dualities as arising from a single construction, unifying methods developed separately across reinforcement learning, control, and decision theory.

Community

00

Publication notes

Author note
27 pages, 2 figures