MEGA Hub

Reinforcement Learning on Cost-Constrained Quadrupedal Hardware

Authors

Do you know Javier C. Weddington?You can claim authorship or link another user.Do you know Bence P. Ölveczky?You can claim authorship or link another user.Do you know Stephen A. Baccus?You can claim authorship or link another user.

Abstract

Deploying learned control policies on low-cost robotic platforms introduces transport latencies and noisy motor feedback that systematically widens the sim-to-real gap. The chasm of simulation to deployment in hardware lies in the delay of the actuator reaching the commanded position. On platforms such as the Mini Pupper 2, a measured > $50 ms transport delay transforms the locomotion task from a standard Markov decision process into a partially observable one. In this paper, we take a biologically inspired approach of handling noisy and delayed feedback to close the sim-to-real gap, thereby expanding the capability of reinforcement learning on cost-constrained hardware. Using a low-cost quadrupedal hardware platform, we find that using a forward model of the average actuator delay, paired with a time-aware neural network results in robust locomotion. Additionally, our time-aware neural network learned a central pattern generator (CPG): a self-sustaining rhythmic gait that is robust to +320 ms latency perturbations, mirroring the CPGs found in the spinal cords of vertebrates. We posit that temporal self-organization may be a general strategy for cost-constrained locomotion.

Community

00

Publication notes

Author note
Sim-to-real transfer, locomotion, reinforcement learning, central pattern generator