MEGA Hub

HSTGFormer: Hyper Spatial-Temporal Graph Transformer for 3D Human Pose Estimation

Authors

Do you know Ruochen Li?You can claim authorship or link another user.Do you know Shuang Chen?You can claim authorship or link another user.Do you know Wenke E?You can claim authorship or link another user.Do you know Farshad Arvin?You can claim authorship or link another user.Do you know Amir Atapour-Abarghouei?You can claim authorship or link another user.

Abstract

Transformer-based methods have achieved strong performance in monocular 3D human pose estimation, but most existing approaches organise spatial and temporal reasoning as separate stages, which may weaken unified spatial-temporal interdependencies inherent in human motion and compress frame-level structural information before temporal modelling. In this paper, we propose HSTGFormer, a graph-enhanced Transformer framework that reformulates spatial-temporal reasoning as localised coupled graph aggregation over joint-time nodes. Specifically, HSTGFormer introduces a Hyper Spatial-Temporal Graph (HSTG), which decomposes global spatial-temporal reasoning into local spatial-temporal receptive fields around individual joint-time nodes by extending per-frame skeleton graphs into temporal neighbourhoods, thereby enabling structure-aware coupled reasoning while preserving local structural motion information. It further incorporates an Adaptive Dual-Scale Temporal Graph (ADSTG) to capture joint-specific temporal dependencies over complementary short- and long-range windows. A lightweight node-wise fusion module further adaptively integrates the two graph representations for each joint-time node. Experiments on Human3.6M and MPI-INF-3DHP show that HSTGFormer achieves strong accuracy with high computational efficiency.

Community

00

Publication notes

Author note
Accepted to BMVC 2026, full paper