MEGA Hub

Fine-Grained Action Recognition with Cross-Attentive Latent Sparse Experts

Authors

Do you know Imtiaz Ul Hassan?You can claim authorship or link another user.Do you know Tasweer Ahmad?You can claim authorship or link another user.Do you know Nik Bessis?You can claim authorship or link another user.Do you know Ardhendu Behera?You can claim authorship or link another user.

Abstract

Fine-grained human action recognition (FHAR) must distinguish visually similar actions that differ mainly in body configuration, timing, or local appearance. RGB representations retain visual context but often suppress joint-level geometry, whereas skeleton representations encode kinematics but discard dense spatial detail. We introduce FineX, which factorizes fine-grained cues into RGB appearance, pose heatmap geometry, and skeletal-graph topology. Pairwise cross-attention enables symmetric, stream-preserving information exchange, followed by a streamwise latent sparse Mixture-of-Experts that routes each representation to a content-dependent subset of shared experts, regularized by a load-balancing objective. FineX achieves state-of-the-art results on Gym99, Gym288, and Diving48. On the long-tailed Gym288, it raises mean class accuracy from 68.6% to 76.2% (+7.6 points) without textual supervision or large-scale vision-language pre-training, demonstrating the benefit of structured visual-pose-graph fusion and conditional expert refinement for FHAR.

Community

00