MEGA Hub

Inter-X++: A Comprehensive Benchmark for Multimodal Human-Human Interaction Analysis

Authors

Do you know Liang Xu?You can claim authorship or link another user.Do you know Chengqun Yang?You can claim authorship or link another user.Do you know Zili Lin?You can claim authorship or link another user.Do you know Xintao Lv?You can claim authorship or link another user.Do you know Yichao Yan?You can claim authorship or link another user.Do you know Xin Jin?You can claim authorship or link another user.Do you know Zhibo Chen?You can claim authorship or link another user.Do you know Xiaokang Yang?You can claim authorship or link another user.Do you know Wenjun Zeng?You can claim authorship or link another user.

Abstract

The capability to perceive and synthesize human-human interactions is fundamental to developing intelligent digital human systems. However, existing datasets and modeling approaches are fundamentally constrained by low-fidelity kinematics, the omission of dexterous hand gestures and a severe lack of rich multimodal annotations. Furthermore, fragmented interaction representations and inconsistent evaluation protocols also impede fair and rigorous benchmarking. To systematically address these bottlenecks, we present Inter-X++, a comprehensive and large-scale benchmark designed to empower versatile HHI analysis. Captured via a novel hybrid motion capture system, Inter-X++ provides 11,388 high-fidelity interaction sequences and over 8.1M frames, featuring precise whole-body movements and detailed finger articulations. Meanwhile, we enrich the data foundation with multifaceted annotations, including hierarchical fine-grained textual descriptions, interaction categories, causal interaction orders, the relationship and personality of the subjects, as well as vertex-level contact maps and physically regularized constraints. Leveraging these elaborate annotations, we formulate a unified testing ground comprising four categories of downstream tasks that symmetrically span both generative and perceptive paradigms. To eliminate benchmarking ambiguities, we systematically standardize the interaction representations and evaluation protocols. Finally, we go beyond dataset construction to propose OpenHHI, a single and unified HHI representation and modeling framework that jointly optimizes interaction reconstruction and semantic understanding. Extensive experiments reveal that OpenHHI achieves state-of-the-art performance on both generation and perception tasks. This definitively proves that our unified representation successfully bridges interaction understanding and generation simultaneously.

Community

00

Publication notes

Author note
24 pages, 10 figures