MEGA Hub

OpenLongTail: Generative Scaling of Long-Tail Driving Data

Authors

Do you know Lulin Liu?You can claim authorship or link another user.Do you know Nuo Chen?You can claim authorship or link another user.Do you know Yan Wang?You can claim authorship or link another user.Do you know Bangya Liu?You can claim authorship or link another user.Do you know Wenyan Cong?You can claim authorship or link another user.Do you know Hezhen Hu?You can claim authorship or link another user.Do you know Boris Ivanovic?You can claim authorship or link another user.Do you know Hao Wang?You can claim authorship or link another user.Do you know Ziyao Zeng?You can claim authorship or link another user.Do you know Xinyu Gong?You can claim authorship or link another user.Do you know Yang Zhou?You can claim authorship or link another user.Do you know Zixiang Xiong?You can claim authorship or link another user.Do you know Dilin Wang?You can claim authorship or link another user.Do you know Zhangyang Wang?You can claim authorship or link another user.Do you know Weisong Shi?You can claim authorship or link another user.Do you know Ruohan Zhang?You can claim authorship or link another user.Do you know Marco Pavone?You can claim authorship or link another user.Do you know Zhiwen Fan?You can claim authorship or link another user.

Abstract

Scaling robust driving policies is fundamentally bottlenecked by the scarcity of edge cases in curated datasets. While the real world continuously captures these critical events, such long-tail events remain underutilized when collected from heterogeneous sources. Specifically, diverse but valuable in-the-wild long-tail videos lack the full view coverage required for training policy models, often missing multi-view poses or originating solely from monocular dash cameras. This modality gap prevents these ubiquitous observations from being converted into scalable training data for long-tail generalization. We introduce OpenLongTail, an open-source generative data engine for scaling autonomous driving policies under long-tail events. To transform heterogeneous data sources into view-aligned and temporally coherent multi-view assets that are useful for policy learning, we develop a pose-informed extrapolative view synthesis pipeline that generates the missing views. We further enhance cross-view consistency and the temporal alignment for the newly generated views by injecting Plücker ray geometry into the scalable generation engine. By synthesizing heterogeneous long-tail data, we observe a significant improvement in closed-loop driving robustness in handling long-tail events. By measuring the extrapolative view synthesis and pose metrics, we validate the effectiveness of OpenLongTail in visual fidelity, cross-view consistency, and ego-trajectory recovery.

Community

00