MEGA Hub

GeoRoute: Geometry-Aware Hybrid Inference for Traffic Future-Frame Prediction

Authors

Do you know Khang Minh Le?You can claim authorship or link another user.Do you know Hieu Dinh Trung Pham?You can claim authorship or link another user.Do you know Luu Thanh Danh?You can claim authorship or link another user.Do you know Nam-Tien Le?You can claim authorship or link another user.Do you know Hieu Anh Ngo?You can claim authorship or link another user.Do you know Phuong Huu Vu Tran?You can claim authorship or link another user.Do you know Son Nguyen Minh Le?You can claim authorship or link another user.Do you know Nguyen Trong Nghia?You can claim authorship or link another user.Do you know Tu Tran Thi Cam?You can claim authorship or link another user.Do you know Huy Minh Nhat Nguyen?You can claim authorship or link another user.Do you know Cuong Tuan Nguyen?You can claim authorship or link another user.

Abstract

Long-horizon future-frame prediction is important for autonomous driving, traffic surveillance, and intelligent transportation systems, yet remains challenging due to temporal ghosting, geometry drift, and inconsistent object motion. Recent latent video diffusion models have achieved impressive visual quality, but directly applying them to structured traffic scenes often leads to unstable geometry and degraded temporal coherence over extended horizons. We present a training-free inference framework that stabilizes reliable static structure in pretrained video predictions through multi-frame temporal context and view-conditioned routing. For front-camera videos, our method refines generated futures with a multi-frame depth-layered renderer that projects static geometry from observed history frames while preserving dynamic regions from the generative base model. For heterogeneous traffic views, a frozen vision-language model infers a coarse camera group from the observed clip and selects a specialized motion-based predictor. The framework requires neither retraining nor fine-tuning of the underlying video model and can be applied directly to pretrained generators. We validate the proposed framework on the AI City Challenge Track 5 benchmark, where our final system achieves competitive performance among the top-ranked teams. These results demonstrate that geometry-aware inference-time refinement and view-conditioned hybrid inference can improve static-geometry stability and low-level structural fidelity without changing the original model architecture.

Community

00

Publication notes

Author note
accepted to the ECCV 2026 AI City Challenge Workshop