MEGA Hub

Unified Video Dense Prediction from Disjoint Data

Authors

Do you know Yihong Sun?You can claim authorship or link another user.Do you know Seoung Wug Oh?You can claim authorship or link another user.Do you know Jiahui Huang?You can claim authorship or link another user.Do you know Bharath Hariharan?You can claim authorship or link another user.Do you know Joon-Young Lee?You can claim authorship or link another user.

Abstract

Scene understanding requires simultaneous prediction about geometry, appearance, and semantics. However, existing task-specific annotations are fragmented across incompatible, domain-specific datasets. Current unified systems circumvent this by restricting training to fully co-annotated data, or by incurring the large computational cost of pseudo-labeling. To mitigate this, we introduce UniD, a unified video model that jointly predicts eight dense scene properties-depth, surface normals, semantic segmentation, boundaries, human parts, albedo, shading, and materials-all learned from disjoint, domain-specific datasets. We propose a simple yet effective distillation step in which per-task experts supervise a unified backbone through lightweight task projectors, eliminating the need for annotation overlap or pseudo-labeling. Our key insight is that the strong visual priors of a pretrained diffusion model are sufficient to bridge the domain gaps introduced by disjoint training sources, enabling robust generalization to scene-task combinations never seen during training. UniD achieves competitive performance against per-task specialists and multi-task baselines, with strong generalization to out-of-distribution scenarios and enhanced temporal and cross-task consistency. Code and video results are available at https://unid-video.github.io/.

Community

00

Publication notes

Author note
ECCV 2026