MEGA Hub

KeyFrame-Compass: Towards Comprehensive Evaluation of Keyframe-Conditioned Video Generation

Authors

Do you know Yuqi Tang?You can claim authorship or link another user.Do you know Tengfei Liu?You can claim authorship or link another user.Do you know Yizheng Lai?You can claim authorship or link another user.Do you know Yuran Wang?You can claim authorship or link another user.Do you know Yang Shi?You can claim authorship or link another user.Do you know Wanshun Su?You can claim authorship or link another user.Do you know Zhuoran Zhang?You can claim authorship or link another user.Do you know Qixun Wang?You can claim authorship or link another user.Do you know Xiaohan Zhang?You can claim authorship or link another user.Do you know Xinlei Yu?You can claim authorship or link another user.Do you know Xuehai Bai?You can claim authorship or link another user.Do you know Xuanyu Zhu?You can claim authorship or link another user.Do you know Bohan Zeng?You can claim authorship or link another user.Do you know Bozhou Li?You can claim authorship or link another user.Do you know Shujie Li?You can claim authorship or link another user.Do you know Yifan Dai?You can claim authorship or link another user.Do you know Yujie Wei?You can claim authorship or link another user.Do you know Shixuan Liu?You can claim authorship or link another user.Do you know Haotian Wang?You can claim authorship or link another user.Do you know Jialu Chen?You can claim authorship or link another user.Do you know Yuanxing Zhang?You can claim authorship or link another user.

Abstract

Video generation increasingly relies on keyframe-based workflows, where creators specify a sequence of reference images to guide generation. Although recent models support multi-keyframe conditioning, it remains unclear whether they can faithfully reproduce the prescribed keyframes while maintaining overall video quality. We present KeyFrame-Compass, the first comprehensive benchmark for evaluating keyframe-conditioned video generation. The benchmark contains 386 carefully curated samples spanning three application domains, two video structures, two prompt granularities, two conditioning formats, and four keyframe densities, enabling controlled analysis under diverse generation settings. We further introduce an automated evaluation framework that jointly measures keyframe execution and overall video quality. Specifically, we decompose keyframe execution into six complementary metrics covering presence, fidelity, temporal ordering, localization, persistence, and uniqueness, while assessing overall video quality through evidence-grounded MLLM judgments augmented with specialized perception models. Experiments on nine representative video generation systems reveal several fundamental limitations. Current models exhibit a clear trade-off between faithful keyframe execution and natural video synthesis. Their performance further degrades as keyframe constraints become denser and most open-source models also fail to interpret storyboard-grid inputs as temporally ordered keyframe sequences.

Community

00