MEGA Hub

End-to-end Conditional Diffusion for Realistic and Controllable Visual Traffic Scenario Generation

Authors

Do you know Jingzheng Li?You can claim authorship or link another user.Do you know Yufei Ge?You can claim authorship or link another user.Do you know Zhijun Chen?You can claim authorship or link another user.Do you know Qianren Mao?You can claim authorship or link another user.Do you know Zizhe Wang?You can claim authorship or link another user.Do you know Binhang Qi?You can claim authorship or link another user.Do you know Bing Li?You can claim authorship or link another user.Do you know Keyu Chen?You can claim authorship or link another user.Do you know Baochang Zhang?You can claim authorship or link another user.Do you know Xianglong Liu?You can claim authorship or link another user.Do you know Philip S Yu?You can claim authorship or link another user.

Abstract

Generating closed-loop traffic scenarios that are both realistic and controllable is crucial for evaluating autonomous driving systems, especially under rare safety-critical interactions. Existing learning-based methods often struggle to balance controllability and realism, offering either limited fine-grained control over traffic behavior or controllable scenarios at the expense of behavioral plausibility. This paper presents E2E-CDiff, an end-to-end conditional diffusion framework for controllable and realistic scenario generation. Conditioned on front-view visual observations, E2E-CDiff jointly denoises future motion states and executable low-level controls for route-interacting background vehicles. This unified state-action generation mitigates the planning-control mismatch in conventional two-stage trajectory-then-controller pipelines. Differentiable guidance further regulates speed, enforces drivable-area compliance, and supports collision-avoidance or collision-seeking behaviors, enabling both naturalistic and safety-critical scenario generation. Experiments on Bench2Drive show that E2E-CDiff achieves a favorable controllability-realism trade-off compared with representative reinforcement- and imitation-learning baselines, while its collision-guided variant induces challenging interactions across multiple autonomous driving systems. E2E-CDiff also performs competitively as a learning-based ego planner, demonstrating the generality of end-to-end state-action diffusion.

Community

00