MEGA Hub

OmniMech: All-in-one Multimodal Mechanical Benchmark for 3D Reconstruction

Authors

Do you know Taiting Lu?You can claim authorship or link another user.Do you know Runze Liu?You can claim authorship or link another user.Do you know Ziwei Dong?You can claim authorship or link another user.Do you know Sisong Bei?You can claim authorship or link another user.Do you know Jingying Zeng?You can claim authorship or link another user.Do you know Mingjia Wang?You can claim authorship or link another user.Do you know Zhenghao Li?You can claim authorship or link another user.Do you know Kaiyuan Lin?You can claim authorship or link another user.Do you know Yi-Shan Wu?You can claim authorship or link another user.Do you know Yangshoudu Zheng?You can claim authorship or link another user.Do you know Hongxing Pan?You can claim authorship or link another user.Do you know Kai Zhang?You can claim authorship or link another user.Do you know Guoliang Shi?You can claim authorship or link another user.Do you know Ling Ma?You can claim authorship or link another user.Do you know Yifan Yang?You can claim authorship or link another user.Do you know Jiaying Lu?You can claim authorship or link another user.Do you know Qi He?You can claim authorship or link another user.Do you know Sung-Liang Chen?You can claim authorship or link another user.Do you know Yi-Chao Chen?You can claim authorship or link another user.Do you know Yincheng Jin?You can claim authorship or link another user.Do you know Mahanth Gowda?You can claim authorship or link another user.

Abstract

Recent vision-language models (VLMs) can generate executable CAD programs from images, but existing methods mainly target coarse, general-purpose 3D objects and rarely address the fine-grained geometry and millimeter-level tolerances required in industrial mechanical design. We introduce OmniMech, the first million-scale benchmark for evaluating VLMs on executable CAD generation from industrial manufacturing data. OmniMech contains more than 251,000 fully dimensioned and toleranced 2D orthographic drawings, paired with native CAD models, multi-view renderings, mesh, STEP and B-rep representations, and rich semantic annotations. The benchmark includes four tasks: (1) parametric CAD program synthesis from engineering drawings; (2) diagram-to-3D reasoning for geometrically and structurally consistent reconstruction; (3) annotation-grounded reasoning over dimensions, symbols, feature callouts, and manufacturing constraints; and (4) tool-augmented agentic reasoning using visualization, measurement, CAD execution, and verification tools. Experiments show that current VLMs and CAD-specialized models still struggle with executable program synthesis, fine-grained 3D reconstruction, and reliable enforcement of dimensions and tolerances. We will release the benchmark data, evaluation code, and tool interfaces to support future research.

Community

00