MEGA Hub

Towards Physics-Faithful Generation of Scientific Diagrams

Authors

Do you know Minghui Zhang?You can claim authorship or link another user.Do you know Jinxin Shi?You can claim authorship or link another user.Do you know Yifan Chang?You can claim authorship or link another user.Do you know Liangliang Zhao?You can claim authorship or link another user.Do you know Yuandong Pu?You can claim authorship or link another user.Do you know Qian Yu?You can claim authorship or link another user.Do you know Ming Hu?You can claim authorship or link another user.Do you know Hanxiao Zhang?You can claim authorship or link another user.Do you know Yun Gu?You can claim authorship or link another user.Do you know Yirong Chen?You can claim authorship or link another user.Do you know Yu Qiao?You can claim authorship or link another user.Do you know Bo Zhang?You can claim authorship or link another user.Do you know Xiangchao Yan?You can claim authorship or link another user.Do you know Bin Fu?You can claim authorship or link another user.Do you know Yihao Liu?You can claim authorship or link another user.

Abstract

Text-to-image generation has reached photorealistic quality, yet state-of-the-art systems remain unreliable at producing scientific diagrams, whose value depends not on appearance but on physical faithfulness: correct force directions, valid coordinate systems, consistent thermodynamic states, and equations matching the depicted scenario. Trained on web imagery with physically shallow captions, generic models produce diagrams that look plausible but are physically wrong, harmful in education and scientific communication. We present Princigram, a physics-faithful scientific-diagram generator, and its data pipeline. Our central advance is Structured Physical Chain-of-Thought (SP-CoT): a per-subdiscipline schema that decomposes a physics diagram into an explicit multi-step reasoning chain across six subdisciplines, from scene identification through force or process analysis to governing laws and synthesis. Unlike free-form chain-of-thought, SP-CoT follows a fixed schema with strict fidelity rules that separate visually grounded facts from physically inferred reasoning and type all mathematics symbolically; it serves both as dense training supervision and, at inference, as a structured "thinking" prompt. With it we curate and structurally annotate 4.3 million physics images, of which 115,037 carry expert-level annotation, and adapt a unified multimodal backbone. We further introduce VeriphyT2IBench, whose questions are derived from each held-out diagram's own structured annotation: each diagram becomes an item-specific bank of binary questions about its objects, forces, and states, so a judge model's score decomposes into named physical facts rather than one holistic number. On the physics subset of GenExam and on VeriphyT2IBench, Princigram shows that explicit physics-structured supervision improves the physical faithfulness of generated scientific diagrams.

Community

00