MEGA Hub

Chemical Chain-of-Thought Functions as a Hallucination-Prone Molecular Scratchpad

Authors

Do you know Jiatong Li?You can claim authorship or link another user.Do you know Yuxuan Ren?You can claim authorship or link another user.Do you know Weida Wang?You can claim authorship or link another user.Do you know Xiaoyong Wei?You can claim authorship or link another user.Do you know Yatao Bian?You can claim authorship or link another user.

Abstract

Chemical reasoning language models are expected to derive molecular answers through faithful chain-of-thought (CoT). However, across four reasoning model families and twelve chemistry tasks, hallucination is widespread and largely decoupled from answer correctness: correct answers often coexist with fabricated structural claims absent from the relevant molecules. Yet this does not make the reasoning trace computationally irrelevant. Attribution analyses suggest a shared scratchpad function expressed in model-specific forms: Chem-R and ether-0 rely on fragmented SMILES drafts, whereas ChemDFM-R emphasizes scaffold, positional, and naming cues. Notably, perturbing Chem-R's SMILES sketches degrades generation, showing that structural drafts can be causally load-bearing even when verbal structural claims are largely inert. Together, these results show that chemical CoT is neither a faithful explanation nor merely a post-hoc rationalization, but a hallucination-prone molecular scratchpad. This finding cautions against treating CoT as direct evidence of faithful reasoning and motivates process-level supervision beyond answer-only evaluation.

Community

00

Publication notes

Author note
16 pages, 6 figures