MEGA Hub

StructureClaw: Traceable LLM Agents and an Executable Benchmark for Structural Engineering Workflows

Authors

Do you know Sizhong Qin?You can claim authorship or link another user.Do you know Yi Gu?You can claim authorship or link another user.Do you know Yao Jiang?You can claim authorship or link another user.Do you know Ao Cai?You can claim authorship or link another user.Do you know Changjian Zhou?You can claim authorship or link another user.Do you know Shaoxuan Shuai?You can claim authorship or link another user.Do you know Jiachang Wang?You can claim authorship or link another user.Do you know Tianhao Shen?You can claim authorship or link another user.Do you know Yueqiang Li?You can claim authorship or link another user.Do you know Xinhao Li?You can claim authorship or link another user.Do you know Li Zeng?You can claim authorship or link another user.Do you know Yueshi Chen?You can claim authorship or link another user.Do you know Dachen Gao?You can claim authorship or link another user.Do you know Genrong Xu?You can claim authorship or link another user.Do you know Wenjie Liao?You can claim authorship or link another user.Do you know Xinzheng Lu?You can claim authorship or link another user.

Abstract

Addressing a structural-engineering request requires more than a single answer; it requires a chain of interdependent artifacts: interpreted requirements, a computable model, validation records, solver outputs, code-check records, and a final report. Evaluations centered on question answering or script generation rarely verify this complete evidence chain and may therefore reward fluent outputs even when the underlying engineering workflow is incomplete, internally inconsistent, or non-executable. To address this limitation, we present StructureClaw, an artifact-centered workbench in which LLM agents operate through governed engineering skills, typed tools, shared artifact state, and local analysis backends. We also introduce StructureClaw-Bench, an executable benchmark of 150 controlled scenarios spanning standard workflow execution, interactive robustness, and multimodal structural-model reconstruction. A scenario succeeds only when all required artifact- and execution-level assertions pass in a single run. Across ten agent-model configurations, each evaluated on the same 50 standard cases, the average Success Rate rises from 56.8% with the generic-skill baseline to 88.6% with the full automatic workflow. The interactive and multimodal evaluations identify two prominent remaining challenges: safe handling of invalid numerical inputs and fixture-consistent reconstruction of structural models. These findings show that artifact-centered evaluation can expose workflow-level failures that are difficult to identify from final responses alone, providing a more rigorous basis for evaluating and improving structural-engineering agents. The code and benchmark are available at https://github.com/structureclaw/structureclaw.

Community

00

Publication notes

Author note
21 pages, 13 figures