MEGA Hub

ReguSim: Evaluating LLM Agent Rule Grounding in Financial Compliance

Authors

Do you know Yiyang Luo?You can claim authorship or link another user.Do you know Yihang Jiang?You can claim authorship or link another user.Do you know Qijun Xie?You can claim authorship or link another user.Do you know Liang Lan?You can claim authorship or link another user.Do you know Lin Willian Cong?You can claim authorship or link another user.Do you know Anyi Rao?You can claim authorship or link another user.Do you know Yunya Song?You can claim authorship or link another user.

Abstract

LLM agents in financial markets may cite rules yet still submit orders that violate executable constraints or misread surveillance evidence. We introduce ReguSim, a controlled financial-compliance environment, and ReguBench, a target-marked monitoring benchmark, to separate four artifacts: stated reasoning, attempted action, execution enforcement, and monitor evidence. In trader runs with DeepSeek V4 Pro and Gemini 3.5 Flash, visible rules reduce but do not eliminate rejected actions, and incentive or persona framing shifts behavior. A bridge study shows that trader rationales can mislead an independent monitor unless enforcement evidence is shown. In monitoring, simple structured baselines either match or exceed prompt-only LLMs. The results frame financial compliance evaluation as an audit of rule-grounded actions and evidence use, rather than a single compliance score.

Community

00