MEGA Hub

ToolHazard: Scaling Adversarial Environments for Security Evaluation and Alignment of LLM-based Agents

Authors

Do you know Yutao Mou?You can claim authorship or link another user.Do you know Pengfei Yang?You can claim authorship or link another user.Do you know Zhe Yin?You can claim authorship or link another user.Do you know Zhangchi Xue?You can claim authorship or link another user.Do you know Xiaotian Luan?You can claim authorship or link another user.Do you know Dingyao Yu?You can claim authorship or link another user.Do you know Tong Zhang?You can claim authorship or link another user.Do you know Shikun Zhang?You can claim authorship or link another user.Do you know Wei Ye?You can claim authorship or link another user.

Abstract

Large language model (LLM) agents integrated with external tools are vulnerable to indirect prompt injections embedded in environmental states. However, existing studies largely rely on manually implemented or reused environments, stochastic LLM-based tool simulation, and predefined injection locations, limiting scalable security research across broader domains. To bridge this gap, we propose **ToolHazard**, a scalable adversarial environment synthesis framework that reduces human engineering and supports expansion with additional seed domains and compute. Through an Environment Simulator, an Attacker Agent, and a User Simulator, ToolHazard synthesizes executable stateful environments, discovers viable injection points and generates environment-specific payloads, and constructs state-grounded long-horizon tasks. Based on ToolHazard, we build **ToolHazard-Bench** for stress-testing agents under complex workflows and diverse environmental attacks. Experiments reveal substantial agent vulnerabilities and show that injection timing and placement affect attack effectiveness. Moreover, ToolHazard-generated alignment data improves security on both ToolHazard-Bench and AgentDojo while preserving benign task utility.

Community

00

Publication notes

Author note
Work in Progress