MEGA Hub

Pragmatic Attack Surface: Vulnerabilities of Implicit Context in Large Language Models

Authors

Do you know Bocheng Chen?You can claim authorship or link another user.Do you know Han Zi?You can claim authorship or link another user.Do you know Roucheng Ou?You can claim authorship or link another user.Do you know Yawei Liu?You can claim authorship or link another user.Do you know Minyue Chen?You can claim authorship or link another user.Do you know Zimo Qi?You can claim authorship or link another user.Do you know Rongrong Wang?You can claim authorship or link another user.Do you know Guangliang Liu?You can claim authorship or link another user.

Abstract

In the era of large language models (LLMs), attackers often manipulate natural language to elicit unsafe or harmful outputs, creating a new natural language attack surface unique to LLM-based systems, where attacks directly exploit explicit linguistic cues in user prompts to bypass the safety mechanism of LLMs. However, such attacks can often be mitigated by existing safety alignment algorithms. On the other hand, human language is inherently grounded in pragmatics, necessitating typical context to interpret language, e.g., world knowledge, social norms. However, such contexts are often implicit because they are not directly expressed in human language and are not sufficiently leveraged in safety alignment, creating a fundamental mismatch between human language interpretation and safety alignment approaches. In this paper, we demonstrate that this mismatch exposes vulnerabilities in LLMs. We refer to this vulnerability as the pragmatic attack surface, which can be exploited to achieve high attack success rates. The experimental results demonstrate that our proposed approach outperforms baseline attack methods across various open-source and closed-source models by a substantial margin.

Community

00