MEGA Hub

AEGIS: Awareness-Enhanced Guidance for Iterative Safeguard

Authors

Do you know Kyungwon Park?You can claim authorship or link another user.Do you know Sangmin Lee?You can claim authorship or link another user.Do you know Heejae Chon?You can claim authorship or link another user.Do you know Hyungu Kang?You can claim authorship or link another user.

Abstract

Span-level rationales are often assumed to improve controllability in text detoxification, but it remains unclear when such guidance helps and when it introduces trade-offs. We present Awareness-Enhanced Guidance for Iterative Safeguard (AEGIS) as an exploratory framework for studying span-guided multilingual detoxification across English, Mandarin Chinese, and Korean. AEGIS combines span-level detector outputs with frozen generator backbones, allowing harmful spans, intensity labels, and target attributes to be provided as structured guidance during rewriting. Rather than claiming state-of-the-art detoxification performance, we analyze how span guidance affects the balance between toxicity reduction and meaning preservation across generator families, model scales, and languages. Our results suggest that span-guided detoxification is conditionally useful: explicit rationales change the trade-off between toxicity reduction and meaning preservation, but their effects depend strongly on the generator backbone and the linguistic context. These findings highlight both the promise and the limitations of span-level control signals for multilingual detoxification.

Community

00

Publication notes

Author note
11 pages, 3 figures, 9 tables. Preprint