MEGA Hub

Introducing the Privacy-HSD Trade-off: Hate Speech Detection, but not at the Cost of Privacy

Authors

Do you know Stephen Meisenbacher?You can claim authorship or link another user.Do you know Vlad Garbuz?You can claim authorship or link another user.Do you know Chirill Donos?You can claim authorship or link another user.Do you know Maxim Dnestreanschii?You can claim authorship or link another user.Do you know Gabriel Creanga?You can claim authorship or link another user.Do you know Andreea-Elena Bodea?You can claim authorship or link another user.Do you know Thomas Lampert?You can claim authorship or link another user.Do you know Jana Diesner?You can claim authorship or link another user.

Abstract

Hate speech is a real and timely threat that affects a large portion of online users, especially youth and minority groups. While building reliable and robust automatic hate speech detection (HSD) systems is paramount, we argue that this must also be balanced with the individual right to privacy. Exploring the intersection of HSD and privacy, we demonstrate that HSD systems might unintentionally achieve performance at the cost of encoding authorship, posing a threat to privacy. Building on these findings, we establish the notion of a privacy-HSD trade-off, which demands a careful balance. We benchmark a series of text privatization methods, as well as our newly proposed domain-specific AgnoSpeech technique, showing that balancing privacy and HSD is difficult but feasible. The findings make a strong case for more research on the trade-offs between privacy and HSD, both of which have tangible implications for the safeguarding of online participation.

Community

00

Publication notes

Author note
13 pages, 1 figure, 3 tables. Accepted to WOAH 2026