MEGA Hub

Shieldstral

Authors

Do you know Antonia Calvi?You can claim authorship or link another user.Do you know Avinash Sooriyarachchi?You can claim authorship or link another user.Do you know Giada Pistilli?You can claim authorship or link another user.Do you know Guillaume Lample?You can claim authorship or link another user.Do you know Maarten Buyl?You can claim authorship or link another user.Do you know Maximilian Augustin?You can claim authorship or link another user.Do you know Maximilian Müller?You can claim authorship or link another user.Do you know Pierre Stock?You can claim authorship or link another user.Do you know Tom Bewley?You can claim authorship or link another user.Do you know Wassim Bouaziz?You can claim authorship or link another user.Do you know Yimu Pan?You can claim authorship or link another user.

Abstract

We introduce Shieldstral, a 3B-parameter policy-adaptive multimodal safety classifier that matches or outperforms models nearly 7$\times$ its size on text safety benchmarks and sets a new state of the art on multimodal safety classification. Shieldstral formulates content moderation as a binary question-answering task. This simple formulation unifies diverse moderation tasks into a single yes/no problem, enabling heterogeneous safety datasets with divergent taxonomies to be consolidated under one training framework. We present the data construction recipe, covering curation and generation of approximately 54.1M samples and a fine-grained evaluation set to evaluate policy adaptability. Together, these enable a small adaptive model to match or outperform much larger models.

Community

00