MEGA Hub

Detecting CSAM Text-to-Image LoRAs From Weights

Authors

Do you know David Demitri Africa?You can claim authorship or link another user.Do you know Cate Heine?You can claim authorship or link another user.Do you know Nadine Staes-Polet?You can claim authorship or link another user.Do you know Kimberly Mai?You can claim authorship or link another user.

Abstract

Low-rank adaptation (LoRA) fine-tuning has made it cheap and easy to customize open-weight image generation models for specific tasks, including the production of child sexual abuse material (CSAM). Existing moderation relies on metadata or generated outputs, but metadata can be deceptive and generating outputs may itself be unacceptable or illegal. We show that a safer signal lives in the weights. The top-left singular vectors of a LoRA's updates form a compact, inference-free fingerprint ($u_1$) of its strongest learned change. Using human-subject age as a benign proxy for CSAM, we find that $u_1$ identifies what a LoRA was trained on, generalizes across base models, and abstains on unrelated benign content. The signal is robust to additive weight noise, rescaling, and precision reduction. These results indicate that harmful LoRAs could be screened directly from their weights without relying on metadata or generating harmful outputs.

Community

00