MEGA Hub

GPT-Red: Automated Red Teaming via Self-Play at Scale

Authors

Do you know Eric Wallace?You can claim authorship or link another user.Do you know Christopher A. Choquette-Choo?You can claim authorship or link another user.Do you know Nikhil Kandpal?You can claim authorship or link another user.Do you know Sam Toyer?You can claim authorship or link another user.Do you know Dylan Hunn?You can claim authorship or link another user.Do you know Stephanie Lin?You can claim authorship or link another user.Do you know Yuxin Wen?You can claim authorship or link another user.Do you know Xiangyu Qi?You can claim authorship or link another user.Do you know Christopher Wolff?You can claim authorship or link another user.Do you know Zizhao Wang?You can claim authorship or link another user.Do you know Milad Nasr?You can claim authorship or link another user.Do you know Sicheng Zhu?You can claim authorship or link another user.Do you know Chuan Guo?You can claim authorship or link another user.Do you know Juan Felipe Cerón Uribe?You can claim authorship or link another user.Do you know Kaiwen Wang?You can claim authorship or link another user.Do you know Aiden Low?You can claim authorship or link another user.Do you know Kai Xiao?You can claim authorship or link another user.Do you know Kai Chen?You can claim authorship or link another user.

Abstract

We introduce GPT-Red, an automated red-teaming agent that is trained to discover novel prompt injection attacks against frontier LLMs. The goal of this model is to evaluate and improve the robustness of our production systems. To this end, we use it to adversarially train GPT-5.6, our most robust model to prompt injections to date. To create GPT-Red, we design a scalable self-play algorithm where the model is tasked with attacking a diverse population of simultaneously-trained defender agents. We train the model on realistic red-teaming environments using compute on the same scale as some of our largest RL post-training runs, making it the single-largest LLM safety training run ever documented. GPT-Red excels at red-teaming: it reliably breaks our past models up to GPT-5.5, it finds more successful attacks than human red-teamers, and it generalizes to held-out environments, defender models, and harnesses. In the future, we expect that as we improve the robustness of each new GPT model, it will in turn will provide better learning signal for even stronger red-teamer agents, thus unlocking a self-improvement flywheel.

Community

00