A Survey of Red Teaming for Generative Artificial Intelligence
Shizhong Wu , Biyu Zhou , Qirui Tang , Wenjie Xiao , Xiaodan Zhang , Xuehai Tang , Yong Peng , Songlin Hu
Strategic Study of CAE ›› : 1 -21.
With the proliferation of generative artificial intelligence (AI) technologies, particularly large language models(LLMs), their associated safety risks have grown increasingly pressing. As a proactive safety methodology for discovering and evaluating potential threats, red teaming has attracted significant attention in academia, industry, and among policymakers. Distinct from existing reviews that mainly focus on a single model type and emphasize the categorization of attack methods, this study expands the scope to encompass multiple system types, including LLMs, multimodal models, agents, and retrieval-augmented generation (RAG). It establishes a novel classification framework centered on the dual dimensions of “attack generation” and “effect evaluation,” while systematically reviewing representative industrial practices. First, we clarify the core concepts and risk taxonomy of generative AI red teaming. Subsequently, from the perspectives of single-turn and multi-turn interactions, we survey automated red teaming attack techniques targeting LLMs, and categorize the corresponding attack effectiveness detection methods into three major classes: discriminative, generative, and edge-side detection. Furthermore, we explore the novel safety threats and innovative testing paradigms introduced by multimodal models, AI agents, and RAG systems. Finally, we examine the persistent challenges in this field, including the ongoing evolution of attack and defense methods and the imperative for multi-stakeholder governance, and outline directions for future research. This study aims to provide technical references and practical insights for China in constructing a safety assessment system for generative AI and enhancing proactive risk defense capabilities.
red teaming / generative artificial intelligence (AI) / AI safety / large language models / automated testing
Funding project: Chinese Academy of Engineering project “Strategic Research on Artificial Intelligence Safety Prevention in the Cyber and Information Domain”(2025-XBZD-08)
/
| 〈 |
|
〉 |