生成式人工智能红队测试技术综述
A Survey of Red Teaming for Generative Artificial Intelligence
随着大语言模型等生成式人工智能技术的广泛应用,其伴生的安全风险日益凸显。红队测试作为主动发现和评估潜在威胁的关键手段,已在学术界、产业界及政策制定部门引起广泛重视。区别于现有综述多聚焦单一模型类型且偏重攻击方法归纳,本文将研究对象从大语言模型扩展至多模态模型、智能体及检索增强生成等多类系统,以“攻击生成”与“效果评判”为主线构建新的分类视角,系统梳理了国内外代表性产业实践。首先,厘清生成式人工智能红队测试的核心概念与风险分类体系;其次,从单轮与多轮交互视角梳理面向大语言模型的自动化红队攻击技术,并将攻击效果评判方法归纳为判别式、生成式与端侧三大类;进而,探讨多模态模型、智能体及检索增强生成系统所带来的新型安全威胁与测试范式创新;最后,深入剖析该领域在攻防技术持续演进及多方协同治理等方面面临的核心挑战,并对未来研究方向进行展望。本研究旨在为我国构建生成式人工智能安全测评体系、增强风险主动防御能力提供系统化技术参考与实践支撑。
With the proliferation of generative artificial intelligence (AI) technologies, particularly large language models(LLMs), their associated safety risks have grown increasingly pressing. As a proactive safety methodology for discovering and evaluating potential threats, red teaming has attracted significant attention in academia, industry, and among policymakers. Distinct from existing reviews that mainly focus on a single model type and emphasize the categorization of attack methods, this study expands the scope to encompass multiple system types, including LLMs, multimodal models, agents, and retrieval-augmented generation (RAG). It establishes a novel classification framework centered on the dual dimensions of “attack generation” and “effect evaluation,” while systematically reviewing representative industrial practices. First, we clarify the core concepts and risk taxonomy of generative AI red teaming. Subsequently, from the perspectives of single-turn and multi-turn interactions, we survey automated red teaming attack techniques targeting LLMs, and categorize the corresponding attack effectiveness detection methods into three major classes: discriminative, generative, and edge-side detection. Furthermore, we explore the novel safety threats and innovative testing paradigms introduced by multimodal models, AI agents, and RAG systems. Finally, we examine the persistent challenges in this field, including the ongoing evolution of attack and defense methods and the imperative for multi-stakeholder governance, and outline directions for future research. This study aims to provide technical references and practical insights for China in constructing a safety assessment system for generative AI and enhancing proactive risk defense capabilities.
| [1] |
OpenAI,Achiam J,Adler S,et al. GPT-4 technical report[PP/OL]. V6. arXiv(2024-03-04)[2026-02-10]. https://doi.org/10.48550/arXiv.2303.08774. |
| [2] |
Touvron H,Lavril T,Izacard G,et al. LLaMA:Open and efficient foundation language models[PP/OL]. arXiv (2023-02-27)[2026-02-10]. https://doi.org/10.48550/arXiv.2302.13971. |
| [3] |
Team G,Anil R,Borgeaud S,et al. Gemini:A family of highly capable multimodal models[PP/OL]. V5. arXiv (2025-05-09)[2026-02-10]. https://doi.org/10.48550/arXiv.2312.11805. |
| [4] |
Wei A,Haghtalab N,Steinhardt J. Jailbroken:How does LLM safety training fail?[C]//The Thirty-Seventh Annual Conference on Neural Information Processing Systems,2023:80079-80110. |
| [5] |
Zhang Z X,Lei L Q,Wu L D,et al. SafetyBench:Evaluating the safety of large language models[C]//The 62nd Annual Meeting of the Association for Computational Linguistics,2024:15537-15553. |
| [6] |
李希陶,吴江,郑庆华, 大语言模型越狱攻击:模型、根因及其攻防演化[J]. 中国科学:信息科学,2025,55(6):1372-1405. |
| [7] |
Li X T,Wu J,Zheng Q H,et al. Jailbreaking large language models:Models,origins,and evolution of attacks and defenses[J]. Scientia Sinica Informationis,2025,55(6):1372-1405. |
| [8] |
Bengio Y,Hinton G,Yao A,et al. Managing extreme AI risks amid rapid progress[J]. Science,2024,384(6698):842-845. |
| [9] |
Center for AI Safety. Statement on AI risk[EB/OL]. (2023-05-30)[2026-02-04]. https://aistatement.com/. |
| [10] |
杨伟平,程豪杰,周百顺, M3-SafetyBench:多领域多场景多维度的大语言模型安全评估体系[J]. 中国科学:信息科学,2025,55(11):2923-2940. |
| [11] |
Yang W P,Cheng H J,Zhou B S,et al. M3-SafetyBench:A comprehensive benchmark for evaluating the safety of large language models across multiple domains,scenarios,and dimensions[J]. Scientia Sinica Informationis,2025,55(11):2923-2940. |
| [12] |
Wei J,Tay Y,Bommasani R,et al. Emergent abilities of large language models[PP/OL]. V2. arXiv(2022-10-26)[2026-02-10]. https://doi.org/10.48550/arXiv.2206.07682. |
| [13] |
Schaeffer R,Miranda B,Koyejo S. Are emergent abilities of large language models a mirage?[PP/OL]. V2. arXiv (2023-05-22)[2026-02-10]. https://doi.org/10.48550/arXiv.2304.15004. |
| [14] |
National Institute of Standards and Technology. Artificial intelligence risk management framework:Generative artificial intelligence profile:NIST AI 600-1[R]. Gaithersburg:National Institute of Standards and Technology,2024:18,50. |
| [15] |
Smuha N A. Regulation 2024/1689 of the eur. Parl. & council of June 13,2024 (EU artificial intelligence act)[J]. International Legal Materials,2025,64(5):1234-1381. |
| [16] |
Bai Y T,Kadavath S,Kundu S,et al. Constitutional AI:Harmlessness from AI feedback[PP/OL]. arXiv(2022-12-15)[2026-02-10]. https://doi.org/10.48550/arXiv.2212.08073. |
| [17] |
Anthropic. Challenges in red teaming AI systems[EB/OL]. (2024-06-12)[2026-02-10]. https://www.anthropic.com/research/challenges-in-red-teaming-ai-systems. |
| [18] |
Reuel A,Hardy A,Smith C,et al. BetterBench:Assessing AI benchmarks,uncovering issues,and establishing best practices[C]//The Thirty-Eighth Annual Conference on Neural Information Processing Systems,2024:21763-21813. |
| [19] |
Eriksson M,Purificato E,Noroozian A,et al. Can we trust AI benchmarks? An interdisciplinary review of current issues in AI evaluation[C]//The AAAI/ACM Conference on AI,Ethics,and Society,2025:850-864. |
| [20] |
中国信息通信研究院人工智能研究所,人工智能关键技术和应用评测工业和信息化部重点实验室. 大模型基准测试体系研究报告(2024年)[R]. 北京:中国信息通信研究院,2024:2. |
| [21] |
Artificial Intelligence Research Institute of China Academy of Information and Communications Technology,Ministry of Industry and Information Technology Key Laboratory of Artificial Intelligence Key Technologies and Applications Evaluation. Large model benchmarking system research report (2024)[R]. Beijing:China Academy of Information and Communications Technology,2024:2. |
| [22] |
Eriksson M,Purificato E,Noroozian A,et al. AI benchmarks:Interdisciplinary issues and policy considerations[C]//ICML Workshop on Technical AI Governance (TAIG),2025:1-13. |
| [23] |
Raji I D,Bender E M,Paullada A,et al. AI and the everything in the whole wide world benchmark[PP/OL]. arXiv (2021-11-26)[2026-02-10]. https://doi.org/10.48550/arXiv.2111.15366. |
| [24] |
Scarfone K,Souppaya M,Cody A,et al. Technical guide to information security testing and assessment:NIST Special Publication 800-115[R]. Gaithersburg:National Institute of Standards and Technology,2008:2-1-2-5,5-2-5-6. |
| [25] |
Strom B E,Applebaum A,Miller D P,et al. MITRE ATT&CK:Design and philosophy:MP180360R1[R]. McLean:The MITRE Corporation,2020:3. |
| [26] |
Applebaum A,Miller D,Strom B,et al. Intelligent,automated red team emulation[C]//The 32nd Annual Conference on Computer Security Applications,2016:363-373. |
| [27] |
European Central Bank. TIBER-EU framework:How to implement the European framework for threat intelligence-based ethical red teaming[EB/OL]. (2025-02-11)[2026-08-18]. https://www.ecb.europa.eu/pub/pdf/other/ecb.tiber_eu_framework_2025~b32eff9a10.en.pdf. |
| [28] |
Sinha A,Lucassen J,Grimes K,et al. What can generative AI red-teaming learn from cyber red-teaming?[R]. Pittsburgh:Software Engineering Institute,Carnegie Mellon University,2025:10-12,19-20. |
| [29] |
Anthropic. Progress from our frontier red team[EB/OL]. (2025-03-19)[2026-01-20]. https://www.anthropic.com/news/strategic-warning-for-ai-risk-progress-and-insights-from-our-frontier-red-team. |
| [30] |
Abuadbba A,Moore K,Goel D,et al. From promise to peril:Rethinking cybersecurity red and blue teaming in the age of LLMs[J]. IEEE Security & Privacy,2026,24(2):53-63. |
| [31] |
Jabbar M S,Al-Azani S,Alotaibi A,et al. Red teaming large language models:A comprehensive review and critical analysis[J]. Information Processing & Management,2025,62(6):104239. |
| [32] |
段俊贤,刘思雨,关霁洋, 生成式可视媒体鉴别与安全[J]. 中国科学:信息科学,2025,55(8):1925-1949. |
| [33] |
Duan J X,Liu S Y,Guan J Y,et al. Survey on generative visual media detection and security[J]. Scientia Sinica Informationis,2025,55(8):1925-1949. |
| [34] |
Lin L Z,Mu H L,Zhai Z N,et al. Against the Achilles’ heel:A survey on red teaming for generative models[J]. Journal of Artificial Intelligence Research,2025,82:687-775. |
| [35] |
Booth H,Souppaya M,Vassilev A,et al. Secure software development practices for generative AI and dual-use foundation models:An SSDF community profile:NIST SP 800-218A[R]. Gaithersburg:National Institute of Standards and Technology,2024:22. |
| [36] |
Japan AI Safety Institute. Guide to red teaming methodology on AI safety (version 1.00)[EB/OL]. (2024-09-25)[2026-08-18]. https://aisi.go.jp/assets/pdf/ai_safety_RT_v1.00_en.pdf. |
| [37] |
OpenAI. Advancing red teaming with people and AI[EB/OL]. (2024-11-21)[2026-01-20]. https://openai.com/index/advancing-red-teaming-with-people-and-ai/. |
| [38] |
Meta AI. Expanding our open source large language models responsibly[EB/OL]. (2024-07-23)[2026-01-20]. https://ai.meta.com/blog/meta-llama-3-1-ai-responsibility/. |
| [39] |
Google. Google’s AI red team:The ethical hackers making AI safer[EB/OL]. (2023-07-19)[2026-01-20]. https://blog.google/innovation-and-ai/technology/safety-security/googles-ai-red-team-the-ethical-hackers-making-ai-safer/. |
| [40] |
NVIDIA. Defining LLM red teaming[EB/OL]. (2025-02-25)[2026-01-20]. https://developer.nvidia.com/blog/defining-llm-red-teaming/. |
| [41] |
Microsoft. Microsoft AI red team building future of safer AI[EB/OL]. (2023-08-07)[2026-01-20]. https://www.microsoft.com/en-us/security/blog/2023/08/07/microsoft-ai-red-team-building-future-of-safer-ai/. |
| [42] |
IBM. What is red teaming for generative AI[EB/OL]. (2024-04-11)[2026-01-20]. https://research.ibm.com/blog/what-is-red-teaming-gen-AI. |
| [43] |
OpenAI. OpenAI red teaming network[EB/OL]. (2023-09-19)[2026-01-20]. https://openai.com/index/red-teaming-network/. |
| [44] |
Verma A,Krishna S,Gehrmann S,et al. Operationalizing a threat model for red-teaming large language models (LLMs)[J/OL]. Transactions on Machine Learning Research,2024-07-21. https://openreview.net/forum?id=sSAp8ITBpC. |
| [45] |
Verma A, Krishna S, Gehrmann S, et al. Operationalizing a threat model for red-teaming large language models (LLMs) [PP/OL]. V2. arXiv (2025-07-10)[2026-01-18]. https://doi.org/10.48550/arXiv.2407.14937. |
| [46] |
United Nations Educational,Scientific,Cultural Organization. Recommendation on the ethics of artificial intelligence[M]. Paris:UNESCO (United Nations Educational,Scientific and Cultural Organization),2022:25-38. |
| [47] |
Bengio Y,Mindermann S,Privitera D,et al. International AI safety report:DSIT 2025/001[R/OL]. London:Department for Science,Innovation and Technology,2025 (2025-01-29)[2026-01-18]. https://www.gov.uk/government/publications/international-ai-safety-report-2025. |
| [48] |
Weidinger L,Uesato J,Rauh M,et al. Taxonomy of risks posed by language models[C]//The 2022 ACM Conference on Fairness,Accountability,and Transparency (FAccT '22),2022:214-229. |
| [49] |
Hendrycks D,Mazeika M,Woodside T. An overview of catastrophic AI risks[PP/OL]. V6. arXiv (2023-10-09)[2026-01-20]. https://doi.org/10.48550/arXiv.2306.12001. |
| [50] |
Derczynski L,Kirk H R,Balachandran V,et al. Assessing language model deployment with risk cards[PP/OL]. arXiv (2023-03-31)[2026-01-18]. https://doi.org/10.48550/arXiv.2303.18190. |
| [51] |
Yu Y M,Liu Y R,Zhang J,et al. Understanding generative AI risks for youth:A taxonomy based on empirical data[PP/OL]. V2. arXiv (2025-02-25)[2026-01-20]. https://doi.org/10.48550/arXiv.2502.16383. |
| [52] |
全国网络安全标准化技术委员会秘书处. 人工智能安全标准体系(V1.0)(征求意见稿)[R]. 北京:全国网络安全标准化技术委员会秘书处,2025:1. |
| [53] |
Secretariat of National Information Security Standardization Technical Committee. Artificial intelligence security standard system (version 1.0) (draft for comments)[R]. Beijing:Secretariat of National Information Security Standardization Technical Committee,2025:1. |
| [54] |
Qi X Y,Zeng Y,Xie T H,et al. Fine-tuning aligned language models compromises safety,even when users do not intend to![C]//The Twelfth International Conference on Learning Representations,2024. |
| [55] |
Sun H,Zhang Z X,Deng J W,et al. Safety assessment of Chinese large language models[PP/OL]. arXiv (2023-04-20)[2026-01-20]. https://doi.org/10.48550/arXiv.2304.10436. |
| [56] |
He J Y,Feng W T,Min Y S,et al. Control risk for potential misuse of artificial intelligence in science[PP/OL]. arXiv (2023-12-11)[2026-01-20]. https://doi.org/10.48550/arXiv.2312.06632. |
| [57] |
Hua W Y,Yang X J,Jin M Y,et al. TrustAgent:Towards safe and trustworthy LLM-based agents[C]// The 2024 Conference on Empirical Methods in Natural Language Processing,2024:10000-10016. |
| [58] |
Papernot N,McDaniel P,Sinha A,et al. SoK:Security and privacy in machine learning[C]//2018 IEEE European Symposium on Security and Privacy (EuroS&P),2018:399-414. |
| [59] |
Yao Y F,Duan J H,Xu K D,et al. A survey on large language model (LLM) security and privacy:The good,the bad,and the ugly[J]. High-Confidence Computing,2024,4(2):100211. |
| [60] |
Liu H C,Wang Y Q,Fan W Q,et al. Trustworthy AI:A computational perspective[J]. ACM Transactions on Intelligent Systems and Technology,2023,14(1):1-59. |
| [61] |
Vassilev A,Oprea A,Fordyce A,et al. Adversarial machine learning:A taxonomy and terminology of attacks and mitigations:NIST AI 100-2e2023[R]. Gaithersburg:National Institute of Standards and Technology,2024:6-13. |
| [62] |
Shayegani E,Al Mamun M A,Fu Y,et al. Survey of vulnerabilities in large language models revealed by adversarial attacks[PP/OL]. arXiv (2023-10-16)[2026-01-20]. https://doi.org/10.48550/arXiv.2310.10844. |
| [63] |
OpenAI. GPT-4o system card[EB/OL]. (2024-08-08)[2026-01-18]. https://cdn.openai.com/gpt-4o-system-card.pdf. |
| [64] |
Purpura A,Wadhwa S,Zymet J,et al. Building safe GenAI applications:An end-to-end overview of red teaming for large language models[C]//The 5th Workshop on Trustworthy NLP (TrustNLP 2025),2025:335-350. |
| [65] |
Shen X Y,Chen Z Y,Backes M,et al. “Do anything now”:Characterizing and evaluating in-the-wild jailbreak prompts on large language models[C]//The 2024 ACM SIGSAC Conference on Computer and Communications Security,2024:1671-1685. |
| [66] |
Liu Y,Deng G L,Xu Z Z,et al. A hitchhiker’s guide to jailbreaking ChatGPT via prompt engineering[C]//The 4th International Workshop on Software Engineering and AI for Data Quality in Cyber-Physical Systems/Internet of Things,2024:12-21. |
| [67] |
Wang J X,Liu Z C,Park K H,et al. Adversarial demonstration attacks on large language models[PP/OL]. V2. arXiv (2023-10-14)[2026-01-20]. https://doi.org/10.48550/arXiv.2305.14950. |
| [68] |
Zhou X Y,Qiang Y,Zade S Z,et al. Hijacking large language models via adversarial in-context learning[PP/OL]. V3. arXiv (2025-05-29)[2026-08-18]. https://doi.org/10.48550/arXiv.2311.09948. |
| [69] |
Anil C,Durmus E,Panickssery N,et al. Many-shot jailbreaking[C]//The Thirty-Eighth Annual Conference on Neural Information Processing Systems,2024:129696-129742. |
| [70] |
Yong Z X,Menghini C,Bach S H. Low-resource languages jailbreak GPT-4[PP/OL]. V2. arXiv(2024-01-27)[2026-01-20]. https://doi.org/10.48550/arXiv.2310.02446. |
| [71] |
Li J,Liu Y,Liu C Y,et al. A cross-language investigation into jailbreak attacks in large language models[PP/OL]. arXiv (2024-01-30)[2026-01-20]. https://doi.org/10.48550/arXiv.2401.16765. |
| [72] |
Jiang F Q,Xu Z C,Niu L Y,et al. ArtPrompt:ASCII art-based jailbreak attacks against aligned LLMs[C]//The 62nd Annual Meeting of the Association for Computational Linguistics,2024:15157-15173. |
| [73] |
Yuan Y L,Jiao W X,Wang W X,et al. GPT-4 is too smart to be safe:Stealthy chat with LLMs via cipher[C]//The Twelfth International Conference on Learning Representations,2024. |
| [74] |
Ren Q B,Gao C,Shao J,et al. CodeAttack:Revealing safety generalization challenges of large language models via code completion[C]//The 62nd Annual Meeting of the Association for Computational Linguistics,2024:11437-11452. |
| [75] |
Perez E,Huang S,Song F,et al. Red teaming language models with language models[C]//The 2022 Conference on Empirical Methods in Natural Language Processing,2022:3419-3448. |
| [76] |
Liu X G,Xu N,Chen M H,et al. AutoDAN:Generating stealthy jailbreak prompts on aligned large language models[C]//The Twelfth International Conference on Learning Representations,2024:1-21. |
| [77] |
Deng G L,Liu Y,Li Y K,et al. MASTERKEY:Automated jailbreaking of large language model chatbots[C]//The 2024 Network and Distributed System Security Symposium (NDSS),2024:1-18. |
| [78] |
Yu J H,Lin X W,Yu Z,et al. LLM-fuzzer:Scaling assessment of large language model jailbreaks[C]//The 33rd USENIX Security Symposium (USENIX Security 24),2024:4657-4674. |
| [79] |
Jha P,Arora A,Ganesh V. LLM stinger:Jailbreaking LLMs using RL fine-tuned LLMs(Student Abstract)[C]//The AAAI Conference on Artificial Intelligence,2025:29393-29395. |
| [80] |
Lee S,Ni S W,Wei C,et al. xJailbreak:Representation space guided reinforcement learning for interpretable LLM jailbreaking[PP/OL]. V2. arXiv (2025-01-30)[2026-01-20]. https://doi.org/10.48550/arXiv.2501.16727. |
| [81] |
Li H Y,Ye J W,Wu J,et al. JailPO:A novel black-box jailbreak framework via preference optimization against aligned LLMs[C]//The AAAI Conference on Artificial Intelligence,2025:27419-27427. |
| [82] |
Yang X K,Tang X H,Han J Z,et al. The dark side of trust:Authority citation-driven jailbreak attacks on large language models[PP/OL]. arXiv(2024-11-18)[2026-01-20]. https://doi.org/10.48550/arXiv.2411.11407. |
| [83] |
Yang X K,Zhou B Y,Tang X H,et al. Exploiting synergistic cognitive biases to bypass safety in LLMs[C]//The AAAI Conference on Artificial Intelligence,2026:2200-2208. |
| [84] |
Zou A,Wang Z F,Carlini N,et al. Universal and transferable adversarial attacks on aligned language models[PP/OL]. V2. arXiv (2023-12-20)[2026-01-20]. https://doi.org/10.48550/arXiv.2307.15043. |
| [85] |
Jia X J,Pang T Y,Du C,et al. Improved techniques for optimization-based jailbreaking on large language models[C]//The Thirteenth International Conference on Learning Representations,2025:1-22. |
| [86] |
Liao Z Y,Sun H. AmpleGCG:Learning a universal and transferable generative model of adversarial suffixes for jailbreaking both open and closed LLMs[C]//The First Conference on Language Modeling,2024:1-33. |
| [87] |
Guo X G,Yu F X,Zhang H,et al. COLD-attack:Jailbreaking LLMs with stealthiness and controllability[C]//The 41st International Conference on Machine Learning,2024:16974-17002. |
| [88] |
Sitawarin C,Mu N,Wagner D,et al. PAL:Proxy-guided black-box attack on large language models[PP/OL]. arXiv (2024-02-15)[2026-01-20]. https://doi.org/10.48550/arXiv.2402.09674. |
| [89] |
Chao P,Robey A,Dobriban E,et al. Jailbreaking black box large language models in twenty queries[C]//2025 IEEE Conference on Secure and Trustworthy Machine Learning (SaTML),2025:23-42. |
| [90] |
Mehrotra A,Zampetakis M,Kassianik P,et al. Tree of attacks:Jailbreaking black-box LLMs automatically[C]//The Thirty-Eighth Annual Conference on Neural Information Processing Systems,2024:61065-61105. |
| [91] |
Zeng Y,Lin H P,Zhang J W,et al. How johnny can persuade LLMs to jailbreak them:Rethinking persuasion to challenge AI safety by humanizing LLMs[C]//The 62nd Annual Meeting of the Association for Computational Linguistics,2024:14322-14350. |
| [92] |
Li X,Zhou Z K,Zhu J N,et al. DeepInception:Hypnotize large language model to be jailbreaker[PP/OL]. V5. arXiv (2024-11-28)[2026-01-20]. https://doi.org/10.48550/arXiv.2311.03191. |
| [93] |
Weng Z X,Jin X L,Jia J Y,et al. Foot-in-the-door:A multi-turn jailbreak for LLMs[C]//The 2025 Conference on Empirical Methods in Natural Language Processing,2025:1939-1950. |
| [94] |
Zhou Z H,Xiang J Y,Chen H P,et al. Speak out of turn:Safety vulnerability of large language models in multi-turn dialogue[PP/OL]. V2. arXiv (2024-10-30)[2026-01-20]. https://doi.org/10.48550/arXiv.2402.17262. |
| [95] |
Yu E X,Li J,Liao M,et al. CoSafe:Evaluating large language model safety in multi-turn dialogue coreference[C]//The 2024 Conference on Empirical Methods in Natural Language Processing,2024:17494-17508. |
| [96] |
Jiang Y F,Aggarwal K,Laud T,et al. Red Queen:Exposing latent multi-turn risks in large language models[C]//The 63rd Annual Meeting of the Association for Computational Linguistics,2025:25554-25591. |
| [97] |
Zhang J C,Zhou Y,Liu Y X,et al. Holistic automated red teaming for large language models through top-down test case generation and multi-turn interaction[C]//The 2024 Conference on Empirical Methods in Natural Language Processing,2024:13711-13736. |
| [98] |
Yang X K,Zhou B Y,Tang X H,et al. Chain of attack:Hide your intention through multi-turn interrogation[C]//The 63rd Annual Meeting of the Association for Computational Linguistics,2025:9881-9901. |
| [99] |
Wang F X,Duan R J,Xiao P,et al. MRJ-agent:An effective jailbreak agent for multi-round dialogue[PP/OL]. V2. arXiv (2025-01-07)[2026-01-20]. https://doi.org/10.48550/arXiv.2411.03814. |
| [100] |
Xiong X Q,Li O X,Liu Z,et al. TROJail:Trajectory-level optimization for multi-turn large language model jailbreaks with process rewards[C]//The 64th Annual Meeting of the Association for Computational Linguistics,2026:48086-48109. |
| [101] |
Guo W Y,Li J,Wang W Y,et al. MTSA:Multi-turn safety alignment for LLMs through multi-round red-teaming[C]//The 63rd Annual Meeting of the Association for Computational Linguistics,2025:26424-26442. |
| [102] |
Hartvigsen T,Gabriel S,Palangi H,et al. ToxiGen:A large-scale machine-generated dataset for adversarial and implicit hate speech detection[C]//The 60th Annual Meeting of the Association for Computational Linguistics,2022:3309-3326. |
| [103] |
Lees A,Tran V Q,Tay Y,et al. A new generation of perspective API:Efficient multilingual character-level transformers[C]//The 28th ACM SIGKDD Conference on Knowledge Discovery and Data Mining,2022:3197-3207. |
| [104] |
Hosseini H,Kannan S,Zhang B S,et al. Deceiving Google’s perspective API built for detecting toxic comments[PP/OL]. arXiv (2017-02-27)[2026-01-20]. https://doi.org/10.48550/arXiv.1702.08138. |
| [105] |
Markov T,Zhang C,Agarwal S,et al. A holistic approach to undesired content detection in the real world[C]//The AAAI Conference on Artificial Intelligence,2023:15009-15018. |
| [106] |
OpenAI. Moderation[EB/OL]. [2026-01-22]. https://platform.openai.com/docs/guides/moderation. |
| [107] |
Mazeika M,Phan L,Yin X W,et al. HarmBench:A standardized evaluation framework for automated red teaming and robust refusal[C]//The 41st International Conference on Machine Learning,2024:35181-35224. |
| [108] |
Wang Y X,Li H N,Han X D,et al. Do-not-answer:A dataset for evaluating safeguards in LLMs[PP/OL]. V2. arXiv (2023-09-04)[2026-01-20]. https://doi.org/10.48550/arXiv.2308.13387. |
| [109] |
Hu D,Bao Y N,Wei L W,et al. Supervised adversarial contrastive learning for emotion recognition in conversations[C]//The 61st Annual Meeting of the Association for Computational Linguistics,2023:10835-10852. |
| [110] |
Hu D,Wei L W,Liu Y X,et al. Structured probabilistic coding[C]//The AAAI Conference on Artificial Intelligence,2024:12491-12501. |
| [111] |
Chen Z Y,Yu H M,Wu X,et al. Libra:Large Chinese-based safeguard for AI content[C]//Natural Language Processing and Chinese Computing:14th National CCF Conference,2025:567-580. |
| [112] |
Inan H,Upasani K,Chi J F,et al. Llama guard:LLM-based input-output safeguard for human-AI conversations[PP/OL]. arXiv (2023-12-07)[2026-01-20]. https://doi.org/10.48550/arXiv.2312.06674. |
| [113] |
Zeng W J,Liu Y C,Mullins R,et al. ShieldGemma:Generative AI content moderation based on gemma[PP/OL]. V2. arXiv (2024-08-04)[2026-01-20]. https://doi.org/10.48550/arXiv.2407.21772. |
| [114] |
Souly A,Lu Q Y,Bowen D,et al. A StrongREJECT for empty jailbreaks[C]//The Thirty-Eighth Annual Conference on Neural Information Processing Systems,2024:125416-125440. |
| [115] |
Li L J,Dong B W,Wang R H,et al. SALAD-bench:A hierarchical and comprehensive safety benchmark for large language models[C]//The 62nd Annual Meeting of the Association for Computational Linguistics,2024:3923-3954. |
| [116] |
Luo H C,Gu J D,Liu F Y,et al. An image is worth 1000 lies:Adversarial transferability across prompts on vision-language models[C]//The Twelfth International Conference on Learning Representations,2024. |
| [117] |
Dou Z H,Hu X,Yang H B,et al. Adversarial attacks to multi-modal models[C]//The 1st ACM Workshop on Large AI Systems and Models with Privacy and Safety Analysis,2024:35-46. |
| [118] |
Yin Z Y,Ye M C,Zhang T R,et al. VLATTACK:Multimodal adversarial attacks on vision-language tasks via pre-trained models[C]//The Thirty-Seventh Annual Conference on Neural Information Processing Systems,2023:52936-52956. |
| [119] |
Cheng H,Xiao E J,Yang J Y,et al. Transfer attack for bad and good:Explain and boost adversarial transferability across multimodal large language models[C]//The 33rd ACM International Conference on Multimedia,2025:5010-5019. |
| [120] |
Ye M,Rong X K,Huang W K,et al. A survey of safety on large vision-language models:Attacks,defenses and evaluations[PP/OL]. arXiv (2025-02-14)[2026-01-20]. https://doi.org/10.48550/arXiv.2502.14881. |
| [121] |
Gong Y H,Ran D S,Liu J M,et al. Figstep:Jailbreaking large vision-language models via typographic visual prompts[C]//The AAAI Conference on Artificial Intelligence,2024:18748-18756. |
| [122] |
Bagdasaryan E,Hsieh T Y,Nassi B,et al. Abusing images and sounds for indirect instruction injection in multi-modal LLMs[PP/OL]. V4. arXiv (2023-10-03)[2026-01-20]. https://doi.org/10.48550/arXiv.2307.10490. |
| [123] |
Huang S,Papernot N,Goodfellow I,et al. Adversarial attacks on neural network policies[PP/OL]. arXiv (2017-02-08)[2026-01-20]. https://doi.org/10.48550/arXiv.1702.02284. |
| [124] |
Microsoft. CVE-2025-32711:M365 copilot information disclosure vulnerability[EB/OL]. (2025-06-11)[2026-01-18]. https://msrc.microsoft.com/update-guide/vulnerability/CVE-2025-32711. |
| [125] |
Díaz S,Kern C,Olive K. Google’s approach for secure AI agents:An introduction[EB/OL]. (2025-05)[2026-01-18]. https://storage.googleapis.com/gweb-research2023-media/pubtools/1018686.pdf. |
| [126] |
Debenedetti E,Zhang J,Balunovic M,et al. AgentDojo:A dynamic environment to evaluate prompt injection attacks and defenses for LLM agents[C]//The Thirty-Eighth Annual Conference on Neural Information Processing Systems,2024:82895-82920. |
| [127] |
Zhang H,Huang J,Mei K,et al. Agent security bench (ASB):Formalizing and benchmarking attacks and defenses in LLM-based agents[C]//The Thirteenth International Conference on Learning Representations,2025:1-36. |
| [128] |
Vijayvargiya S,Soni A B,Zhou X H,et al. OpenAgentSafety:A comprehensive framework for evaluating real-world AI agent safety[C]//The Fourteenth International Conference on Learning Representations,2026:1-26. |
| [129] |
Lynch A,Wright B,Larson C,et al. Agentic misalignment:How LLMs could be insider threats[PP/OL]. V2. arXiv (2025-10-16)[2026-01-20]. https://doi.org/10.48550/arXiv.2510.05179. |
| [130] |
Kutasov J,Sun Y Q,Colognese P,et al. SHADE-arena:Evaluating sabotage and monitoring in LLM agents[PP/OL]. V2. arXiv (2025-07-08)[2026-01-20]. https://doi.org/10.48550/arXiv.2506.15740. |
| [131] |
Gu X M,Zheng X S,Pang T Y,et al. Agent S mith:A single image can jailbreak one million multimodal LLM agents exponentially fast[C]//The 41st International Conference on Machine Learning,2024:16647-16672. |
| [132] |
Hammond L,Chan A L,Clifton J,et al. Multi-agent risks from advanced AI[PP/OL]. arXiv (2025-02-19)[2026-01-18]. https://doi.org/10.48550/arXiv.2502.14143. |
| [133] |
Kuntz T,Duzan A,Zhao H,et al. OS-harm:A benchmark for measuring safety of computer use agents[C]//The Thirty-Ninth Annual Conference on Neural Information Processing Systems,2025:50396-50427. |
| [134] |
Yang J Y,Shao S,Liu D R,et al. RiOSWorld:Benchmarking the risk of multimodal computer-use agents[C]//The Thirty-Ninth Annual Conference on Neural Information Processing Systems,2025:9322-9369. |
| [135] |
Liao Z Y,Jones J,Jiang L X,et al. RedTeamCUA:Realistic adversarial testing of computer-use agents in hybrid web-OS environments[C]//The Fourteenth International Conference on Learning Representations,2026:1-46. |
| [136] |
Wu L X,Wang C,Liu T M,et al. From assistants to adversaries:Exploring the security risks of mobile LLM agents[PP/OL]. V2. arXiv (2025-05-20)[2026-01-20]. https://doi.org/10.48550/arXiv.2505.12981. |
| [137] |
Liang X,Niu S M,Li Z Y,et al. SafeRAG:Benchmarking security in retrieval-augmented generation of large language model[C]//The 63rd Annual Meeting of the Association for Computational Linguistics,2025:4609-4631. |
| [138] |
Zou W,Geng R P,Wang B H,et al. PoisonedRAG:Knowledge corruption attacks to retrieval-augmented generation of large language models[C]//The 34th USENIX Security Symposium,2025:3827-3844. |
| [139] |
Perçin S,Su X,Syed Q S,et al. Investigating the robustness of retrieval-augmented generation at the query level[C]//The Fourth Workshop on Generation,Evaluation and Metrics (GEM2),2025:439-457. |
| [140] |
Tabassi E. Artificial intelligence risk management framework (AI RMF 1.0):NIST AI 100-1[R]. Gaithersburg:National Institute of Standards and Technology,2023:1-5. |
| [141] |
International Network of AI Safety Institutes. Joint statement on risk assessment of advanced AI systems[EB/OL]. (2024-11-20)[2026-01-18]. https://www.nist.gov/document/joint-statement-risk-assessment-advanced-ai-systems-international-network-aisis. |
| [142] |
Brown T B,Mann B,Ryder N,et al. Language models are few-shot learners[C]//The Thirty-Fourth Annual Conference on Neural Information Processing Systems,2020:1877-1901. |
| [143] |
Chen M,Tworek J,Jun H,et al. Evaluating large language models trained on code[PP/OL]. V2. arXiv (2021-07-14)[2026-01-20]. https://doi.org/10.48550/arXiv.2107.03374. |
| [144] |
Ahmad L,Agarwal S,Lampe M,et al. OpenAI’s approach to external red teaming for AI models and systems[PP/OL]. arXiv (2025-01-24)[2026-01-20]. https://doi.org/10.48550/arXiv.2503.16431. |
| [145] |
OpenAI. DALL·E 2 Preview:Risks and limitations[EB/OL]. (2022-04-11)[2026-01-18]. https://github.com/openai/dalle-2-preview/blob/main/system-card.md. |
| [146] |
OpenAI. DALL·E 3 system card[R]. San Francisco:OpenAI,2023:9-13. |
| [147] |
Sharma M,Tong M,Mu J,et al. Constitutional classifiers:Defending against universal jailbreaks across thousands of hours of red teaming[PP/OL]. arXiv (2025-01-31)[2026-01-20]. https://doi.org/10.48550/arXiv.2501.18837. |
| [148] |
Anthropic. Activating AI safety level 3 protections[EB/OL]. (2025-05-22)[2026-01-18]. https://www.anthropic.com/news/activating-asl3-protections. |
| [149] |
Weidinger L,Mellor J F J,Pegueroles B G,et al. STAR:SocioTechnical approach to red teaming language models[C]//The 2024 Conference on Empirical Methods in Natural Language Processing,2024:21516-21532. |
| [150] |
Fabian D,Crisp J. Why red teams play a central role in helping organizations secure AI systems[EB/OL]. (2023-07)[2026-01-18]. https://services.google.com/fh/files/blogs/google_ai_red_team_digital_final.pdf. |
| [151] |
Radharapu B,Robinson K,Aroyo L,et al. AART:AI-assisted red-teaming with diverse data generation for new LLM-powered applications[C]//The 2023 Conference on Empirical Methods in Natural Language Processing:Industry Track,2023:380-395. |
| [152] |
Deng B Y,Wang W J,Feng F L,et al. Attack prompt generation for red teaming and defending large language models[C]//The 2023 Conference on Empirical Methods in Natural Language Processing,2023:2176-2189. |
| [153] |
Grattafiori A, Evtimov I, Bitton J, et al. Taming the beast: Inside the LLaMA 3 red team process[EB/OL]. 2024[2026-01-18]. https://www.youtube.com/watch? v=UQaNjwLhAmo. |
| [154] |
Ge S Y,Zhou C T,Hou R,et al. MART:Improving LLM safety with multi-round automatic red-teaming[C]//The 2024 Conference of the North American Chapter of the Association for Computational Linguistics:Human Language Technologies,2024:1927-1937. |
| [155] |
Bhatt M,Chennabasappa S,Nikolaidis C,et al. Purple llama CyberSecEval:A secure coding benchmark for language models[PP/OL]. arXiv (2023-12-07)[2026-01-20]. https://doi.org/10.48550/arXiv.2312.04724. |
| [156] |
NVIDIA. NVIDIA AI red team:An introduction[EB/OL]. (2023-06-14)[2026-01-20]. https://developer.nvidia.com/blog/nvidia-ai-red-team-an-introduction. |
| [157] |
NVIDIA. Strength in numbers:NVIDIA and generative red team challenge unleash thousands to vet security at DEF CON[EB/OL]. (2023-08-10)[2026-01-20]. https://blogs.nvidia.com/blog/nvidia-generative-red-team-challenge/. |
| [158] |
NVIDIA. Securing LLM systems against prompt injection[EB/OL]. (2023-08-03)[2026-01-20]. https://developer.nvidia.com/blog/securing-llm-systems-against-prompt-injection/. |
| [159] |
NVIDIA. Secure LLM tokenizers to maintain application integrity[EB/OL]. (2024-06-27)[2026-01-20]. https://developer.nvidia.com/blog/secure-llm-tokenizers-to-maintain-application-integrity/. |
| [160] |
Derczynski L,Galinkin E,Martin J,et al. Garak:A framework for security probing large language models[PP/OL]. arXiv (2024-06-16)[2026-01-20]. https://doi.org/10.48550/arXiv.2406.11036. |
| [161] |
Rebedea T,Dinu R,Sreedhar M N,et al. NeMo guardrails:A toolkit for controllable and safe LLM applications with programmable rails[C]//The 2023 Conference on Empirical Methods in Natural Language Processing:System Demonstrations,2023:431-445. |
| [162] |
Microsoft. Responsible AI transparency report:How we build,support our customers,and grow[R]. Redmond:Microsoft Corporation,2024:9. |
| [163] |
Munoz G D L,Minnich A J,Lutz R,et al. PyRIT:A framework for security risk identification and red teaming in generative AI system[PP/OL]. arXiv (2024-10-01)[2026-01-20]. https://doi.org/10.48550/arXiv.2410.02828. |
| [164] |
Microsoft. Introducing AI red teaming agent:Accelerate your AI safety and security journey with Azure AI Foundry[EB/OL]. (2025-04-04)[2026-02-08]. https://devblogs.microsoft.com/foundry/ai-red-teaming-agent-preview/. |
| [165] |
Wang X,Chen Y H,Li J C,et al. OpenRT:An open-source red teaming framework for multimodal LLMs[PP/OL]. V2. arXiv (2026-01-10)[2026-01-20]. https://doi.org/10.48550/arXiv.2601.01592. |
| [166] |
Huang Y,Sun L C,Wang H R,et al. Position:TrustLLM:Trustworthiness in large language models[C]//The 41st International Conference on Machine Learning,2024:20166-20270. |
中国工程院咨询项目“网信领域人工智能安全防范战略研究”(2025-XBZD-08)
Funding project: Chinese Academy of Engineering project “Strategic Research on Artificial Intelligence Safety Prevention in the Cyber and Information Domain”(2025-XBZD-08)
/
| 〈 |
|
〉 |