生成式人工智能系统的安全风险评估框架构建研究
Construction of Safety Risk Assessment Framework for Generative Artificial Intelligence Systems
生成式人工智能(GenAI)技术快速演进引发了安全挑战,构建相应的风险评估框架成为精准治理的迫切需求。本文基于生态风险理论的“压力源 ‒ 暴露 ‒ 效应”分析范式,将GenAI视为社会技术生态系统的压力源,构建了包含算法忠实度、技术统治度、行政治理能力3个核心维度的安全风险评估框架;提出了初步的评估指标体系,为将相应理论框架转化为可操作的测量工具提供了基础。其中,算法忠实度细分为认知忠实度(后向、前向)与情感忠实度,技术统治度考量基础设施渗透率和用户采信程度,行政治理能力基于政策布局理论评估管理部门宏观调节水平。以医疗临床辅助场景为例,对所提框架开展了探索性应用,展示了后向生成忠实度、非正式渗透风险在特定情境下的重要性以及框架在该场景中的应用逻辑与适用性。进一步阐述了未来GenAI系统安全风险治理体系的方法创新与突破方向,即构建动态、可计算的风险演化模型,发展智能化、自适应的评估验证技术,推动面向场景的模块化与差异化评估。整体上,安全风险评估框架有助于刻画GenAI风险的形成机制与传导路径,为不同场景下的差异化风险研究提供了理论支撑与模块化评估工具。
The rapid evolution of generative artificial intelligence (GenAI) has given rise to new safety challenges, making it urgent to develop risk assessment frameworks that can support precise governance. Drawing on the stressor‒exposure‒effect analytical paradigm grounded in ecological risk theory, this study conceptualizes GenAI as a stressor within a sociotechnical ecosystem and proposes a safety risk assessment framework comprising three core dimensions: algorithmic fidelity, technological dominance, and administrative governance capacity. It further develops a preliminary system of assessment indicators, providing a basis for translating the theoretical framework into an operational measurement tool. Within this framework, algorithmic fidelity is subdivided into cognitive fidelity (including both backward-looking and forward-looking fidelity) and affective fidelity. Technological dominance considers the penetration of infrastructure and the degree of user acceptance or reliance, while the administrative governance capacity is used to assess the macro-level regulatory capacity of governing authorities based on the policy arrangement approach. Taking clinical decision support in healthcare as an example, this study conducts an exploratory application of the proposed GenAI safety risk assessment framework. The analysis preliminarily illustrates the importance of backward-looking generative fidelity and informal penetration risk in specific contexts, as well as the framework's application logic and potential applicability. The study further discusses directions for methodological innovation and breakthroughs in future GenAI safety risk governance, including the construction of dynamic and computable models of risk evolution, development of intelligent and adaptive assessment and validation technologies, and promotion of scenario-oriented modular and differentiated assessment. Overall, the proposed safety risk assessment framework helps characterize the formation mechanisms and transmission pathways of GenAI risks, providing theoretical support and modular assessment tools for differentiated risk research across diverse application scenarios.
| [1] |
Luna J,Tan I,Xie X F,et al. Navigating governance paradigms:A cross-regional comparative study of generative AI governance processes & principles[C]//Proceedings of the AAAI/ACM Conference on AI,Ethics,and Society,2024:917-931. |
| [2] |
Paul R. European artificial intelligence “trusted throughout the world”:Risk‐based regulation and the fashioning of a competitive common AI market[J]. Regulation & Governance,2024,18(4):1065-1082. |
| [3] |
Cheong I,Caliskan A,Kohno T. Safeguarding human values:Rethinking US law for generative AI’s societal impacts[J]. AI and Ethics,2025,5(2):1433-1459. |
| [4] |
Wach K,Duong C D,Ejdys J,et al. The dark side of generative artificial intelligence:A critical analysis of controversies and risks of ChatGPT[J]. Entrepreneurial Business and Economics Review,2023,11(2):7-30. |
| [5] |
Hannon B,Kumar Y,Sorial P,et al. From vulnerabilities to improvements—A deep dive into adversarial testing of AI models[C]//2023 Congress in Computer Science,Computer Engineering,& Applied Computing (CSCE),2023:2645-2649. |
| [6] |
Gordon G,Rieder B,Sileno G. On mapping values in AI governance[J]. Computer Law & Security Review,2022,46:105712. |
| [7] |
UNESCO. Recommendation on the ethics of artificial intelligence[EB/OL]. (2023-05-16)[2026-06-15]. https://www.unesco.org/en/articles/recommendation-ethics-artificial-intelligence. |
| [8] |
Krafft T D,Zweig K A,König P D. How to regulate algorithmic decision-making:A framework of regulatory requirements for different applications[J]. Regulation and Governance,2022,16(1):119-136. |
| [9] |
European Union. Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence[EB/OL]. (2024-06-13)[2025-12-20]. https://eur-lex.europa.eu/eli/reg/2024/1689/oj. |
| [10] |
Jones R N. An environmental risk assessment/management framework for climate change impact assessments[J]. Natural Hazards,2001,23(2-3):197-230. |
| [11] |
Larned S T,Schallenberg M. Stressor-response relationships and the prospective management of aquatic ecosystems[J]. New Zealand Journal of Marine and Freshwater Research,2019,53(4):489-512. |
| [12] |
Jarvis L,Rosenfeld J,Gonzalez-Espinosa P C,et al. A process framework for integrating stressor-response functions into cumulative effects models[J]. Science of the Total Environment,2024,906:167456. |
| [13] |
Tredinnick L,Laybats C. The dangers of generative artificial intelligence[J]. Business Information Review,2023,40(2):46-48. |
| [14] |
Hacker P,Engel A,Mauer M. Regulating ChatGPT and other large generative AI models[C]//Proceedings of the 2023 ACM Conference on Fairness Accountability and Transparency,2023:1112-1123. |
| [15] |
Weidinger L,Uesato J,Rouh M,et al. Taxonomy of risks posed by language models[C]//Proceedings of the 2022 ACM Conference on Fairness. Accountability and Transparency,2022:214-229. |
| [16] |
Yankouskaya A,Liebherr M,Ali R. Can ChatGPT be addictive? A call to examine the shift from support to dependence in AI conversational large language models[J]. Human-Centric Intelligent Systems,2025,5(1):77-89. |
| [17] |
Taeihagh A. Governance of generative AI[J]. Policy and Society,2025,44(1):1-22. |
| [18] |
Mohamed M A H,Al-Mhdawi M K S,Ojiako U,et al. Generative AI in construction risk management:A bibliometric analysis of the associated benefits and risks[J]. Urbanization,Sustainability and Society,2025,2(1):198-230. |
| [19] |
Duffourc M,Gerke S. Generative AI in health care and liability risks for physicians and safety concerns for patients[J]. Jama,2023,330(4):313. |
| [20] |
Namiot D,Ilyushin E. On cyber risks of generative artificial intelligence[J]. International Journal of Open Information Technologies,2024,12(10):109-119. |
| [21] |
Wu F,Shen T,Bäck T,et al. Knowledge-empowered,collaborative,and co-evolving AI models:The post-LLM roadmap[J]. Engineering,2025,44:87-100. |
| [22] |
Abu-Rasheed H,Weber C,Fathi M. Knowledge graphs as context sources for LLM-based explanations of learning recommendations[C]//2024 IEEE Global Engineering Education Conference (EDUCON),2024:1-5. |
| [23] |
Martino A,Iannelli M,Truong C. Knowledge injection to counter large language model (LLM) hallucination[M]. Cham:Springer,2023:182-185. |
| [24] |
Yang J. Integrated application of LLM model and knowledge graph in medical text mining and knowledge extraction[J]. Social Medicine and Health Management,2024,5(2):56-62. |
| [25] |
Wang Y Q. Generative AI in operational risk management:Harnessing the future of finance[J]. SSRN Electronic Journal,2023:1-10. |
| [26] |
Yankouskaya A,Babiker A,Rizvi S,et al. LLM-D12:A dual-dimensional scale of instrumental and relational dependencies on large language models[J]. ACM Transactions on the Web,2025:3765895. |
| [27] |
Eigner E,Händler T. Determinants of LLM-assisted decision-making[PP/OL]. arXiv (2024-02-27)[2025-12-20]. https://doi.org/10.48550/arXiv.2402.17385. |
| [28] |
Li Z W,Zhang Z,Wang M W,et al. From assistance to reliance:Development and validation of the large language model dependence scale[J]. International Journal of Information Management,2025,83:102888. |
| [29] |
Liu S H,Hu X,Xia X,et al. An empirical study of vulnerable package dependencies in LLM repositories[PP/OL]. arXiv (2025-08-29)[2025-12-20]. https://doi.org/10.48550/arXiv.2508.21417. |
| [30] |
Li L Q,Wang Z F,Jose J M,et al. LLM supporting knowledge tracing leveraging global subject and student specific knowledge graphs[J]. Information Fusion,2026,126:103577. |
| [31] |
US Congress. Candidate voice fraud prohibition act[EB/OL].(2023-07-13)[2025-12-20]. https://www.congress.gov/bill/118th-congress/house-bill/4611. |
| [32] |
US Congress. Preventing deep fake scams act[EB/OL]. (2025-06-18)[2025-12-20]. https://www.congress.gov/bill/119th-congress/senate-bill/2117/all-info. |
| [33] |
Department for Science,Innovation and Technology. A pro-innovation approach to AI regulation[EB/OL]. (2023-03-29)[2025-12-20]. https://www.gov.uk/government/publications/ai-regulation-a-pro-innovation-approach. |
| [34] |
国家互联网信息办公室. 生成式人工智能服务管理暂行办法[EB/OL]. (2023-07-10)[2025-12-20]. https://www.cac.gov.cn/2023-07/13/c_1690898327029107.htm. |
| [35] |
Cyberspace Administration of China. Measures for the administration of generative artificial intelligence services[EB/OL]. (2023-07-10)[2025-12-20]. https://www.cac.gov.cn/2023-07/13/c_1690898327029107.htm. |
| [36] |
European Parliament. Artificial intelligence act:MEPs adopt landmark law[EB/OL]. (2024-03-13)[2026-06-15]. https://www.europarl.europa.eu/news/en/press-room/20240308IPR19015/artificial-intelligence-act-meps-adopt-landmark-law. |
| [37] |
Kreps S,McCain R M,Brundage M. All the news that’s fit to fabricate:AI-generated text as a tool of media misinformation[J]. Journal of Experimental Political Science,2022,9(1):104-117. |
| [38] |
Xu D N,Fan S J,Kankanhalli M. Combating misinformation in the era of generative AI models[C]//Proceedings of the 31st ACM International Conference on Multimedia,2023:9291-9298. |
| [39] |
Argyle L P,Busby E C,Fulda N,et al. Out of one,many:Using language models to simulate human samples[J]. Political Analysis,2023,31(3):337-351. |
| [40] |
Wan Y X,Pu G,Sun J,et al. “Kelly is a warm person,Joseph is a role model”:Gender biases in LLM-generated reference letters[PP/OL]. V5. arXiv (2023-12-01)[2025-12-20]. https://doi.org/10.48550/arXiv.2310.09219. |
| [41] |
Sheng E,Chang K W,Natarajan P,et al. The woman worked as a babysitter:On biases in language generation[C]//Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing,2019:3405-3410. |
| [42] |
Merritt S M,Ako-Brew A,Bryant W J,et al. Automation-induced complacency potential:Development and validation of a new scale[J]. Frontiers in Psychology,2019,10:225. |
| [43] |
Arts B,Leroy P. Institutional dynamics in environmental governance[M]. Dordrecht:Springer,2006:45-68. |
| [44] |
Birkstedt T,Minkkinen M,Tandon A,et al. AI governance:Themes,knowledge gaps and future agendas[J]. Internet Research,2023,33(7):133-167. |
| [45] |
Hu X M,Chen J Z,Li X C,et al. Do large language models know about facts?[PP/OL]. arXiv (2023-10-08)[2025-12-20]. https://doi.org/10.48550/arXiv.2310.05177. |
| [46] |
Wang B S,Yue X,Sun H. Can ChatGPT defend its belief in truth? Evaluating LLM reasoning via debate[PP/OL]. V2. arXiv (2023-10-10)[2025-12-20]. https://doi.org/10.48550/arXiv.2305.13160. |
| [47] |
Tian Y F,Ravichander A,Qin L H,et al. MacGyver:Are large language models creative problem solvers?[PP/OL]. V4. arXiv (2025-02-22)[2025-12-20]. https://doi.org/10.48550/arXiv.2311.09682. |
| [48] |
Thorne J,Vlachos A,Christodoulopoulos C,et al. FEVER:A large-scale dataset for Fact Extraction and VERification[PP/OL]. V3. arXiv (2018-12-18)[2025-12-20]. https://doi.org/10.48550/arXiv.1803.05355. |
| [49] |
Schuster T,Fisch A,Barzilay R. Get your vitamin C! Robust fact verification with contrastive evidence[PP/OL]. arXiv (2021-03-15)[2025-12-20]. https://doi.org/10.48550/arXiv.2103.08541. |
| [50] |
Politifact fact-checking dataset[DB/OL]. [2025-12-20]. https://www.kaggle.com/datasets/shivkumarganesh/politifact-factcheck-data. |
| [51] |
Cobbe K,Kosaraju V,Bavarian M,et al. Training verifiers to solve math word problems[PP/OL]. V2. arXiv (2021-11-18)[2025-12-20]. https://doi.org/10.48550/arXiv.2110.14168. |
| [52] |
Saparov A,He H. Language models are greedy reasoners:A systematic formal analysis of chain-of-thought[PP/OL]. V4. arXiv (2023-03-02)[2025-12-20]. https://doi.org/10.48550/arXiv.2210.01240. |
| [53] |
Rao S,Tetreault J. Dear sir or madam,may I introduce the GYAFC dataset:Corpus,benchmarks and metrics for formality style transfer[C]//Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics:Human Language Technologies,2018:129-140. |
| [54] |
Socher R,Perelygin A,Wu J,et al. Recursive deep models for semantic compositionality over a sentiment treebank[C]//Proceedings of the 2013 Conference on Empirical Methods in Natural Language Processing,2013:1631-1642. |
| [55] |
European Commission. Ethics guidelines for trustworthy AI[EB/OL]. (2019-04-08)[2025-12-20]. https://digital-strategy.ec.europa.eu/en/library/ethics-guidelines-trustworthy-ai. |
| [56] |
Statt N. Google dissolves AI ethics board just one week after forming it[EB/OL]. (2019-04-05)[2025-12-20]. https://www.theverge.com/2019/4/4/18296113/google-ai-ethics-board-ends-controversy-kay-coles-james-heritage-foundation. |
| [57] |
Morgan D. Anticipatory regulatory instruments for AI systems:A comparative study of regulatory sandbox schemes[C]//Proceedings of the 2023 AAAI/ACM Conference on AI,Ethics,and Society,2023:980-981. |
| [58] |
Autio C,Schwartz R,Dunietz J,et al. Artificial intelligence risk management framework:Generative artificial intelligence profile:NIST AI 600-1[R]. Gaithersburg:National Institute of Standards and Technology,2024. |
| [59] |
Gan Y Y,Yang Y,Ma Z,et al. Navigating the risks:A survey of security,privacy,and ethics threats in LLM-based agents[PP/OL]. arXiv (2024-11-14)[2025-12-20]. https://doi.org/10.48550/arXiv.2411.09523. |
| [60] |
Goetz L,Trengove M,Trotsyuk A,et al. Unreliable LLM bioethics assistants:Ethical and pedagogical risks[J]. The American Journal of Bioethics,2023,23(10):89-91. |
| [61] |
Mou Y T,Zhang S K,Ye W. SG-bench:Evaluating LLM safety generalization across diverse tasks and prompt types[C]//Advances in Neural Information Processing Systems 37,2024:123032-123054. |
| [62] |
Sarker I H. LLM potentiality and awareness:A position paper from the perspective of trustworthy and responsible AI modeling[J]. Discover Artificial Intelligence,2024,4(1):40. |
| [63] |
Zhou X Y,Wang W D,Lu L,et al. SafeAgent:Safeguarding LLM agents via an automated risk simulator[PP/OL]. V2. arXiv (2025-07-18)[2025-12-20]. https://doi.org/10.48550/arXiv.2505.17735. |
| [64] |
Belmoukadam O,De Jonghe J,Ajridi S,et al. AdversLLM:A practical guide to governance,maturity and risk assessment for LLM-based applications[J]. International Journal on Cybernetics & Informatics,2024,13(6):51-68. |
| [65] |
Hua W Y,Yang X J,Jin M Y,et al. TrustAgent:Towards safe and trustworthy LLM-based agents[C]//Findings of the Association for Computational Linguistics:EMNLP 2024,2024:10000-10016. |
| [66] |
Ji J M,Liu M,Dai J,et al. BeaverTails:Towards improved safety alignment of LLM via a human-preference dataset[C]//Advances in Neural Information Processing Systems 36,2023:24678-24704. |
| [67] |
Peng S,Chen P Y,Hull M,et al. Navigating the safety landscape:Measuring risks in finetuning large language models[C]//Advances in Neural Information Processing Systems 37,2024: 95692-95715. |
| [68] |
Wei A,Haghtalab N,Steinhardt J. Jailbroken:How does LLM safety training fail?[C]//Advances in Neural Information Processing Systems 36,2023:80079-80110. |
| [69] |
Yang Z Y,Raman S S,Shah A,et al. Plug in the safety chip:Enforcing constraints for LLM-driven robot agents[C]//2024 IEEE International Conference on Robotics and Automation (ICRA),2024:14435-14442. |
| [70] |
Röttger P,Pernisi F,Vidgen B,et al. SafetyPrompts:A systematic review of open datasets for evaluating and improving large language model safety[C]//Proceedings of the AAAI Conference on Artificial Intelligence,2025:27617-27627. |
| [71] |
Bird C,Ungless E,Kasirzadeh A. Typology of risks of generative text-to-image models[C]//Proceedings of the 2023 AAAI/ACM Conference on AI,Ethics,and Society,2023:396-410. |
中国工程院咨询项目“国家级大模型监管保险箍模式研究”(2025-XZ-08)
教育部哲学社会科学重大课题研究项目(24JZD040)
国家自然科学基金项目(72293583)
国家自然科学基金项目(72293580)
/
| 〈 |
|
〉 |