Safety Alignment of Large Language Models: The Offensive-Defensive Game Mechanism under Multi-round Adversarial Prompting
Keywords:
Large Language Model, Safety Alignment, Multi-Round Adversarial Prompting, Offensive and Defensive Game MechanismAbstract
With the widespread application of large language models in various fields, the issue of safety alignment with value has become increasingly crucial. The adversarial game mechanism under multiple rounds of adversarial prompts provides a new perspective and approach for the safety alignment of large language models. This paper deeply analyzes the connotation and characteristics of multiple rounds of adversarial prompts, elaborates on the basic principles of the adversarial game mechanism, discusses its significant role in the safety alignment of large language models, analyzes the challenges faced, and proposes corresponding countermeasures, aiming to provide theoretical support for the safe and reliable application of large language models.Downloads
Published
2026-07-12
How to Cite
Gao, X. (2026). Safety Alignment of Large Language Models: The Offensive-Defensive Game Mechanism under Multi-round Adversarial Prompting. CPS Digital Library - Series of Conferences, 1, 33–37. Retrieved from https://seriesofconference.com/index.php/SCJ/article/view/266
Issue
Section
Articles
License
Copyright (c) 2026 Xinyue Gao

This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivatives 4.0 International License.






