Safety Alignment of Large Language Models: The Offensive-Defensive Game Mechanism under Multi-round Adversarial Prompting

Authors

  • Xinyue Gao Xi’an University of Posts & Telecommunications, Xi’an, 710116, China

Keywords:

Large Language Model, Safety Alignment, Multi-Round Adversarial Prompting, Offensive and Defensive Game Mechanism

Abstract

With the widespread application of large language models in various fields, the issue of safety alignment with value has become increasingly crucial. The adversarial game mechanism under multiple rounds of adversarial prompts provides a new perspective and approach for the safety alignment of large language models. This paper deeply analyzes the connotation and characteristics of multiple rounds of adversarial prompts, elaborates on the basic principles of the adversarial game mechanism, discusses its significant role in the safety alignment of large language models, analyzes the challenges faced, and proposes corresponding countermeasures, aiming to provide theoretical support for the safe and reliable application of large language models.

Downloads

Published

2026-07-12

How to Cite

Gao, X. (2026). Safety Alignment of Large Language Models: The Offensive-Defensive Game Mechanism under Multi-round Adversarial Prompting. CPS Digital Library - Series of Conferences, 1, 33–37. Retrieved from https://seriesofconference.com/index.php/SCJ/article/view/266