Papers
Topics
Authors
Recent
Search
2000 character limit reached

Secure Tug-of-War (SecTOW)

Updated 3 July 2026
  • Secure Tug-of-War (SecTOW) is an iterative defense-attack framework designed to enhance MLLM security by countering jailbreaking inputs.
  • It employs Group Relative Policy Optimization in alternate attacker and defender training rounds to automatically expand jailbreak data and refine refusal strategies.
  • Robust quality monitoring and rule-based rewards yield drastic ASR reductions and low ORR, ensuring both security and general utility.

Secure Tug-of-War (SecTOW) is an iterative defense-attack training framework for enhancing the security of multimodal LLMs (MLLMs) against jailbreaking inputs—maliciously crafted image-query pairs designed to bypass safety constraints. Unlike traditional guardrail-type solutions, which apply external filters but leave internal vulnerabilities unaddressed, or standard supervised fine-tuning (SFT), which often induces excessive refusals on benign queries, SecTOW employs reinforcement learning (specifically, Group Relative Policy Optimization) to conduct an adversarial, end-to-end retraining of both the defender and the attacker. The result is a white-box security protocol that automatically expands jailbreak data, robustly enhances refusal strategies, and monitors both performance and reward exploitation to maintain strong general utility (Dai et al., 29 Jul 2025).

1. System Architecture and Iterative Process

At its core, SecTOW alternates between training two independently parameterized MLLMs—a defender (MDM_D) and an auxiliary attacker (MAM_A)—in a multiround cycle.

  • Defender (MDM_D): Processes input tuples (I,Q)(I, Q) (image and query), producing a response R=MD(I,Q)R = M_D(I, Q). Its objective is to safely refuse or answer malicious inputs while retaining informativeness for benign queries.
  • Attacker (MAM_A): Utilizes a “think-then-attack” prompt to generate or refine jailbreak queries, conditioned on images from a general dataset.

Training advances through KK iterations, each with two phases:

  1. Attack Phase: MAM_A is updated via GRPO using current folds of jailbreak (DJ(k)D_J^{(k)}) and general data (DG(k)D_G^{(k)}), generating raw jailbreaks that are filtered for high quality.
  2. Defense Phase: MAM_A0 retrains via GRPO on a union of newly discovered jailbreaks and benign samples, updating its refusal strategy.

Data partitioning is performed by splitting both the initial, scarce jailbreak dataset MAM_A1 and a much larger general set MAM_A2 into MAM_A3 folds, cycling through each fold per iteration. Cold-start phases (SFT for MAM_A4 and initial GRPO for MAM_A5) address reward sparsity at the outset (Dai et al., 29 Jul 2025).

2. Reinforcement Learning Formalism and GRPO

Both defender and attacker operate as autoregressive policies MAM_A6, generating token sequences over the respective state and action spaces:

  • State space MAM_A7: MAM_A8, where MAM_A9 is the sequence of previously generated tokens.
  • Action space MDM_D0: MLLM vocabulary.

Policies are optimized to maximize expected return:

MDM_D1

where trajectories MDM_D2 accrue cumulative reward MDM_D3.

SecTOW implements Group Relative Policy Optimization (GRPO), which eschews a value baseline in favor of intra-group relative advantage:

  • For each input, MDM_D4 rollouts MDM_D5 produce rewards MDM_D6; mean MDM_D7 and advantage MDM_D8 are computed.
  • The policy is updated using the surrogate loss:

MDM_D9

where (I,Q)(I, Q)0 is a KL regularization coefficient, often tuned dynamically.

GRPO is favored for its simplicity (no learned value function), direct regularization, and variance reduction within large-scale MLLM training (Dai et al., 29 Jul 2025).

3. Reward Design: Defender and Attacker Objectives

Distinct reward functions are utilized for each role, built on simple, deterministic criteria rather than model-based or generative labeling.

  • Defender Reward: For each sample (I,Q)(I, Q)1, assign binary label (I,Q)(I, Q)2 (“reject required” or “pass through”). Rule-based detection (I,Q)(I, Q)3 establishes if the model’s response matches a refusal phrase. Reward function:

(I,Q)(I, Q)4

This yields stable, low-noise feedback.

  • Attacker Reward: For each (image, query) (I,Q)(I, Q)5, a composite criterion is used. First, pass the query through the defender and judge if the response is unsafe (via an external safety evaluator). Second, enforce template adherence in the generated query:

(I,Q)(I, Q)6

where both scores are binary indicators.

These rule-based mechanisms reduce the need for complex generative annotations and enable robust scaling to synthetic samples (Dai et al., 29 Jul 2025).

4. Quality Monitoring to Mitigate Reward Gaming

To avoid degenerate behaviors—such as the attacker generating repetitive, low-diversity queries or the defender excessively rejecting benign inputs (“over-refusal”)—SecTOW incorporates explicit early stopping criteria:

  • Attacker Diversity: On a validation set, computes

(I,Q)(I, Q)7

(using FuzzyWuzzy or other string similarity). If (I,Q)(I, Q)8 drops by (I,Q)(I, Q)9, attacker retraining is halted.

  • Defender Over-Refusal Rate (ORR): On general validation,

R=MD(I,Q)R = M_D(I, Q)0

If ORR exceeds R=MD(I,Q)R = M_D(I, Q)1, defender retraining is stopped to preserve utility.

This systematic monitoring ensures that neither model exploits the rule-based rewards to the detriment of system balance or general purpose performance (Dai et al., 29 Jul 2025).

5. Performance Evaluation and Ablation Analysis

SecTOW was instantiated with Qwen2-VL-7B for both defender and attacker, using Llama-Guard-3 as the external safety evaluator. Benchmarks include:

Key metrics:

Post three SecTOW iterations, results are as follows:

Benchmark Initial ASR / ACC∕ORR Post-SecTOW ASR / ACC∕ORR
JailBreakV-28k 0.1918 0.0061
FigStep 0.3320 0.0000
SafeBench 0.1404 0.0022
MM-SafetyBench 0.6726 0.0298
MMMU 0.5411∕0.0000 0.5422∕0.0078
MMMU-Pro 0.4116∕0.0000 0.4145∕0.0006

Compared to SFT, which reduced ASR but increased ORR above 20%, SecTOW achieves lower ASR while keeping ORR beneath R=MD(I,Q)R = M_D(I, Q)2. Relative to other dense defenses (MMS, MIRage), SecTOW reduces ASR by at least 88% on JailBreakV-28k and neutralizes all FigStep attacks (Dai et al., 29 Jul 2025).

Ablation studies demonstrate the necessity of iteration, defender and attacker monitoring, and cold start phases for optimal robustness and usability. Notably, disabling diversity monitoring or defender ORR checks led to substantial ASR and ORR degradations, respectively.

6. Implementation Notes and Adaptation Strategies

SecTOW’s reproducibility and adaptability to diverse MLLM contexts are underpinned by modular and lightweight design choices:

  • Model and Evaluator Choices: Qwen2-VL-7B and Llama-Guard-3; suitable substitutions can be made for other architectures (e.g., Flamingo, LLaVA).
  • Iteration Count: R=MD(I,Q)R = M_D(I, Q)3; attacker allocation: 80% training, 20% validation for diversity.
  • Cold Start: Defender SFT on full seed data; attacker early GRPO warm-up.
  • GRPO Settings: Learning rate R=MD(I,Q)R = M_D(I, Q)4, group size R=MD(I,Q)R = M_D(I, Q)5, KL weight R=MD(I,Q)R = M_D(I, Q)6.
  • Jailbreak Filtering: Retain generated (image, query) pairs if at least R=MD(I,Q)R = M_D(I, Q)7 out of R=MD(I,Q)R = M_D(I, Q)8 defender samples are successful attacks.
  • Integration: Adaptation requires (a) swapping in the target architecture, (b) preparing minimal jailbreak and general sets, (c) implementing GRPO-based training, and (d) leveraging rule-based reward and monitoring routines as specified.

The minimal reliance on learned rewards and annotations, coupled with explicit cycle monitoring, makes SecTOW broadly extensible for robust, scalable MLLM security enhancement (Dai et al., 29 Jul 2025).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Secure Tug-of-War (SecTOW).