Papers
Topics
Authors
Recent
Search
2000 character limit reached

SecInfer: Robust Defense Against Prompt Injection

Updated 14 July 2026
  • SecInfer is a prevention-based defense that leverages inference-time scaling to generate diverse candidate responses, mitigating prompt injection attacks in LLMs.
  • It employs system-prompt-guided sampling and stochastic decoding to explore varied reasoning paths, enhancing robustness under adversarial conditions.
  • The method uses target-task-guided aggregation to select the candidate best aligned with the intended task, outperforming traditional defenses.

SecInfer is a prevention-based defense against prompt injection attacks in LLMs that is built on inference-time scaling, an approach that allocates additional compute during inference to improve model behavior. Rather than relying on fine-tuning alone, SecInfer combines system-prompt-guided sampling with target-task-guided aggregation so that multiple candidate responses are generated along diverse reasoning paths and then filtered according to alignment with the intended task. In the formulation reported in "SecInfer: Preventing Prompt Injection via Inference-time Scaling," the method is evaluated against both existing and adaptive prompt injection attacks and is described as outperforming state-of-the-art defenses as well as existing inference-time scaling approaches (Liu et al., 29 Sep 2025).

1. Conceptual setting and threat model

Prompt injection attacks exploit a model’s tendency to follow attacker-controlled instructions embedded within untrusted user data, causing the model to execute the injected task rather than the intended target task. SecInfer is positioned as a prevention-based defense for this setting, rather than a detection-only mechanism that simply discards contaminated samples (Liu et al., 29 Sep 2025).

A central premise of the method is that conventional prevention-based defenses, including prompt pre-processing and fine-tuning for security, have limited effectiveness against strong attacks, especially optimization-based prompt injection attacks. The work also argues that standard inference-time scaling methods such as self-consistency or majority voting are insufficient because injected prompts can systematically bias all sampled outputs. This establishes the specific problem that SecInfer addresses: not merely improving average reasoning quality, but using inference-time scaling to increase robustness under adversarial prompt contamination.

The framework assumes knowledge of the intended target task. That assumption is operationalized directly in the aggregation stage, where the defense attempts to identify which candidate response best serves the original task rather than the attacker’s objective. This suggests that SecInfer is best understood as a task-conditioned robustness mechanism rather than a generic sampling heuristic.

2. Inference-time scaling as a security mechanism

Inference-time scaling in SecInfer refers to allocating additional computation and sampling during inference to improve security under prompt injection. The method uses this additional compute to generate multiple candidate outputs for a given input, vary the generation process through system prompts and stochastic decoding, and then apply a task-aware selection mechanism (Liu et al., 29 Sep 2025).

The paper contrasts this use of inference-time scaling with prior scaling techniques such as self-consistency, best-of-NN, in-context learning, and iterative refinement. In those approaches, more samples or more reasoning steps are typically used to improve task accuracy or reasoning reliability. SecInfer instead uses additional inference-time computation to diversify candidate trajectories under adversarial conditions and then select outputs that remain aligned with the intended task.

This design implies a security-through-diversification principle. The method does not assume that a single generation path is reliable once prompt injection is present. Instead, it attempts to ensure that at least one candidate remains task-faithful and that the aggregation mechanism can recover it. A plausible implication is that diversity alone is not sufficient; diversity must be coupled to a selection mechanism that explicitly encodes the target task.

3. System-prompt-guided sampling

The first stage of SecInfer is system-prompt-guided sampling. The method maintains a set of carefully designed Chain-of-Thought system prompts, denoted by

P={p1,p2,…,pn}.\mathcal{P} = \{p_1, p_2, \ldots, p_n\}.

For each sample, one prompt pki∈Pp_{k_i} \in \mathcal{P} is selected, concatenated with the original target-task instruction sts_t and the possibly contaminated user data xcx_c, and passed to the backend LLM ff to produce a candidate response (Liu et al., 29 Sep 2025).

The sampling step is written as

r^i=f(pki ∥ st ∥ xc),\hat{r}_i = f(p_{k_i} \,\|\, s_t \,\|\, x_c),

and is repeated NN times to obtain a candidate set

R={r^1,…,r^N}.\mathcal{R} = \{\hat{r}_1, \ldots, \hat{r}_N\}.

The use of a system-prompt pool is the distinguishing feature of this stage. The goal is not merely stochastic variation, but exploration of diverse reasoning paths induced by different system prompts. The paper further describes the use of stochastic decoding, including temperature sampling and top-kk sampling, which can be combined to increase diversity. With temperature P={p1,p2,…,pn}.\mathcal{P} = \{p_1, p_2, \ldots, p_n\}.0, the autoregressive token distribution is written as

P={p1,p2,…,pn}.\mathcal{P} = \{p_1, p_2, \ldots, p_n\}.1

The motivation for this design is explicit: naive random sampling can still fail because prompt injection may bias the generation distribution in a correlated way across samples. System-prompt-guided sampling is therefore intended to reduce correlation among failure modes. In the reported ablations, combining system prompts with stochastic decoding yields the best diversity and robustness, and increasing the number of samples P={p1,p2,…,pn}.\mathcal{P} = \{p_1, p_2, \ldots, p_n\}.2 improves robustness quickly, with attack success rate approaching near zero at P={p1,p2,…,pn}.\mathcal{P} = \{p_1, p_2, \ldots, p_n\}.3.

4. Target-task-guided aggregation

The second stage of SecInfer is target-task-guided aggregation, which selects the candidate most likely to accomplish the intended task. This stage differs depending on whether the output space is closed-domain or open-domain (Liu et al., 29 Sep 2025).

For closed-domain tasks, such as multiple-choice settings, each candidate response is mapped into the task’s output set, for example by keyword spotting, and majority voting is applied over valid mapped responses. In this setting, the aggregation mechanism uses the known target-task label space as a filter.

For open-domain tasks, the procedure is more elaborate. Candidate responses are first embedded,

P={p1,p2,…,pn}.\mathcal{P} = \{p_1, p_2, \ldots, p_n\}.4

and then clustered using a non-parametric method such as agglomerative clustering. For each cluster P={p1,p2,…,pn}.\mathcal{P} = \{p_1, p_2, \ldots, p_n\}.5, the centroid is

P={p1,p2,…,pn}.\mathcal{P} = \{p_1, p_2, \ldots, p_n\}.6

and the representative response is chosen as the candidate closest to that centroid,

P={p1,p2,…,pn}.\mathcal{P} = \{p_1, p_2, \ldots, p_n\}.7

The set of cluster representatives is then passed, together with the original target instruction P={p1,p2,…,pn}.\mathcal{P} = \{p_1, p_2, \ldots, p_n\}.8, to a judge LLM P={p1,p2,…,pn}.\mathcal{P} = \{p_1, p_2, \ldots, p_n\}.9, which selects the representative best aligned with the intended task or abstains: pki∈Pp_{k_i} \in \mathcal{P}0

This aggregation strategy is central to SecInfer’s distinction from standard best-of-pki∈Pp_{k_i} \in \mathcal{P}1 or majority-vote methods. The paper’s position is that candidate selection must be conditioned on the intended task, because prompt injection can cause the majority of responses to be attacker-aligned. The use of an LLM-as-a-Judge is therefore not merely a ranking component; it is the mechanism that converts response diversity into task-conditioned security.

5. Empirical evaluation

The reported evaluation covers both open-weight and closed-source backends, specifically LLaMA3.1-8B-Instruct, Qwen3-8B, GPT-4o, and GPT-4.1, and spans 6 target tasks, 8 injected tasks, 7 standard attacks, and 6 adaptive attacks (Liu et al., 29 Sep 2025). The main metrics are utility pki∈Pp_{k_i} \in \mathcal{P}2, utility under attack pki∈Pp_{k_i} \in \mathcal{P}3, and attack success rate pki∈Pp_{k_i} \in \mathcal{P}4.

The paper states that SecInfer dramatically improves security, with pki∈Pp_{k_i} \in \mathcal{P}5 substantially higher than for baseline defenses and pki∈Pp_{k_i} \in \mathcal{P}6 reduced to near-zero for all attack types when SecInfer is applied. By contrast, no-defense and alternative defenses are reported as having high pki∈Pp_{k_i} \in \mathcal{P}7 under strong attacks. Closed-source models are also described as robust under SecInfer, with utility and security on par with open-weight models.

The baseline set includes prompt pre-processing defenses such as paraphrasing, retokenization, delimiters, and sandwich or instructional defense; detection via DataSentinel; fine-tuning-based defense via SecAlign++; and inference-time scaling baselines such as self-consistency, best-of-pki∈Pp_{k_i} \in \mathcal{P}8, CoT, ICL, and iterative refinement. According to the reported comparison, prompt pre-processing defenses have limited effectiveness, detection methods can achieve low pki∈Pp_{k_i} \in \mathcal{P}9 but very low sts_t0 because contaminated samples are discarded, SecAlign++ performs reasonably on weak attacks but is broken by optimization-based attacks, and generic inference-time scaling baselines remain inadequate because filtered outputs are often attacker-chosen.

The evaluation also extends to agent settings using InjecAgent and AgentDojo. In those settings, SecInfer is reported to reduce agent prompt injection success from 40–60% to near-zero without harming agent task completion as measured by sts_t1. The paper further notes that SecInfer is comparable in inference cost to other scaling-based approaches and is highly parallelizable.

6. Adaptive attacks, limitations, and interpretation

SecInfer is evaluated against adaptive attacks that target the defense itself, including attempts to disable reasoning, optimize for multiple contaminated candidates, and attack the judge model or manipulate response selection (Liu et al., 29 Sep 2025). The reported result is that sts_t2 remains low even under these adaptive conditions, although sts_t3 can decrease in extreme settings, such as when the attacker is allowed to erase or entirely control input content.

The principal limitation identified in the work concerns cases where the attacker-chosen task and the intended task are of the same type, for example when both are sentiment analysis tasks. In that scenario, the paper states that no prompt-based defense can distinguish maliciously crafted content from genuine content, and the problem reduces to adversarial example generation. This limitation places a clear boundary on what SecInfer claims to solve: it addresses misalignment between intended and injected tasks, not arbitrary semantic ambiguity when both tasks share the same output structure.

A related misconception addressed by the work is the idea that more sampling alone is sufficient. SecInfer argues that additional samples without task-guided aggregation do not reliably defend against prompt injection because the attack can correlate failure across samples. The method’s contribution is therefore the combination of two specific ingredients: diversity induced by system-prompt-guided sampling and explicit target-task alignment enforced during aggregation.

7. Position within LLM security research

SecInfer frames inference-time scaling as a security primitive rather than only a capability-enhancement technique. Its core claim is that additional inference-time compute can be traded for robustness against prompt injection if the compute is used to generate diverse candidate reasoning paths and to aggregate them with explicit reference to the intended task (Liu et al., 29 Sep 2025).

Within that framing, the method occupies a distinct position relative to prevention-based fine-tuning defenses and detection-based filtering systems. It is prevention-based, but it does not depend on retraining the target model to become secure. It is scaling-based, but it does not assume that majority behavior among samples is trustworthy. It is aggregation-heavy, but the aggregation is task-aware rather than purely statistical.

This suggests a broader interpretation of SecInfer as an instance of inference-time control for security-sensitive generation. The reported results support that interpretation in the specific context studied: prompt injection attacks against LLMs and LLM agents, across multiple model families, target tasks, and adaptive attack settings.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to SecInfer.