Isolate the Contributions of Individual Claude Code Safety-Prompt Sections

Isolate the contribution of individual safety-oriented sections in the Claude Code client harness to its overall refusal rate for suspicious MCP tool workflows.

Background

The paper observes that Claude Code refuses the three-channel fragmentation attack more reliably than clients such as Cursor and Codex CLI. The authors hypothesize that this behavior results from the cumulative effect of multiple safety-oriented components in Claude Code’s system prompt, including action-caution guidance, tool-approval instructions, and descriptions of the permission framework, rather than from a single instruction. Determining the individual contribution of each prompt section would clarify which components cause the observed cross-client safety gap and could inform the design of more effective client-side defenses.

References

Isolating the contribution of individual prompt sections to the overall refusal rate remains an open question for future work.

— Measuring and Exploiting Implicit Trust in LLM Tool-Calling Pipelines  (2609.18217 - Ediga et al., 16 Sep 2026) in Section 4, subsection “Cross-Client Safety Gap” (label: sec:client-gap)