Papers
Topics
Authors
Recent
Search
2000 character limit reached

When Not to Imitate: Boundary-Aware Skill Memory for Reliable Tool-Use LLM Agents

Published 23 Aug 2026 in cs.CL | (2608.22339v1)

Abstract: Extracting skills from past successes is critical for the efficient evolution of LLM agents. Prevailing agent self-evolution paradigms typically rely on a core assumption: equipping LLMs with skill memories derived from successful trajectories will monotonically improve their problem-solving capabilities. However, probe analyses reveal that extracting skills solely from successful trajectories traps the model in a \textbf{Skill Imitation Trap}. For tasks that resemble past successes but require different tools, retrieving more skills paradoxically increases the model's confidence in wrong tool calls---procedure skills raise the wrong-tool margin by 47%47\% over a memory-free baseline. To overcome this limitation, we propose \textbf{Boundary-Aware Skill Memory} (BASM), which augments each skill with explicit boundary fields---applicability conditions, risk cues, avoidance rules, and recovery notes. These fields transform each retrieved skill from an unconditional action template into state-conditioned guidance: the agent applies the skill when its conditions hold, suppresses inapplicable tool calls when they do not, and issues targeted repairs when execution fails. Across three agent benchmarks and four model scales, BASM consistently outperforms success-distilled skill-memory baselines: it improves task success rate by up to 23.8%23.8\% on AppWorld, accuracy by up to 5.0%5.0\% on BFCL, and reduces attack success rate by 4.6%4.6\% on AgentDojo, while simultaneously reducing average AppWorld steps by up to 6.6%6.6\% relative to the memory-free baseline.

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.

Tweets

Sign up for free to view the 2 tweets with 90 likes about this paper.