Runtime Behavioral Monitoring for Agent Skills

Develop runtime monitoring approaches that can reliably distinguish malicious agent actions from legitimate ones in deployments of the Agent Skills framework without relying on a formal behavioral specification and while maintaining low false positive rates.

Background

Because Agent Skills specify behavior in natural language and operate over an unbounded action space, there is no formal behavioral specification to anchor runtime checks.

The authors note that approaches like anomaly detection, behavioral fingerprints, and LLM-based intent classification remain unvalidated at scale, while excessive false positives could render monitoring impractical.

References

Developing runtime monitoring approaches that can distinguish malicious agent actions from legitimate ones---without a formal behavioral specification and without generating prohibitive false positive rates---is an open challenge.

Towards Secure Agent Skills: Architecture, Threat Taxonomy, and Security Analysis  (2604.02837 - Li et al., 3 Apr 2026) in Section 7.2, Open Challenges (C3: Runtime Behavioral Monitoring)

Even with these improvements, preventing multi-context attacks at an acceptable cost remains an open problem.

Red-Teaming Auto Mode: Improving Blocking Classifiers Against Malign Coding Agents  (2609.19587 - Remedios et al., 17 Sep 2026) in Abstract, page 1

Because SkillShift uses no explicit injection and preserves valid task outputs, whether such violations of Skill Policy Integrity can be detected remains unclear.

A Finger on the Scale: Covert Policy Steering through Agentic Skills  (2609.02564 - Li et al., 2 Sep 2026) in Section 2, Related Work, subsection “Skill Attacks and Detection Gap”