SkillTester: Evaluating Agent Skills
- SkillTester is a benchmarking and quality-assurance tool for agent skills that evaluates utility through paired baseline comparisons and measures security using controlled probes.
- It employs a paired evaluation methodology that compares task performance with and without skill activation to reveal efficiency gains and highlight potential risks.
- The framework outputs concise, evidence-rich metrics including utility scores, security scores, and status labels to support informed skill selection and risk management.
SkillTester is a benchmarking and quality-assurance tool for agent skills, understood as reusable skill packages that agents can load to extend their capabilities. It is designed to shift skill selection away from weak proxies such as popularity, reputation, or documentation quality and toward structured evidence about two dimensions: utility, meaning whether a skill materially helps an agent perform a task better than a matched no-skill baseline, and security, meaning whether the skill behaves safely under controlled adversarial or risky conditions (Wang et al., 28 Mar 2026). In this formulation, a skill is not merely documentation; it may include instructions, scripts, setup steps, and tool-use guidance that alter both capability and risk. SkillTester is therefore presented as a comparative quality-assurance harness for skills in an “agent-first” software engineering environment (Wang et al., 28 Mar 2026).
1. Origins, scope, and problem setting
SkillTester is motivated by the rapid growth of skill ecosystems across systems such as Claude Code, OpenAI Codex, GitHub Copilot, and OpenClaw, alongside the absence of systematic evaluation for the skills distributed within those ecosystems (Wang et al., 28 Mar 2026). The framework is explicitly aimed at supporting skill-selection and enablement decisions by making both benefit and risk observable in a controlled way. The paper argues that marketplace discovery tools and rankings, including those from marketplaces such as skills.sh, may help users find skills but do not establish whether a skill is actually effective or safe (Wang et al., 28 Mar 2026).
The security rationale is central to the framework’s scope. The paper lists concrete failure modes for malicious or unsafe skills, including unwanted tool invocation, shell or code execution, unauthorized file/network access, data exfiltration, and prompt-injection-style misuse through downstream tools (Wang et al., 28 Mar 2026). SkillTester addresses these risks by treating skills as executable interventions in an agent’s behavior rather than as passive metadata.
A plausible implication is that SkillTester belongs to a broader research movement toward structured evaluation of systems that extend or guide model behavior. In adjacent work, SKATE evaluates LLMs through model-generated verifiable tasks (Gould et al., 8 Aug 2025), while RUM assesses software testing skills through a hybrid rule+LLM pipeline (Wang et al., 18 Aug 2025). SkillTester differs in target and unit of analysis: it evaluates agent skills, not general reasoning ability or student submissions.
2. Guiding principles and public-facing outputs
The framework is organized around two explicit principles: the comparative utility principle and the user-facing simplicity principle (Wang et al., 28 Mar 2026).
The comparative utility principle requires that skill utility be evaluated relative to a matched no-skill baseline, not in isolation. The operative question is not whether a task can be solved with the skill enabled, but what changes when the same model and environment are run with and without the skill. This baseline is therefore a required reference point for utility evaluation (Wang et al., 28 Mar 2026).
The user-facing simplicity principle states that, although the system may collect many internal details, the public output should remain compact and readable. The paper associates this with minimalist presentation and progressive disclosure, and it reduces the published result to a small set of interpretable outputs rather than a large collection of raw submetrics (Wang et al., 28 Mar 2026).
| Public output | Meaning |
|---|---|
| Utility score | Comparative benefit relative to matched no-skill execution |
| Security score | Aggregate pass rate over controlled security probes |
| Security status label | Pass / Caution / Risky |
This design separates internal evidence collection from external reporting. The paper explicitly states that the framework is meant to be evidence-rich internally, but simple externally (Wang et al., 28 Mar 2026).
3. Paired utility evaluation
SkillTester’s utility evaluation is a paired baseline vs. with-skill procedure. Each task is executed under two matched conditions: a Baseline condition, in which the same task, model, and environment are used but skill use is disabled, and a With-skill condition, in which the same task and environment are used with skill use enabled (Wang et al., 28 Mar 2026). This produces counterfactual evidence about what the skill changes.
Task authoring is derived from the skill itself. The paper specifies that evaluators should compare the claims in SKILL.md with what the skill resources actually do, then convert those claims and gaps into executable tasks with explicit objectives and pass criteria (Wang et al., 28 Mar 2026). The resulting utility task set is divided into Common functional tasks, which cover representative intended use cases, and Edge functional tasks, which cover boundary, failure-handling, and unusual cases (Wang et al., 28 Mar 2026).
A key mechanism is the invocation gate: the skill must actually be invoked for a task to receive utility credit (Wang et al., 28 Mar 2026). This prevents the ambient capability of the underlying model from being mistakenly attributed to the skill. The paper also specifies a common execution schema for raw records, including task identity, success/failure status, invocation outcome, token usage, elapsed time, and security probe pass/fail counts (Wang et al., 28 Mar 2026).
The methodological choice to compare matched executions rather than isolated successes is one of the framework’s central claims. The paper argues that a task may succeed even without the skill, and in that case the skill is useful only if it improves efficiency or otherwise changes the outcome (Wang et al., 28 Mar 2026).
4. Security probe suite and threat organization
Security is evaluated separately from utility through a controlled security probe suite, not through baseline-vs.-with-skill comparison (Wang et al., 28 Mar 2026). The paper states that this separation is deliberate because utility and security answer different questions and should not be conflated.
The public probe taxonomy is divided into three groups:
| Security group | Probe themes |
|---|---|
| Abnormal behavior control | unsafe execution requests, uncontrolled retries, misleading safety claims, tool misuse, risky code or shell behavior |
| Permission boundary | unauthorized file access, unauthorized network access, unauthorized process/tool access, privilege abuse |
| Sensitive data protection | credential leakage, unsafe logging, exfiltration, mishandling sensitive context |
These three groups are described as a compact reporting layer over a broader threat space, and the paper states that they are mapped loosely to OWASP agentic application threat categories (Wang et al., 28 Mar 2026). The stated purpose is practical exposure of risky behaviors under controlled conditions rather than exhaustive proof of safety.
The paper also notes possible future extensions to the security family set, including memory/context poisoning, insecure inter-agent communication, cascading failures, and rogue-agent behavior (Wang et al., 28 Mar 2026). This suggests that the current taxonomy is intentionally compact rather than complete.
5. Scoring model
SkillTester normalizes execution artifacts into a utility score , a security score , and a security status label (Wang et al., 28 Mar 2026). The published utility score is the average over valid functional tasks:
Each task score is gated by invocation:
where is 1 if the skill is invoked on task , otherwise 0 (Wang et al., 28 Mar 2026).
The task value depends on the paired baseline/with-skill outcomes:
Here is the skill-path success indicator and is the baseline-path success indicator (Wang et al., 28 Mar 2026). This yields three cases: skill failure gives 0, skill-only success gives 100, and dual success is scored by relative efficiency.
When both baseline and skill succeed, SkillTester compares token cost and elapsed time. With smoothing constant :
0
The token and time sub-scores are
1
2
with
3
The combined efficiency score is
4
The mapping from efficiency to utility credit is
5
The default parameters are
6
Under these settings, equal-cost success yields approximately 50, better efficiency yields a score above 50 up to 100, and worse efficiency still receives partial credit with a floor of 20 (Wang et al., 28 Mar 2026).
Security is computed independently. For each group 7 among abnormal behavior control, permission boundary, and sensitive data protection, the group score is the pass rate:
8
where 9 is the number of passed probes and 0 is the total number of probes in group 1 (Wang et al., 28 Mar 2026). The overall security score is the unweighted mean:
2
The paper maps this score to a three-level label:
3
with default threshold
4
The status rule embodies an intentionally strict policy: Pass requires perfect security probe performance, Caution denotes good but imperfect performance, and Risky applies below 80 (Wang et al., 28 Mar 2026).
6. System role, deployment, related work, and limitations
The paper describes SkillTester as more than a benchmark: it is a quality-assurance layer around skill ecosystems, intended for an “agent-first” world in which evaluation harnesses become as important as the underlying agent models (Wang et al., 28 Mar 2026). The public service is deployed at https://skilltester.ai, and the broader project is maintained at https://github.com/skilltester-ai/skilltester (Wang et al., 28 Mar 2026).
Its main claims are methodological rather than benchmark-centric. The paper argues that skill utility should be measured comparatively, not absolutely; that security requires a separate evaluation track; that a small set of public outputs is sufficient for practical decisions; that task-level efficiency matters even when both baseline and with-skill runs succeed; and that controlled probe suites are a practical way to expose risk (Wang et al., 28 Mar 2026). This places SkillTester closer to an auditable decision framework than to a single success-rate leaderboard.
The framework also has explicit boundaries. The paper states that SkillTester is not exhaustive certification, formal verification, full marketplace governance, runtime monitoring/adaptation, or a replacement for broader enterprise controls (Wang et al., 28 Mar 2026). The current security taxonomy is described as not exhaustive, and the default parameters 5 are presented as a transparent starting point, not final optimal calibration (Wang et al., 28 Mar 2026).
In a wider context, SkillTester can be read alongside other evaluation systems that operationalize skills through structured evidence. RUM extends assessment from objective script checking to subjective analysis of test cases, scripts, and reports in software testing education (Wang et al., 18 Aug 2025). TESTQUEST treats locator robustness and Page Object quality as skill-shaping targets in a gamified web-testing workflow (Olianas et al., 30 May 2025). KTester improves LLM-based unit-test generation by injecting project-specific and testing-domain knowledge (Li et al., 18 Nov 2025). These systems differ in domain, but they share an orientation toward operational, artifact-grounded evaluation rather than reputation-based inference. This suggests a broader shift toward measurable capability assessment across software testing and agent tooling, with SkillTester specializing that shift for the utility and security of agent skills.