Papers
Topics
Authors
Recent
Search
2000 character limit reached

DEFENDCLI: Command-Line Provenance EDR

Updated 8 July 2026
  • DEFENDCLI is a provenance-based Endpoint Detection and Response system that centers on command-line activity to provide granular attack correlation.
  • It reduces noise in conventional provenance graphs by integrating command-line semantics, exposing obfuscation and low-frequency attacker behaviors.
  • The system combines modules like Attack-Clause Sketch and Attack-Evidence Awareness with advanced embedding and graph scoring techniques for real-time, analyst-usable investigation.

DEFENDCLI is a provenance-based Endpoint Detection and Response system that makes command-line activity a primary object of attack provenance examination rather than ancillary process metadata. Introduced in “DEFENDCLI: {Command-Line} Driven Attack Provenance Examination” (Wu et al., 18 Aug 2025), it is presented as the first provenance-graph EDR system to operate explicitly at the command-line level. Its central claim is that provenance alone often yields graph structures too noisy for operational triage, whereas provenance augmented with command-line semantics can expose obfuscation, low-frequency attacker behavior, and attack correlation with greater precision. The system evaluates anomalies across three coupled levels—unusual system process calls, suspicious command-line executions, and infrequent external network connections—and reports approximately 1.6x precision improvement over the strongest research baseline on DARPA Engagement-3 datasets, alongside a 2.3x precision improvement in industrial real-time testing (Wu et al., 18 Aug 2025).

1. Research motivation and problem formulation

DEFENDCLI is motivated by an operational failure mode in provenance-based detection: alerts may be technically correct yet insufficiently actionable. The paper states that, in enterprise deployment experience, only 0.07% of the original alerts from a provenance-based detector were truly actionable, because analysts were given large causal graphs without concentrated evidence about which command execution mattered and why (Wu et al., 18 Aug 2025). The system is therefore designed not merely for anomaly detection, but for analyst-usable reconstruction of attack activity.

The paper organizes its critique of prior work around four concerns: interoperability, reliability, flexibility, and practicability. These are then refined into four command-line-centric deficiencies: restricted obfuscation detection (OD), weak attack correlation (AC), poor handling of low-frequency events (LFD), and limited context awareness (CA). In this framing, process-graph structure alone is insufficient because the same process chain may be benign or malicious depending on its exact arguments, shell payloads, helper invocations, or network destinations. DEFENDCLI therefore treats command lines as semantically dense forensic objects, not merely attributes to be stored for later manual inspection (Wu et al., 18 Aug 2025).

A core premise is that EDR should move beyond binary benign/malicious judgments. The paper argues that modern attacks frequently manifest in gray zones where the decisive evidence lies in the interaction between process lineage, command syntax, and external connectivity. This suggests a shift from graph-centric anomaly scoring toward command-line-aware provenance differentiation, where graph structure remains essential but is continually reweighted by the semantics of the executed commands.

2. Process-centric provenance representation

DEFENDCLI adopts a refined command-line-focused provenance structure in which the graph is centered on uniform process nodes. Rather than building a heterogeneous provenance graph with explicit file, command, and network nodes, it stores command-line details and network information as attributes associated with process nodes, while omitting explicit file nodes to reduce redundancy and graph explosion (Wu et al., 18 Aug 2025).

The architecture is divided into two main modules: Attack-Clause Sketch and Attack-Evidence Awareness. The first performs graph construction, attack-irrelevance reduction, causal inference weighting, and interphase attack association. The second evaluates known malicious command-line patterns, differentiates unknown command lines, retrieves suspicious paths, and supports ranking and reporting. This process-centric representation is central to the paper’s argument that actionable EDR output depends on narrowing the graph to the portions that are both causally meaningful and directly verifiable by analysts (Wu et al., 18 Aug 2025).

Attack-irrelevance reduction occurs at two levels. At the node level, subgraphs without command-line elements are removed. At the edge level, loops and cycles are removed so that the provenance structure becomes a Directed Acyclic Graph. The rationale is explicitly pragmatic: command-line-bearing subgraphs are more useful for incident verification, and cyclic dependencies complicate causal reasoning. This design choice deprioritizes complete system reconstruction in favor of an attack-investigation graph biased toward analyst relevance.

The resulting graph is therefore not a universal provenance ledger. It is a task-shaped investigative substrate in which process interactions remain the backbone, but the discriminative content lies in attached command lines, associated network endpoints, and their role in weighted path retrieval.

3. Multi-level differentiation and attack-evidence awareness

The paper’s principal analytical contribution is its three-level differentiation framework:

  1. PC: unusual system process calls
  2. CL: suspicious command-line executions
  3. NC: infrequent external network connections

These three levels are treated as complementary rather than alternative signals (Wu et al., 18 Aug 2025).

At the process-call level, DEFENDCLI examines OS-level process interaction patterns such as unusual parent-child transitions, process relationships inconsistent with normal software behavior, or interaction chains suggestive of staged execution. This preserves the classic value of provenance graphs: causal reconstruction. However, the paper argues that such structure is rarely sufficient on its own, especially when adversaries hide behind legitimate binaries.

At the command-line level, the system performs its main differentiation step. Process chains that are only weakly suspicious structurally can become highly suspicious once their concrete command lines are examined. The paper emphasizes that some command-line artifacts are difficult to avoid even under evasion. A Turla-style chain, for example, may involve apparently legitimate programs such as Explorer.exe, Powershell, csc.exe, and WMI.exe, but the associated commands, task discovery calls such as tasklist, and linked external connection context can disclose malicious intent (Wu et al., 18 Aug 2025).

At the network level, DEFENDCLI uses infrequent external connections, represented mainly as IP:Port metadata, as corroborating evidence. Although the system does not rely on deep packet inspection, it links network destinations back to the responsible process lineage and command sequence. This is particularly relevant for command-and-control or exfiltration scenarios, where a rare outbound connection attached to a suspicious command chain substantially increases confidence.

Known-fact evaluation then refines these signals. DEFENDCLI checks process chains and command lines against Sigma rules and a curated defense-evasion rule repository using regular expressions. If a Sigma rule matches, the associated chain is assigned High-Risk Impact and an RSRS value. Evasion-rule matches receive Medium-Risk or Low-Risk values. The default risk scores are:

{H=2.5,  M=2.0,  L=1.5}.\{H = 2.5,\; M = 2.0,\; L = 1.5\}.

This rule layer is not the entire system; it functions as a structured prior that reweights graph evidence before anomaly ranking (Wu et al., 18 Aug 2025).

4. Graph scoring, command-line differentiation, and InfoPath ranking

DEFENDCLI’s causal scoring layer combines graph centrality with rule-based refinement and command-line anomaly analysis. The paper uses PageRank and shortest-path betweenness centrality as explainable graph measures. PageRank is written as:

PR(v)=(1−d)+d∑u∈In(v)PR(u)L(u),PR(v) = (1 - d) + d \sum_{u \in \text{In}(v)} \frac{PR(u)}{L(u)},

where dd is typically $0.85$, In(v)\text{In}(v) is the set of incoming neighbors, and L(u)L(u) is the outgoing degree of node uu (Wu et al., 18 Aug 2025).

The paper states that PageRank is inverted in interpretation: instead of emphasizing prominence, its normalized form is used to highlight rareness, so less-connected nodes become more suspicious. PageRank and betweenness are combined at the edge level through:

EW(m,n)=PR(m)+PR(n)+CB(m)+CB(n),EW_{(m,n)} = PR_{(m)} + PR_{(n)} + CB_{(m)} + CB_{(n)},

where CBCB denotes shortest-path betweenness centrality. Known command-line evidence then refines this edge weight as:

{H=2.5,  M=2.0,  L=1.5}.\{H = 2.5,\; M = 2.0,\; L = 1.5\}.0

Because suspicious edges receive larger refined weights but shortest-path algorithms minimize cost, DEFENDCLI inverts them during path retrieval:

{H=2.5,  M=2.0,  L=1.5}.\{H = 2.5,\; M = 2.0,\; L = 1.5\}.1

and, after community detection, further adjusts by community score:

{H=2.5,  M=2.0,  L=1.5}.\{H = 2.5,\; M = 2.0,\; L = 1.5\}.2

The community structure is produced with Leiden clustering rather than Louvain, because the paper argues that Leiden avoids weakly connected or disjoint communities and therefore improves interphase attack association (Wu et al., 18 Aug 2025).

The attack-relevant paths extracted from this weighted graph are called InfoPaths. They are retrieved with Dijkstra, with preference for source-to-sink traversal when zero in-degree and zero out-degree nodes are available. Final ranking is based on effective path length and the diversity of system processes involved.

Unknown command-line differentiation is performed separately from rule matching. DEFENDCLI embeds command-line instances using Word2Vec, Doc2Vec, and FastText, applies SimHash to compare their binary signatures, and uses this to choose the embedding that best separates command-line instances in a given environment (Wu et al., 18 Aug 2025). The resulting vectors are then evaluated by multiple unsupervised models, including Isolation Forest, One-Class SVM, Elliptic Envelope, KMeans, Gaussian Mixture, and a Local Outlier Factor / LoOP-related method, with a binary voting mechanism selecting anomalous InfoPaths. The paper’s interpretation is that this approach is more robust than relying on a single anomaly model or on structural graph deviation alone.

5. Empirical evaluation and operational performance

DEFENDCLI is evaluated on DARPA Transparent Computing Engagement-3 datasets—E3-TRACE, E3-THEIA, and E3-CADETS—and on an industrial real-time telemetry setting on Microsoft Azure using an “Industrial Sysmon Security Dataset.” The evaluation compares it with research baselines including Holmes, Unicorn, DepComm, and Nodlink, and with commercial EDR products including Microsoft Defender, Symantec EDR, and Kaspersky EDR (Wu et al., 18 Aug 2025).

The paper distinguishes three evaluation levels: InfoPath-level, node-level, and TTPs-level. The headline precision claims refer mainly to node-level results, because the major operational problem is not coarse attack visibility but reducing irrelevant nodes while preserving attack reconstruction.

Setting Nodlink precision DEFENDCLI precision
E3-TRACE 0.25 0.41
E3-THEIA 0.23 0.38
E3-CADETS 0.14 0.20
Industrial A1 0.04 0.08
Industrial A2 0.19 0.38
Industrial A3 0.05 0.18

These node-level results correspond to the paper’s summary that DEFENDCLI improves precision by approximately 1.6x over Nodlink on DARPA E3 and by 2.3x in industrial real-time scenarios (Wu et al., 18 Aug 2025). Recall remains at or near {H=2.5,  M=2.0,  L=1.5}.\{H = 2.5,\; M = 2.0,\; L = 1.5\}.3 in these comparisons, so the improvement is primarily in precision and therefore in actionability. The paper’s interpretation is that both Nodlink and DEFENDCLI can detect attacks at a broad TTP level, but DEFENDCLI provides a much tighter, command-line-centered explanation of what happened.

The industrial evaluation further reports that DEFENDCLI detects attack instances missed by several commercial systems. In the three industrial scenarios, Microsoft Defender was bypassed in all three, Symantec EDR failed on scenarios 2 and 3, and Kaspersky EDR was bypassed by scenario 3, whereas DEFENDCLI and Nodlink both retained full TTP-level detection, with DEFENDCLI yielding superior node-level precision (Wu et al., 18 Aug 2025). This result is central to the paper’s practical argument: command-line-aware provenance can surface attacks that remain invisible to systems emphasizing signatures, coarse behavior classes, or opaque anomaly scoring.

Operationally, the paper reports runtime on a 64GB RAM, 16-core VM of about 5 seconds per 5,000 records for E3-TRACE and E3-THEIA, and about 7 seconds per 5,000 records for E3-CADETS. In industrial real-time deployment on Azure DataLake with Kafka-based streaming and ETL, the reported end-to-end response time is about 2 seconds, including network latency (Wu et al., 18 Aug 2025). These measurements support the claim that DEFENDCLI is intended for live EDR workflows rather than purely offline forensic analysis.

6. Interpretation, adjacent developments, and limitations

DEFENDCLI’s primary contribution is conceptual as much as technical: it argues that command-line context changes both detection and investigation. In the paper’s view, traditional provenance systems often reveal only that suspicious process structure exists, whereas DEFENDCLI attempts to show which exact command execution rendered that structure malicious, how it related to the surrounding process chain, and whether rare outbound communication corroborated the interpretation (Wu et al., 18 Aug 2025). This makes the system more explainable than opaque AI detectors and more specific than graph-only anomaly systems.

The paper is careful, however, not to present DEFENDCLI as a universal detector. It explicitly states that the system is not a silver bullet and should be combined with antivirus, network intrusion detection, and broader SOC tooling (Wu et al., 18 Aug 2025). Several limitations are also acknowledged. The method depends on audit-log observability; if command, process, or network telemetry is unavailable or incomplete, detection quality degrades. The paper does not fully specify the command-line parsing and normalization grammar, including shell-specific canonicalization, variable expansion handling, or quoting semantics. Some key components—such as the exact community score, the precise binary voting rule, the formal SimHash-based embedding-selection criterion, and the final InfoPath ranking equation—are described operationally but not given as explicit mathematical definitions. The adaptive interface also indicates deployment sensitivity: the major hyperparameters {H=2.5,  M=2.0,  L=1.5}.\{H = 2.5,\; M = 2.0,\; L = 1.5\}.4, {H=2.5,  M=2.0,  L=1.5}.\{H = 2.5,\; M = 2.0,\; L = 1.5\}.5, {H=2.5,  M=2.0,  L=1.5}.\{H = 2.5,\; M = 2.0,\; L = 1.5\}.6, and the risk-score defaults {H=2.5,  M=2.0,  L=1.5}.\{H = 2.5,\; M = 2.0,\; L = 1.5\}.7 must be tuned against environmental drift, and the paper notes the expected trade-offs between detection rate and false alarm rate (Wu et al., 18 Aug 2025).

Later work on agent and CLI security addresses different threat models but points in a related direction. VIGIL treats tool outputs and runtime feedback as an untrusted “tool stream” that should not directly control action commitment (Lin et al., 9 Jan 2026). MOSAIC formalizes CLI command-composition risk as dangerous producer-consumer dependencies over shared OS state (Wu et al., 3 Jul 2026). OpenClaw PRISM distributes runtime policy enforcement across lifecycle hooks for tool-augmented agents (Li, 12 Mar 2026). DualView extends untrusted-data tracking into file system, shell, network, and inter-agent channels by maintaining separate AgentView and Human View environments (Kim et al., 4 Jul 2026). These systems operate in different settings, but they suggest a broader implication: command execution, command output, and tool-mediated runtime state are increasingly being treated as first-class security surfaces, not mere byproducts of higher-level agent reasoning.

DEFENDCLI’s enduring significance lies in that shift. By making command-line activity the central explanatory unit inside provenance examination, it reframes endpoint detection from generic behavioral anomaly detection toward causally grounded, semantically interpretable attack evidence. In the paper’s own terms, the gain is not only higher precision, but a more reliable path from system telemetry to an analyst’s actionable conclusion (Wu et al., 18 Aug 2025).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to DEFENDCLI.