---
title: Graphlet Degree Sequence Analysis
url: https://www.emergentmind.com/topics/graphlet-degree-sequence
type: topic
---

# Graphlet Degree Sequence Analysis

Searching arXiv for recent and foundational papers on graphlet degree sequence and related graphlet methodologies.
Graphlet Degree Sequence (GDS) is a graph invariant built from per-vertex counts of small induced subgraphs in which the vertex occupies a specified rooted or orbit position. In the formulation developed for graph reconstruction, a graphlet is a connected rooted induced subgraph \((G,r)\), and the \(\le k\)-graphlet degree sequence of a vertex \(v\) is the vector of counts of all rooted graphlet isomorphism classes of size at most \(k\) that occur with root mapped to \(v\). The corresponding graphlet degree distribution (gdd) is the matrix obtained by stacking these vectors over all vertices. In that sense, GDS generalizes the ordinary degree sequence from edge incidence to rooted small-subgraph incidence, and in recent work it is treated not only as a local structural signature but also as an invariant with graph-theoretic reconstruction power [2508.19189].

## 1. Formal framework

Let \(H=(V,E)\) be a finite simple undirected graph with \(n=|V|\). In the rooted formulation, a graphlet is a pair \((G,r)\) where \(G\) is a connected subgraph of \(H\), \(|V(G)|<n\), and \(r\in V(G)\). The root is part of the isomorphism type: \((G,r)\cong(H,s)\) means not only that \(G\cong H\), but also that the isomorphism maps \(r\) to \(s\). Thus vertices belonging to different automorphism orbits of the same underlying graph induce different rooted graphlets [2508.19189].

For a vertex \(v\in V(H)\), the graphlet degree with respect to \((G,r)\) is the number of rooted induced embeddings of \((G,r)\) in \(H\) with \(r\) mapped to \(v\). If \(G_i\) denotes the \(i\)-th rooted graphlet isomorphism class under a fixed indexing, then the \(\le k\)-graphlet degree sequence of \(v\) is
\[
\big(\#_H(G_i,v)\ \big|\ i\in \vartheta(G'_{\le k})\big),
\]
while the whole-graph \(\le k\)-graphlet degree distribution is the matrix
\[
D \in \mathbb{N}^{|V(H)| \times |\vartheta(G'_{\le k})|}, \qquad D_{i,j} = \#_H(G_j,i).
\]
The row-level object is the gds; the matrix-level object is the gdd. Because graphlets are proper connected rooted subgraphs, the largest admissible size in an \(n\)-vertex graph is \(n-1\), so “graphlets up to size \(n-1\)” means all connected rooted induced subgraphs on \(1,\dots,n-1\) vertices [2508.19189].

This rooted definition is stricter than a purely unrooted small-subgraph census. The root preserves positional information, and the indexing system distinguishes graphlets not only by underlying subgraph type but also by root placement within that type. A GDS is therefore best understood as a local rooted induced-subgraph profile.

## 2. Informational role and relation to adjacent descriptors

A central point in the literature is that GDS strictly refines degree information. In the rooted framework, ordinary degree is exactly the \(2\)-graphlet information, since it counts only graphlets whose underlying graph is \(K_2\). The recent reconstruction paper states explicitly that the \(2\)-gdd, i.e. the degree sequence, is not a good representation because many nonisomorphic graphs share it. GDS strengthens this by recording how a vertex participates in paths, triangles, tree-like configurations, cyclic configurations, and more generally every rooted connected induced pattern up to the chosen size bound [2508.19189].

The same distinction separates GDS from motif counts. Motif distributions are unrooted whole-graph counts; GDS retains where each pattern occurs. Aggregating over roots recovers motif-style information, but the row-by-row rooted representation is strictly richer. One example shows two graphs with the same \(\le 4\) motif distribution but different \(\le 4\)-gdd, demonstrating that rooted positional information survives aggregation losses that erase vertex roles [2508.19189].

Across the graphlet literature, the terminology is not completely uniform. Orbit-based work on static graphlets usually speaks of graphlet degree vectors (GDVs), graphlet degree distributions (GDDs), and graphlet degree agreement (GDA), while directed and temporal extensions sometimes use alternative names such as “signature vector” or shift attention from static counts to orbit transitions [1707.04572]. In directed networks, a reduced per-node “signature vector” built from directed graphlets starting at a node plays the role of a directed analogue of a GDV, although the paper does not explicitly define a directed GDS or GDD [1603.05843]. By contrast, the neighborhood degree list is a degree-based local invariant that refines the degree sequence and joint degree matrix but remains degree-centric rather than graphlet-centric; in the supplied comparison, it is presented as weaker than graphlet orbit counts [1507.08212].

| Paper | Setting | Core per-node object |
|---|---|---|
| [2508.19189] | Undirected, rooted induced graphlets up to \(n-1\) | \(\le k\)-gds and \(\le k\)-gdd |
| [1603.05843] | Directed, weakly connected 2- and 3-node graphlets | 16-dimensional signature vector |
| [1707.04572] | Temporal snapshots, undirected 4-node graphlet orbits | GDV/GDD baseline plus orbit transitions |
| [1604.08691] | Large directed and undirected graphs | Estimated node graphlet orbit degrees |

This variety suggests that “Graphlet Degree Sequence” names a family of closely related node-level subgraph-incidence descriptors rather than a single universally standardized object.

## 3. Reconstruction-theoretic content

The most extensive recent theory treats GDS and GDD as reconstruction-relevant invariants rather than only descriptive statistics. For a connected graph \(H\) on \(n\) vertices, the \((n-1)\)-vertex rooted graphlets correspond to rooted versions of vertex-deleted subgraphs whenever deletion leaves the graph connected. That observation underlies several exact results [2508.19189].

First, the paper proves extremal bounds. If \(D^{n\times k}\) is the \((\le n-1)\)-gdd of \(G\), then
\[
D_{i,j} \le {n-1 \choose |G_j|-1},
\]
with equality for \(G\cong K_n\). Summing over coordinates gives
\[
\sum D_{i,*}=\sum_{j\le k} D_{i,j} \le \sum_{p\le n-1} {n-1 \choose p-1}.
\]
At the opposite extreme, among connected graphs on \(n\ge 3\),
\[
\min_{G\in G_n}\ \max_{i\in[n],\,j\in[k]} D^G_{i,j}=2,
\]
and this minimum is achieved by \(P_n\) and \(C_n\). These bounds quantify how concentrated or sparse rooted subgraph multiplicities can be.

Second, sufficiently large graphlets determine smaller ones under connectivity assumptions. If \(H\) is connected and \(2\)-vertex connected, then the \((n-1)\)-gdd determines the \((\le n-2)\)-gdd. More generally, if \(H\) is \(k\)-vertex connected, then the \((n-k+1)\)-graphlet degree sequence determines the \((\le n-k)\)-graphlet degree sequence. The counting principle is explicit: a rooted graphlet on \(\ell\) vertices appears in exactly \(n-\ell\) different \((n-1)\)-graphlets in the \(2\)-connected case, and in
\[
{n-\ell \choose k-1}
\]
graphlets of size \(n-k+1\) in the \(k\)-connected case.

Third, graphlet counts characterize connectivity itself. The paper proves that \(H\) is \(k\)-vertex connected if and only if
\[
\sum_{(G,j)\in G'_{n-k+1}} \#_H((G,j),v) = {n-1 \choose k-1}
\qquad\text{for every } v\in V(H).
\]
For \(k=2\), a corollary states that if \(D\) is the \((n-1)\)-gds, then
\[
\sum_{i=1}^m D_{v,i}=n-1
\]
holds for all vertices when there is no articulation, for exactly one vertex when there is exactly one articulation, and for no vertex when there are multiple articulations. In that sense, large rooted graphlets detect articulation structure exactly.

Fourth, the \((n-1)\)-gdd of a \(2\)-vertex connected graph determines its reconstruction deck as defined by Kelly. This is a strong comparison point with classical reconstruction theory, because the rooted enhancement retains more information than the unrooted vertex-deleted deck.

Finally, the paper gives two explicit reconstruction theorems. Trees on \(n\ge 3\) vertices are uniquely reconstructible from their \((\le n-1)\)-gds. More generally, a \(2\)-vertex connected graph \(H\) is reconstructible from its \((n-1)\)-gdd if it contains a vertex \(v\) such that \(H\setminus\{v\}\) is rigid and an additional orbit condition controls ambiguities among vertices whose deletion yields isomorphic rigid graphs. A simpler sufficient condition is that such vertices have identical neighborhoods in \(H\); in that case the theorem’s orbit condition holds, and the number of such vertices is at most two.

## 4. Constructive mechanisms and identities

The reconstruction results are accompanied by explicit mechanisms. For \(2\)-connected graphs, the key observation is that every vertex deletion remains connected, so every deleted graph appears as an \((n-1)\)-graphlet. Because the graphlets are rooted, the \((n-1)\)-gdd does not merely record the multiset of deleted subgraphs; it records, for each surviving vertex, its rooted role inside each deleted subgraph. This makes the \((n-1)\)-gdd a rooted enhancement of the deck [2508.19189].

The tree algorithm is especially transparent. A path-end graphlet is a rooted graphlet whose underlying graph is \(P_k\) and whose root is an endpoint. For a vertex \(v\), define \(lp(v)\) as the size of the largest path-end graphlet rooted at \(v\). The paper proves
\[
lp(v)=\varepsilon(v),
\]
so the tree center is obtained by minimizing \(lp(v)\). At the center \(c\), one then considers trunked tree graphlets \((G,c)\) with \(G\) a tree and \(\deg_G(c)=1\). The inclusion-wise maximal such graphlets correspond exactly to the components of the forest induced by \(V(T)\setminus\{c\}\). A greedy subtraction procedure repeatedly extracts the largest remaining trunked tree graphlet from the center’s row of the gdd, adds its underlying branch to the reconstruction, subtracts the rooted counts contributed by that branch, and iterates until no trunked tree graphlet remains. The branch structure is therefore localized directly at the center, which is why the proof is described as much easier than classical reconstruction from the unrooted deck.

For the asymmetric rigid-subgraph theorem, the constructive steps are also explicit. One first identifies a rigid \((n-1)\)-graphlet \(F=H\setminus\{v\}\) from a nonzero gdd column. Rooted data then identifies, for each \(w\in V(F)\), which rooted copy of \(F\) it corresponds to. If several deletions yield the same rigid graph, the missing vertices are detected by a count discrepancy: for a vertex \(w\notin\{v^1,\dots,v^k\}\), the relevant rooted count is \(k\), while for \(w=v^i\) it is \(k-1\). Finally, adjacency between the restored vertex and each vertex of \(F\) is recovered by degree comparison between \(H\) and \(F\).

The same paper also records several exact local identities, including
\[
\deg(v) = \sum_{u\in N_H(v)} \#_H(G_0,u) - \#_H(G_1,v) - 2\#_H(G_3,v),
\]
which couples ordinary degree to rooted path and triangle counts, and
\[
\sum_{e=\{u,v\}\in E(H)} c(u,v) = \sum_{v\in V(G)} \#_H(G_3,v),
\]
a double-counting identity for edge-triangle incidences. These formulas show that small graphlet coordinates are not independent; they satisfy nontrivial combinatorial consistency constraints.

## 5. Directed, temporal, and scalable variants

In directed networks, the core idea survives but the combinatorics expand sharply. One paper defines a directed graphlet as an induced weakly connected subgraph and classifies neighbors of a node \(i\) into in-, out-, and reciprocal sets \(S_i^-\), \(S_i^+\), and \(S_i^\circ\). Although the paper is orbit-aware and notes that directed graphlets on up to 4 nodes induce 1695 orbit combinations, its main method compresses the description to a per-node signature vector
\[
F_i=[d_i,W_i,T_i]^T
\]
of dimension \(16\): \(3\) degree features, \(6\) wedge-degree features, and \(7\) triangle-degree features. The raw counts are \(3+9+27=39\), then reduced by grouping isomorphic wedge and triangle classes. Triangle counts satisfy
\[
T_i(\alpha,\beta,\gamma) = \sum_{j\in S_i^\gamma} |S_i^\alpha \cap S_j^\beta|,
\]
and wedge counts are obtained by subtracting triangle closures from all candidate wedges. The whole-network summary is not an explicit directed GDS but a graphlet correlation matrix, obtained by computing Pearson correlations between the \(16\) signature coordinates across vertices [1603.05843].

Temporal graphlet methodology shifts the emphasis from static participation counts to role changes across snapshots. A temporal network is represented as a sequence of snapshots, and all \(4\)-node undirected graphlet orbits are enumerated in each snapshot. For each fixed node set appearing across consecutive snapshots, the method counts how node orbit memberships change, storing the results in an orbit-transition matrix \(u\mathcal{T}_4\). This is a network-level summary of orbit-to-orbit dynamics rather than a per-node GDS. The proposed comparison measure, Orbit-Transition-Agreement (OTA), is therefore best viewed as a temporal complement to GDV/GDD-style static descriptors: static GDS captures marginal orbit participation, whereas orbit-transition matrices capture conditional temporal evolution of those roles [1707.04572].

For large static graphs, sampling-based estimation becomes necessary. One paper studies node graphlet orbit degrees directly. For a node \(v\), if \(C_v^{(i)}\) is the set of connected induced subgraphs touching \(v\) in orbit \(i\), then
\[
d_v^{(i)} = |C_v^{(i)}|
\]
is the node’s graphlet orbit degree. The paper proposes SAND for undirected 3- and 4-node orbit degrees and SAND-3D for directed 3-node orbit degrees. If \(K\) samples are drawn, \(m_i\) of them land in orbit \(i\), and each orbit-\(i\) sample has probability \(p_i\), then
\[
\hat d_v^{(i)}=\frac{m_i}{Kp_i}
\]
is an unbiased estimator, with explicit variance and covariance formulas. In this framework, GDS is not the direct output, but it is immediate to obtain orbit-wise graphlet degree sequences by collecting \(\hat d_v^{(i)}\) across vertices [1604.08691].

## 6. Limits, ambiguities, and open problems

Despite its expressive power, GDS is not universally complete. The recent reconstruction paper is explicit that \((\le n-1)\)-gdd is not proved unique for all graphs, and the general graphlet reconstruction conjecture remains open [2508.19189]. Equal full \((\le n-1)\)-gds values can occur for vertices in different ambient graphs: one construction begins with a path on \(n-2\) vertices and ends it either with a triangle or a fork, producing distinguished vertices whose full rooted graphlet profiles coincide. This shows that even complete per-vertex graphlet information need not identify the surrounding graph uniquely.

The same work stresses that uniqueness of gdd is independent of reconstructibility. A gdd could in principle be unique without providing an easy reconstruction method, and constructive recovery may remain difficult even when uniqueness holds. The paper also notes lower bounds on how large \(k\) must be for uniqueness-type hopes, via Nýdl’s result on graphs with identical small induced subgraphs.

Connectivity assumptions are essential in many results because the graphlets are connected. If deleting a vertex disconnects the graph, only the component containing the root is visible in rooted graphlet data, so information present in the full vertex-deleted graph can be lost. This explains why several recovery theorems require \(2\)-vertex connectivity.

Realizability is likewise incomplete. The paper discusses the problem of deciding whether a given matrix is a valid \(\{2,3\}\)-gdd and derives necessary combinatorial identities, but it does not solve general realizability. This places GDS alongside classical degree-sequence theory in one respect: local count data admit strong constraints, yet the full characterization problem remains harder than the definition suggests.

A final source of ambiguity is terminological rather than mathematical. Static undirected work often frames the subject in terms of GDV, GDD, and GDA; directed work may speak of signature vectors; temporal work may elevate orbit-transition matrices instead of static degree sequences [1707.04572]. This suggests that the concept is stable at the level of principle—per-node graphlet-role incidence—but variable in its exact formal packaging.

Graphlet Degree Sequence therefore occupies a distinctive position among graph invariants. It is richer than degree-based descriptors, more localized than global motif counts, and, in some regimes, strong enough to encode connectivity and support reconstruction. At the same time, its full uniqueness, realizability, and temporal generalization remain active structural questions rather than closed theory.

Source: https://www.emergentmind.com/topics/graphlet-degree-sequence