Papers
Topics
Authors
Recent
Search
2000 character limit reached

Graphlet Degree Sequence Analysis

Updated 9 July 2026
  • Graphlet Degree Sequence is a graph invariant that counts per-vertex occurrences of specific rooted subgraph patterns, thereby generalizing the traditional degree sequence.
  • It refines standard degree information by capturing detailed positional roles in various configurations such as paths, triangles, and cyclic structures.
  • GDS supports graph reconstruction by linking large rooted graphlets to smaller ones and extends to applications in directed, temporal, and scalable network analysis.

Searching arXiv for recent and foundational papers on graphlet degree sequence and related graphlet methodologies. Graphlet Degree Sequence (GDS) is a graph invariant built from per-vertex counts of small induced subgraphs in which the vertex occupies a specified rooted or orbit position. In the formulation developed for graph reconstruction, a graphlet is a connected rooted induced subgraph (G,r)(G,r), and the k\le k-graphlet degree sequence of a vertex vv is the vector of counts of all rooted graphlet isomorphism classes of size at most kk that occur with root mapped to vv. The corresponding graphlet degree distribution (gdd) is the matrix obtained by stacking these vectors over all vertices. In that sense, GDS generalizes the ordinary degree sequence from edge incidence to rooted small-subgraph incidence, and in recent work it is treated not only as a local structural signature but also as an invariant with graph-theoretic reconstruction power (Hartman et al., 26 Aug 2025).

1. Formal framework

Let H=(V,E)H=(V,E) be a finite simple undirected graph with n=Vn=|V|. In the rooted formulation, a graphlet is a pair (G,r)(G,r) where GG is a connected subgraph of HH, k\le k0, and k\le k1. The root is part of the isomorphism type: k\le k2 means not only that k\le k3, but also that the isomorphism maps k\le k4 to k\le k5. Thus vertices belonging to different automorphism orbits of the same underlying graph induce different rooted graphlets (Hartman et al., 26 Aug 2025).

For a vertex k\le k6, the graphlet degree with respect to k\le k7 is the number of rooted induced embeddings of k\le k8 in k\le k9 with vv0 mapped to vv1. If vv2 denotes the vv3-th rooted graphlet isomorphism class under a fixed indexing, then the vv4-graphlet degree sequence of vv5 is

vv6

while the whole-graph vv7-graphlet degree distribution is the matrix

vv8

The row-level object is the gds; the matrix-level object is the gdd. Because graphlets are proper connected rooted subgraphs, the largest admissible size in an vv9-vertex graph is kk0, so “graphlets up to size kk1” means all connected rooted induced subgraphs on kk2 vertices (Hartman et al., 26 Aug 2025).

This rooted definition is stricter than a purely unrooted small-subgraph census. The root preserves positional information, and the indexing system distinguishes graphlets not only by underlying subgraph type but also by root placement within that type. A GDS is therefore best understood as a local rooted induced-subgraph profile.

2. Informational role and relation to adjacent descriptors

A central point in the literature is that GDS strictly refines degree information. In the rooted framework, ordinary degree is exactly the kk3-graphlet information, since it counts only graphlets whose underlying graph is kk4. The recent reconstruction paper states explicitly that the kk5-gdd, i.e. the degree sequence, is not a good representation because many nonisomorphic graphs share it. GDS strengthens this by recording how a vertex participates in paths, triangles, tree-like configurations, cyclic configurations, and more generally every rooted connected induced pattern up to the chosen size bound (Hartman et al., 26 Aug 2025).

The same distinction separates GDS from motif counts. Motif distributions are unrooted whole-graph counts; GDS retains where each pattern occurs. Aggregating over roots recovers motif-style information, but the row-by-row rooted representation is strictly richer. One example shows two graphs with the same kk6 motif distribution but different kk7-gdd, demonstrating that rooted positional information survives aggregation losses that erase vertex roles (Hartman et al., 26 Aug 2025).

Across the graphlet literature, the terminology is not completely uniform. Orbit-based work on static graphlets usually speaks of graphlet degree vectors (GDVs), graphlet degree distributions (GDDs), and graphlet degree agreement (GDA), while directed and temporal extensions sometimes use alternative names such as “signature vector” or shift attention from static counts to orbit transitions (Aparício et al., 2017). In directed networks, a reduced per-node “signature vector” built from directed graphlets starting at a node plays the role of a directed analogue of a GDV, although the paper does not explicitly define a directed GDS or GDD (Trpevski et al., 2016). By contrast, the neighborhood degree list is a degree-based local invariant that refines the degree sequence and joint degree matrix but remains degree-centric rather than graphlet-centric; in the supplied comparison, it is presented as weaker than graphlet orbit counts (Barrus et al., 2015).

Paper Setting Core per-node object
(Hartman et al., 26 Aug 2025) Undirected, rooted induced graphlets up to kk8 kk9-gds and vv0-gdd
(Trpevski et al., 2016) Directed, weakly connected 2- and 3-node graphlets 16-dimensional signature vector
(Aparício et al., 2017) Temporal snapshots, undirected 4-node graphlet orbits GDV/GDD baseline plus orbit transitions
(Wang et al., 2016) Large directed and undirected graphs Estimated node graphlet orbit degrees

This variety suggests that “Graphlet Degree Sequence” names a family of closely related node-level subgraph-incidence descriptors rather than a single universally standardized object.

3. Reconstruction-theoretic content

The most extensive recent theory treats GDS and GDD as reconstruction-relevant invariants rather than only descriptive statistics. For a connected graph vv1 on vv2 vertices, the vv3-vertex rooted graphlets correspond to rooted versions of vertex-deleted subgraphs whenever deletion leaves the graph connected. That observation underlies several exact results (Hartman et al., 26 Aug 2025).

First, the paper proves extremal bounds. If vv4 is the vv5-gdd of vv6, then

vv7

with equality for vv8. Summing over coordinates gives

vv9

At the opposite extreme, among connected graphs on H=(V,E)H=(V,E)0,

H=(V,E)H=(V,E)1

and this minimum is achieved by H=(V,E)H=(V,E)2 and H=(V,E)H=(V,E)3. These bounds quantify how concentrated or sparse rooted subgraph multiplicities can be.

Second, sufficiently large graphlets determine smaller ones under connectivity assumptions. If H=(V,E)H=(V,E)4 is connected and H=(V,E)H=(V,E)5-vertex connected, then the H=(V,E)H=(V,E)6-gdd determines the H=(V,E)H=(V,E)7-gdd. More generally, if H=(V,E)H=(V,E)8 is H=(V,E)H=(V,E)9-vertex connected, then the n=Vn=|V|0-graphlet degree sequence determines the n=Vn=|V|1-graphlet degree sequence. The counting principle is explicit: a rooted graphlet on n=Vn=|V|2 vertices appears in exactly n=Vn=|V|3 different n=Vn=|V|4-graphlets in the n=Vn=|V|5-connected case, and in

n=Vn=|V|6

graphlets of size n=Vn=|V|7 in the n=Vn=|V|8-connected case.

Third, graphlet counts characterize connectivity itself. The paper proves that n=Vn=|V|9 is (G,r)(G,r)0-vertex connected if and only if

(G,r)(G,r)1

For (G,r)(G,r)2, a corollary states that if (G,r)(G,r)3 is the (G,r)(G,r)4-gds, then

(G,r)(G,r)5

holds for all vertices when there is no articulation, for exactly one vertex when there is exactly one articulation, and for no vertex when there are multiple articulations. In that sense, large rooted graphlets detect articulation structure exactly.

Fourth, the (G,r)(G,r)6-gdd of a (G,r)(G,r)7-vertex connected graph determines its reconstruction deck as defined by Kelly. This is a strong comparison point with classical reconstruction theory, because the rooted enhancement retains more information than the unrooted vertex-deleted deck.

Finally, the paper gives two explicit reconstruction theorems. Trees on (G,r)(G,r)8 vertices are uniquely reconstructible from their (G,r)(G,r)9-gds. More generally, a GG0-vertex connected graph GG1 is reconstructible from its GG2-gdd if it contains a vertex GG3 such that GG4 is rigid and an additional orbit condition controls ambiguities among vertices whose deletion yields isomorphic rigid graphs. A simpler sufficient condition is that such vertices have identical neighborhoods in GG5; in that case the theorem’s orbit condition holds, and the number of such vertices is at most two.

4. Constructive mechanisms and identities

The reconstruction results are accompanied by explicit mechanisms. For GG6-connected graphs, the key observation is that every vertex deletion remains connected, so every deleted graph appears as an GG7-graphlet. Because the graphlets are rooted, the GG8-gdd does not merely record the multiset of deleted subgraphs; it records, for each surviving vertex, its rooted role inside each deleted subgraph. This makes the GG9-gdd a rooted enhancement of the deck (Hartman et al., 26 Aug 2025).

The tree algorithm is especially transparent. A path-end graphlet is a rooted graphlet whose underlying graph is HH0 and whose root is an endpoint. For a vertex HH1, define HH2 as the size of the largest path-end graphlet rooted at HH3. The paper proves

HH4

so the tree center is obtained by minimizing HH5. At the center HH6, one then considers trunked tree graphlets HH7 with HH8 a tree and HH9. The inclusion-wise maximal such graphlets correspond exactly to the components of the forest induced by k\le k00. A greedy subtraction procedure repeatedly extracts the largest remaining trunked tree graphlet from the center’s row of the gdd, adds its underlying branch to the reconstruction, subtracts the rooted counts contributed by that branch, and iterates until no trunked tree graphlet remains. The branch structure is therefore localized directly at the center, which is why the proof is described as much easier than classical reconstruction from the unrooted deck.

For the asymmetric rigid-subgraph theorem, the constructive steps are also explicit. One first identifies a rigid k\le k01-graphlet k\le k02 from a nonzero gdd column. Rooted data then identifies, for each k\le k03, which rooted copy of k\le k04 it corresponds to. If several deletions yield the same rigid graph, the missing vertices are detected by a count discrepancy: for a vertex k\le k05, the relevant rooted count is k\le k06, while for k\le k07 it is k\le k08. Finally, adjacency between the restored vertex and each vertex of k\le k09 is recovered by degree comparison between k\le k10 and k\le k11.

The same paper also records several exact local identities, including

k\le k12

which couples ordinary degree to rooted path and triangle counts, and

k\le k13

a double-counting identity for edge-triangle incidences. These formulas show that small graphlet coordinates are not independent; they satisfy nontrivial combinatorial consistency constraints.

5. Directed, temporal, and scalable variants

In directed networks, the core idea survives but the combinatorics expand sharply. One paper defines a directed graphlet as an induced weakly connected subgraph and classifies neighbors of a node k\le k14 into in-, out-, and reciprocal sets k\le k15, k\le k16, and k\le k17. Although the paper is orbit-aware and notes that directed graphlets on up to 4 nodes induce 1695 orbit combinations, its main method compresses the description to a per-node signature vector

k\le k18

of dimension k\le k19: k\le k20 degree features, k\le k21 wedge-degree features, and k\le k22 triangle-degree features. The raw counts are k\le k23, then reduced by grouping isomorphic wedge and triangle classes. Triangle counts satisfy

k\le k24

and wedge counts are obtained by subtracting triangle closures from all candidate wedges. The whole-network summary is not an explicit directed GDS but a graphlet correlation matrix, obtained by computing Pearson correlations between the k\le k25 signature coordinates across vertices (Trpevski et al., 2016).

Temporal graphlet methodology shifts the emphasis from static participation counts to role changes across snapshots. A temporal network is represented as a sequence of snapshots, and all k\le k26-node undirected graphlet orbits are enumerated in each snapshot. For each fixed node set appearing across consecutive snapshots, the method counts how node orbit memberships change, storing the results in an orbit-transition matrix k\le k27. This is a network-level summary of orbit-to-orbit dynamics rather than a per-node GDS. The proposed comparison measure, Orbit-Transition-Agreement (OTA), is therefore best viewed as a temporal complement to GDV/GDD-style static descriptors: static GDS captures marginal orbit participation, whereas orbit-transition matrices capture conditional temporal evolution of those roles (Aparício et al., 2017).

For large static graphs, sampling-based estimation becomes necessary. One paper studies node graphlet orbit degrees directly. For a node k\le k28, if k\le k29 is the set of connected induced subgraphs touching k\le k30 in orbit k\le k31, then

k\le k32

is the node’s graphlet orbit degree. The paper proposes SAND for undirected 3- and 4-node orbit degrees and SAND-3D for directed 3-node orbit degrees. If k\le k33 samples are drawn, k\le k34 of them land in orbit k\le k35, and each orbit-k\le k36 sample has probability k\le k37, then

k\le k38

is an unbiased estimator, with explicit variance and covariance formulas. In this framework, GDS is not the direct output, but it is immediate to obtain orbit-wise graphlet degree sequences by collecting k\le k39 across vertices (Wang et al., 2016).

6. Limits, ambiguities, and open problems

Despite its expressive power, GDS is not universally complete. The recent reconstruction paper is explicit that k\le k40-gdd is not proved unique for all graphs, and the general graphlet reconstruction conjecture remains open (Hartman et al., 26 Aug 2025). Equal full k\le k41-gds values can occur for vertices in different ambient graphs: one construction begins with a path on k\le k42 vertices and ends it either with a triangle or a fork, producing distinguished vertices whose full rooted graphlet profiles coincide. This shows that even complete per-vertex graphlet information need not identify the surrounding graph uniquely.

The same work stresses that uniqueness of gdd is independent of reconstructibility. A gdd could in principle be unique without providing an easy reconstruction method, and constructive recovery may remain difficult even when uniqueness holds. The paper also notes lower bounds on how large k\le k43 must be for uniqueness-type hopes, via Nýdl’s result on graphs with identical small induced subgraphs.

Connectivity assumptions are essential in many results because the graphlets are connected. If deleting a vertex disconnects the graph, only the component containing the root is visible in rooted graphlet data, so information present in the full vertex-deleted graph can be lost. This explains why several recovery theorems require k\le k44-vertex connectivity.

Realizability is likewise incomplete. The paper discusses the problem of deciding whether a given matrix is a valid k\le k45-gdd and derives necessary combinatorial identities, but it does not solve general realizability. This places GDS alongside classical degree-sequence theory in one respect: local count data admit strong constraints, yet the full characterization problem remains harder than the definition suggests.

A final source of ambiguity is terminological rather than mathematical. Static undirected work often frames the subject in terms of GDV, GDD, and GDA; directed work may speak of signature vectors; temporal work may elevate orbit-transition matrices instead of static degree sequences (Aparício et al., 2017). This suggests that the concept is stable at the level of principle—per-node graphlet-role incidence—but variable in its exact formal packaging.

Graphlet Degree Sequence therefore occupies a distinctive position among graph invariants. It is richer than degree-based descriptors, more localized than global motif counts, and, in some regimes, strong enough to encode connectivity and support reconstruction. At the same time, its full uniqueness, realizability, and temporal generalization remain active structural questions rather than closed theory.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Graphlet Degree Sequence.