Vertex-Wise Flexible Graph Sampling
- Vertex-wise flexible sampling is a graph-based strategy that allocates sampling budgets heterogeneously across vertices to satisfy exact recovery or rank conditions.
- It enables efficient reconstruction of bandlimited time-vertex signals by allowing flexible per-vertex sampling rates, thereby reducing reconstruction errors and computational costs.
- This framework extends to learning applications, integrating with network embedding and GNN training to adaptively boost performance through strategic, vertex-specific sample selection.
Vertex-wise flexible sampling denotes a family of graph-based sampling schemes in which sampling effort is allowed to vary across vertices rather than being imposed uniformly. In the cited literature, the term appears in several technically distinct settings: time-vertex graph signal processing, generalized graph-signal reconstruction, weighted network embedding, graph neural network training, active few-shot annotation, and design-based graph spatial sampling. Across these settings, the common principle is heterogeneous allocation of samples, measurements, or query probabilities over vertices so as to satisfy recovery conditions, reduce reconstruction error, or improve downstream learning efficiency under resource constraints (Yu et al., 2019, Yamashita et al., 18 Sep 2025, Chen et al., 2017, Oh et al., 2019, Burr et al., 25 Apr 2025).
1. Scope and formal characterizations
A canonical formalization arises in joint time-vertex signal processing. Let be an undirected graph with Laplacian , and let be the cycle graph on time nodes with Laplacian . The joint graph is the Cartesian product with Laplacian
and a joint time-vertex signal is vectorized as . Sampling is represented by a binary operator selecting a subset 0, yielding 1. In the bandlimited setting, exact recovery is possible exactly when the reduced Fourier basis 2 associated with the active spectral coefficients satisfies
3
where 4 is the general bandwidth (Yu et al., 2019).
A second formalization appears in generalized graph sampling. There, a graph signal 5 is sampled through a designed operator 6, producing 7, and reconstructed as 8. For subspace, smoothness, and stochastic priors, the best possible recovery is characterized by a rank condition
9
with 0 determined by the prior and 1. In this formulation, vertex-wise flexibility means that only a limited number of vertices are active, some vertices may be mandatory, and others forbidden (Yamashita et al., 18 Sep 2025).
These formalisms already show that vertex-wise flexibility is not synonymous with unconstrained sampling. The literature repeatedly ties flexibility to rank, projection-bandwidth, coherence, or density conditions. A recurring misconception is that heterogeneous per-vertex sampling eliminates global minimality constraints; the cited results instead show that vertex-specific freedom is admissible only inside sharply defined algebraic or probabilistic bounds (Yu et al., 2019, Sheng et al., 29 Aug 2025).
2. Critical sampling of time-vertex graph signals
For noiseless bandlimited time-vertex signals, Yu et al. distinguish three bandwidth notions: general bandlimitedness (GBL), projection bandwidths 2, and simultaneous bandlimitedness (SBL). A critical sampling set 3 for a GBL signal must satisfy
4
together with 5. The necessary conditions are equally direct: 6 The paper also proves existence of a critical set for every GBL signal by first sampling separately in the time and graph domains, forming 7, and then selecting 8 independent rows by Gaussian elimination. The associated construction procedure has complexity 9, compared with 0 for naïve elimination on the full joint basis (Yu et al., 2019).
The most distinctive point for vertex-wise flexibility is that criticality constrains only the projection 1, not the number of time samples assigned to each vertex. The total sample budget may be distributed heterogeneously across vertices as long as
2
The paper explicitly notes that one may “cheaply sample one node at high rate and others more sparsely,” provided the aggregate budgets and rank condition are respected. In the stated reconstruction formula, once 3 is known to satisfy the rank condition, perfect recovery is
4
The same analysis describes trade-offs among total resources, conditioning of 5, and computation (Yu et al., 2019).
A later sampling theory for jointly bandlimited time-vertex graph signals extends the critical-sampling perspective to continuous-time, infinite-length discrete-time, and finite-length discrete-time models. For a jointly bandlimited signal with joint bandwidth 6 on a graph of 7 vertices, any stable sampling set must satisfy the global lower bound
8
and analogous lower bounds hold for discrete sampling ratios. The theory also provides per-vertex and vertex-subset density bounds through ranks of restricted graph Fourier submatrices, and constructs critical sets by decomposing the joint spectrum into subbands. In subband 9, only 0 vertices need be sampled; the union 1 achieves total density 2 with 3. The paper’s examples make the operational meaning explicit: in a 4 synthetic FTVGS, total ratio 5 achieves perfect recovery and reallocating 6 changes per-vertex rates while preserving 7; on EEG data the multi-band scheme yields 8 versus separate 9 with 0, and on METR-LA traffic it yields critical 1 versus separate 2 with 3 (Sheng et al., 29 Aug 2025).
3. Generalized graph signals and pre-selected vertices
The generalized-sampling framework of Yamashita et al. broadens vertex-wise flexibility beyond pure vertex selection. The signal prior may be subspace-based, smoothness-based, or stochastic. The design variable is the sampling operator 4, and the ideal but nonconvex feasibility problem requires simultaneously that forbidden vertices be inactive, the number of active undecided vertices be bounded, and 5. The vertex set is partitioned into three disjoint subsets: 6 If 7 is the maximum number of active vertices, 8, and 9, then the constraints are
0
with the full-rank condition on 1 (Yamashita et al., 18 Sep 2025).
To make the design tractable, the paper relaxes rank maximization by the nuclear norm and replaces the hard cardinality indicator on undecided vertices by a difference-of-convex penalty. With
2
the final program is
3
It is solved by the General Double-Proximal Gradient for DC programming (GDPGDC). The primal proximal step is explicit: forbidden rows are set to zero, undecided rows undergo group-soft thresholding, and mandatory rows undergo only 4-regularization. The dual proximal step combines singular-value shrinkage for the nuclear norm block with partial-sum group thresholding for the 5 block. Under the cited conditions of Banert–Bȍţ (2019), bounded iterates are guaranteed and cluster points are critical points of the DC program (Yamashita et al., 18 Sep 2025).
This line of work is important because it explicitly interpolates between “vertex-wise sampling,” where samples are raw vertex values, and “fully flexible sampling,” where samples may be arbitrary linear combinations. It also incorporates prior knowledge unavailable to earlier vertex-wise flexible samplers: mandatory inclusion or exclusion of specific vertices. In experiments on 256-node 6-nearest-neighbor sensor graphs with 7, 8, and 9, the method is reported as uniformly superior to SP and AVM and as matching or outperforming GSSS and ScFGSS in noiseless and noisy cases; when mandatory and forbidden vertices are chosen well, it “often achieve[s] a 5–10 dB gain over the next-best method.” On monthly-average Swiss temperatures at 0 locations with 1, 2, and 3, it reduces MSE by 3–7 dB versus ScFGSS and GSSS. The paper also states its limitations plainly: the nuclear norm is only a surrogate for rank, and the active-vertex constraint is enforced through a DC penalty, so the number of active vertices may slightly exceed 4 in practice (Yamashita et al., 18 Sep 2025).
4. Unknown spectral support and transform-domain extensions
When spectral support is unknown, vertex-wise flexibility becomes a question of where sampling is permitted rather than which known spectral coordinates need to be preserved. For finite time-vertex graph signals 5, Sheng et al. propose subset random sampling: first select a subset 6 of rows and a subset 7 of columns, then sample entries only inside the submatrix 8. The model assumes low rank, smoothness, and coherence: 9 together with bounded graph and temporal gradients and coherence bounds on the thin SVD factors. If
0
then rank is preserved in the selected row and row-column submatrices with high probability. A further sample-complexity bound on 1 inside 2 gives exact reconstruction of the original 3 with high probability. The paper emphasizes the design trade-off: rows and time instants outside 4 or 5 are never sampled, so fewer sensors and fewer time windows are needed, but a larger sample budget inside the chosen submatrix is required to satisfy the rank, incoherence, and RIP-type conditions (Sheng et al., 2024).
The corresponding reconstruction strategy has two stages. First, recover 6 by nuclear-norm minimization under the observed-entry constraint. Second, extend to the full 7 by exploiting smoothness, for example through total-variation inpainting with graph and temporal gradients. The same paper notes that the experiment instead uses a joint estimator combining a low-rank surrogate 8, graph- and time-domain sparsity penalties, and an error term (Sheng et al., 2024).
A 2025 extension develops this latter idea into an explicit low-rank, sparsity, and smoothness prior (LSSP) framework. The sampling model selects row and column fractions 9, then samples a fraction 0 inside the resulting submatrix. Reconstruction minimizes the sum of the nonconvex low-rank surrogate, graph- and time-spectral 1 penalties, and temporal smoothness 2, subject to graph/time spectral consistency and exact fitting on the observed set. The paper solves the problem by ADMM with reweighting. On synthetic data with 3, LSSP yields the lowest NRMSE among CCS-ICURC, nonconvex MC, ReLaSP, LIMC, and LRDS; at 4 it reports NRMSE 5, increasing to approximately 6 at 7. On METR-LA traffic with 8 and 9, it again reports the best average NRMSE, including 00 at 01, and identifies a “knee” near 02, corresponding to total sampling of approximately 03 (Sheng et al., 29 Aug 2025).
A different extension changes the transform itself. JFRFT-based sampling replaces the classical joint Fourier transform by the joint time-vertex fractional Fourier transform,
04
or, in vectorized form, 05 with 06. For 07-bandlimited signals with support 08, perfect recovery from a sampling support 09 is characterized by the restricted transform matrix and the pseudoinverse recovery operator
10
Sampling-set design is posed through criteria such as MaxSigMin, A-optimal trace minimization, MinPinv, MaxSig, and MaxVol, and a localized operator 11 enables lower-cost design based on much smaller dense matrices. The paper states that these criteria yield a sampling pattern that can differ from vertex to vertex and time to time according to the local spectral content measured by the fractional transform. In experiments on seasonal U.S. states, all optimized strategies outperform random sampling; on “dog-walking” meshes, MinPinv gives the best runtime/accuracy trade-off; and on spaceborne sea-clutter, sweeping 12 finds 13, reducing NMSE from approximately 14 under JFT to approximately 15 under JFRFT with MinPinv (Zhang et al., 22 May 2025).
5. Learning-oriented sampling in network embeddings and GNNs
In representation learning, vertex-wise flexible sampling generally refers to non-uniform stochastic generation of training pairs or neighborhoods. The “Vertex-Context Sampling” framework for weighted network embedding targets source-context pairs 16 with conditional distribution
17
rather than uniform choice over outgoing neighbors. The same weighted logic extends to higher-order walks by multiplying transition probabilities along the walk. To support this efficiently, the framework uses two levels of Walker alias tables: a global source-vertex table, per-vertex context tables stored back-to-back, and a global negative-sampling table. Preprocessing takes 18, each of VertexSampling, ContextSampling, and NegativeSampling runs in 19 time, and total memory is 20. The framework integrates with DeepWalk, LINE, Walklets, and HPE, and the paper reports improved performance on text9, MovieLens-latest, and KKBOX. On KKBOX, for example, DeepWalk attains 21, uniform VCS-HPE 22, and non-uniform VCS-HPE 23 (Chen et al., 2017).
The GraphSAGE line of work addresses a related issue at the neighborhood-aggregation level. Standard GraphSAGE samples neighbors uniformly; the cited paper argues that this induces high variance in training and inference. Its alternative is a learned importance function for each pair 24, 25, using a one-layer perceptron on concatenated node attributes,
26
The target is an empirical return obtained from the negative classification loss accumulated across GraphSAGE hops, and the regressor is trained by an 27 value-function loss. The resulting scores induce a vertex-specific sampling distribution
28
which is approximated in practice by block-wise argmax to preserve parallelism. On PPI, Reddit, and PubMed, this vertex-wise flexible sampling improves over uniform GraphSAGE: on PPI, 2-layer GraphSAGE with sample size 29 goes from Micro-F1 30 to 31; on PPI 3-layer, from 32 to 33; on Reddit 2-layer, from 34 to 35; and on PubMed 3-layer, from 36 to 37. The reported overhead is about 38 to 39 the wall-clock time of uniform GraphSAGE (Oh et al., 2019).
These learning-oriented formulations differ fundamentally from the reconstruction-oriented literature. They do not seek full-rank recovery operators or critical densities. Instead, they alter the stochastic process that generates training data so that edge weights, neighborhood importance, or return estimates influence where the algorithm spends its sampling budget. The shared feature is still heterogeneity at the vertex level: 40 depends on the source vertex in VCS, and 41 depends on the target-neighbor pair in the GraphSAGE sampler (Chen et al., 2017, Oh et al., 2019).
6. Vertex weighting, active annotation, and spatial designs
Another strand of the literature makes vertex-wise flexibility explicit through application-dependent weighting of the graph-signal space. In the Hilbert-space formulation of graph vertex sampling, the signal space 42 is equipped with an inner product
43
where 44 is symmetric positive-definite. Choosing 45 changes both the graph Fourier basis, via the generalized eigenproblem 46, and the sampling criterion. For a candidate set 47, the 48-th order spectral-proxy cutoff is
49
The practical design is a greedy maximization of 50. The framework is explicitly motivated by vertex importance: 51 recovers the classical case, 52 emphasizes high-degree nodes, and on geometric graphs 53 may be the diagonal matrix of Voronoi-cell areas. In the reported experiment on 54 geometric graphs with 55, the Voronoi-area inner product consistently outperforms 56 and 57 in both the smallest singular value bound and the reconstruction error over 58 trials (Girault et al., 2020).
Active few-shot vertex classification introduces a query-based variant of vertex-wise flexibility. The setting begins with an unlabeled graph 59, a total annotation budget 60, and iterative selection of vertices to be labeled by a human annotator. The paper studies three scenarios: “Balanced Sampling,” which assumes a class oracle and queries one vertex per true class per round; “Unbalanced Sampling,” which drops oracle labels and partitions current embeddings by 61-medoids into 62 groups; and “Unknown Number of Classes,” which first estimates 63 by an elbow method on Deep Graph Infomax embeddings and then uses 64-medoids. Within groups, four strategies are evaluated: Random, Entropy, PageRank, and Medoid sampling. The paper reports that prototypical models outperform discriminative models when fewer than 65 samples per class are available, that removing the class oracle reduces GCN performance by 66 but the prototypical network by only 67 on average, and that moving to the unknown-number-of-classes setting causes a further 68 average decrease for both models. It also reports early-round gains of 69–70 percentage points from label propagation and identifies medoid sampling as the strongest active-learning heuristic (Burr et al., 25 Apr 2025).
Graph spatial sampling provides yet another interpretation. In the Lagged Metropolis–Hastings Walk framework, one specifies target stationary probabilities 71 on vertices of a simple undirected graph and chooses a jump rate 72, a backtrack weight 73, and a preference vector 74 satisfying
75
The proposal depends on both the current and previous vertex, and the Metropolis–Hastings acceptance ratio enforces the target stationary law. By designing the graph 76, one can forbid edges between spatially contiguous units and thereby enforce separation in the sampled set. The paper states that the resulting graph spatial sampling approach can be “more flexible for improving the design efficiency compared to the existing spatial sampling methods,” and reports relative efficiencies several times higher than LPM or GRTS when the graph is designed to match the anticipated spatial trend (Zhang, 2022).
7. Constraints, limitations, and open directions
The literature consistently shows that flexibility is conditional rather than absolute. In critical time-vertex sampling, any qualified set must satisfy 77, 78, and 79, and later jointly bandlimited theory sharpens this into density or ratio bounds such as 80. The same papers provide per-vertex lower bounds, so redistributing samples across vertices is allowed only if the spectral and rank conditions remain intact (Yu et al., 2019, Sheng et al., 29 Aug 2025).
Methods for unknown spectral support introduce different constraints. Subset random sampling relies on low rank, coherence, and smoothness, and compensates for never sampling outside 81 by requiring a larger sample budget inside the chosen submatrix. Its guarantees are high-probability statements rather than deterministic exactness, and the reconstruction pipeline typically combines matrix completion with smoothness-based inpainting or with a joint low-rank/sparsity/smoothness program (Sheng et al., 2024, Sheng et al., 29 Aug 2025).
Optimization-based generalized sampling offers a larger design space but inherits the usual trade-offs of nonconvex relaxations. The nuclear norm is only a surrogate for rank, and the active-vertex cardinality is approximated by a DC penalty rather than enforced exactly, so the number of active vertices may slightly exceed 82. JFRFT-based designs introduce additional freedom through fractional orders 83, but their theory still assumes bandlimitedness and relies on idealized spectral constructions; the multi-band time-vertex theory likewise notes open questions on approximate filters, asynchronous sampling, noise robustness, and model mismatch (Yamashita et al., 18 Sep 2025, Zhang et al., 22 May 2025, Sheng et al., 29 Aug 2025).
Learning-based samplers are constrained less by algebraic recoverability than by model quality. Weighted vertex-context sampling presumes informative edge weights and alias-table preprocessing; RL-based GraphSAGE sampling depends on the learned regressor and current embeddings; active few-shot sampling depends on clustering quality, the quality of label propagation, and the availability of useful low-label inductive biases. The cited results also show that not all models respond equally to loss of oracle information: in the few-shot setting, the GCN is more sensitive than the prototypical model when the class oracle is removed (Chen et al., 2017, Oh et al., 2019, Burr et al., 25 Apr 2025).
Taken together, these works indicate that “vertex-wise flexible sampling” is best understood as an umbrella term for heterogeneous vertex-level allocation mechanisms rather than as a single theorem or algorithmic template. A plausible implication is that future unification will require combining the hard guarantees of sampling theory with the adaptive, data-dependent allocation rules used in contemporary graph learning.