---
title: Vertex-Wise Flexible Graph Sampling
url: https://www.emergentmind.com/topics/vertex-wise-flexible-sampling
type: topic
---

# Vertex-Wise Flexible Graph Sampling

Vertex-wise flexible sampling denotes a family of graph-based sampling schemes in which sampling effort is allowed to vary across vertices rather than being imposed uniformly. In the cited literature, the term appears in several technically distinct settings: time-vertex graph signal processing, generalized graph-signal reconstruction, weighted network embedding, graph neural network training, active few-shot annotation, and design-based graph spatial sampling. Across these settings, the common principle is heterogeneous allocation of samples, measurements, or query probabilities over vertices so as to satisfy recovery conditions, reduce reconstruction error, or improve downstream learning efficiency under resource constraints [1909.02692], [2509.14836], [1711.00227], [1904.12935], [2504.18696].

## 1. Scope and formal characterizations

A canonical formalization arises in joint time-vertex signal processing. Let \(G_G=(V_G,E_G,W_G)\) be an undirected graph with Laplacian \(L_G=U_G\Lambda_GU_G^H\), and let \(G_T\) be the cycle graph on \(T\) time nodes with Laplacian \(L_T=U_T\Lambda_TU_T^H\). The joint graph is the Cartesian product \(G_J=G_T\times G_G\) with Laplacian
\[
L_J=L_T\otimes I_{N_G}+I_T\otimes L_G,
\]
and a joint time-vertex signal \(X\in\mathbb R^{N_G\times T}\) is vectorized as \(x=\mathrm{vec}(X)\). Sampling is represented by a binary operator \(\Psi\) selecting a subset \(S\subseteq V_T\times V_G\), yielding \(x_S=\Psi x\). In the bandlimited setting, exact recovery is possible exactly when the reduced Fourier basis \(\tilde U_J\) associated with the active spectral coefficients satisfies
\[
\mathrm{rank}(\Psi \tilde U_J)=K,
\]
where \(K\) is the general bandwidth [1909.02692].

A second formalization appears in generalized graph sampling. There, a graph signal \(x\in\mathbb R^N\) is sampled through a designed operator \(S^T\in\mathbb R^{M\times N}\), producing \(y=S^T x\), and reconstructed as \(\tilde x=RCy\). For subspace, smoothness, and stochastic priors, the best possible recovery is characterized by a rank condition
\[
\mathrm{rank}(\Phi S^T)=R,
\]
with \(\Phi\) determined by the prior and \(R=\mathrm{rank}\,\Phi\). In this formulation, vertex-wise flexibility means that only a limited number of vertices are active, some vertices may be mandatory, and others forbidden [2509.14836].

These formalisms already show that vertex-wise flexibility is not synonymous with unconstrained sampling. The literature repeatedly ties flexibility to rank, projection-bandwidth, coherence, or density conditions. A recurring misconception is that heterogeneous per-vertex sampling eliminates global minimality constraints; the cited results instead show that vertex-specific freedom is admissible only inside sharply defined algebraic or probabilistic bounds [1909.02692], [2508.21412].

## 2. Critical sampling of time-vertex graph signals

For noiseless bandlimited time-vertex signals, Yu et al. distinguish three bandwidth notions: general bandlimitedness (GBL), projection bandwidths \((K_G,K_T)\), and simultaneous bandlimitedness (SBL). A critical sampling set \(S\subseteq V_T\times V_G\) for a GBL signal must satisfy
\[
|S|=K,\qquad |S_G|=K_G,\qquad |S_T|=K_T,
\]
together with \(\mathrm{rank}(\Psi\tilde U_J)=K\). The necessary conditions are equally direct:
\[
|S|\ge K,\qquad |S_G|\ge K_G,\qquad |S_T|\ge K_T.
\]
The paper also proves existence of a critical set for every GBL signal by first sampling separately in the time and graph domains, forming \(S'=S_T\times S_G\), and then selecting \(K\) independent rows by Gaussian elimination. The associated construction procedure has complexity \(O(N_G^3+T^3+(K_TK_G)^3)\), compared with \(O((TN_G)^3)\) for naïve elimination on the full joint basis [1909.02692].

The most distinctive point for vertex-wise flexibility is that criticality constrains only the projection \(S_G\), not the number of time samples assigned to each vertex. The total sample budget may be distributed heterogeneously across vertices as long as
\[
\sum_{v\in V_G} |S_T(v)| = K,\qquad |\{v:|S_T(v)|>0\}|=K_G.
\]
The paper explicitly notes that one may “cheaply sample one node at high rate and others more sparsely,” provided the aggregate budgets and rank condition are respected. In the stated reconstruction formula, once \(\Psi\) is known to satisfy the rank condition, perfect recovery is
\[
x=\Phi x_S,\qquad \Phi=\tilde U_J(\Psi\tilde U_J)^{-1}.
\]
The same analysis describes trade-offs among total resources, conditioning of \(\Psi\tilde U_J\), and computation [1909.02692].

A later sampling theory for jointly bandlimited time-vertex graph signals extends the critical-sampling perspective to continuous-time, infinite-length discrete-time, and finite-length discrete-time models. For a jointly bandlimited signal with joint bandwidth \(B\) on a graph of \(N\) vertices, any stable sampling set must satisfy the global lower bound
\[
D(S)\ge B/N,
\]
and analogous lower bounds hold for discrete sampling ratios. The theory also provides per-vertex and vertex-subset density bounds through ranks of restricted graph Fourier submatrices, and constructs critical sets by decomposing the joint spectrum into subbands. In subband \(k\), only \(m_k\) vertices need be sampled; the union \(S=\bigcup_k (V^k\times T_k)\) achieves total density \(B/N\) with \(|\bigcup_k V^k|\le B_G\). The paper’s examples make the operational meaning explicit: in a \(4\times4\) synthetic FTVGS, total ratio \(R=7/16\) achieves perfect recovery and reallocating \(V^k\) changes per-vertex rates while preserving \(R\); on EEG data the multi-band scheme yields \(R\approx0.06\) versus separate \(R\approx0.10\) with \(\mathrm{NRMSE}\approx10^{-6}\), and on METR-LA traffic it yields critical \(R\approx0.026\) versus separate \(R\approx0.123\) with \(\mathrm{NRMSE}\approx10^{-5}\) [2508.21412].

## 3. Generalized graph signals and pre-selected vertices

The generalized-sampling framework of Yamashita et al. broadens vertex-wise flexibility beyond pure vertex selection. The signal prior may be subspace-based, smoothness-based, or stochastic. The design variable is the sampling operator \(S^T\), and the ideal but nonconvex feasibility problem requires simultaneously that forbidden vertices be inactive, the number of active undecided vertices be bounded, and \(\mathrm{rank}(\Phi S^T)=R\). The vertex set is partitioned into three disjoint subsets:
\[
S_+ \text{ (mandatory)},\qquad S_- \text{ (forbidden)},\qquad U \text{ (undecided)}.
\]
If \(Z\) is the maximum number of active vertices, \(A=|S_+|\), and \(\Delta=Z-A\), then the constraints are
\[
s_i^T=0\ \forall i\in S_-,\qquad |\{i\in U:\|s_i\|_2>0\}|\le \Delta,
\]
with the full-rank condition on \(\Phi S^T\) [2509.14836].

To make the design tractable, the paper relaxes rank maximization by the nuclear norm and replaces the hard cardinality indicator on undecided vertices by a difference-of-convex penalty. With
\[
G_1(W)=\sum_i \|W_i\|_2,\qquad G_2(W)=\Omega_\Delta(W),
\]
the final program is
\[
\min_{S^T}\ \bigl[\iota_{C_-}(S^T)+\lambda G_1(S_U^T)+(\delta/2)\|S^T\|_F^2\bigr]
-\bigl[\|\Phi S^T\|_*+\lambda G_2(S_U^T)\bigr].
\]
It is solved by the General Double-Proximal Gradient for DC programming (GDPGDC). The primal proximal step is explicit: forbidden rows are set to zero, undecided rows undergo group-soft thresholding, and mandatory rows undergo only \(\ell_2\)-regularization. The dual proximal step combines singular-value shrinkage for the nuclear norm block with partial-sum group thresholding for the \(\Omega_\Delta\) block. Under the cited conditions of Banert–Bȍţ (2019), bounded iterates are guaranteed and cluster points are critical points of the DC program [2509.14836].

This line of work is important because it explicitly interpolates between “vertex-wise sampling,” where samples are raw vertex values, and “fully flexible sampling,” where samples may be arbitrary linear combinations. It also incorporates prior knowledge unavailable to earlier vertex-wise flexible samplers: mandatory inclusion or exclusion of specific vertices. In experiments on 256-node \(k\)-nearest-neighbor sensor graphs with \(M=32\), \(Z=32\), and \(|S_+|=|S_-|=16\), the method is reported as uniformly superior to SP and AVM and as matching or outperforming GSSS and ScFGSS in noiseless and noisy cases; when mandatory and forbidden vertices are chosen well, it “often achieve[s] a 5–10 dB gain over the next-best method.” On monthly-average Swiss temperatures at \(N=110\) locations with \(M=28\), \(Z=28\), and \(|S_+|=|S_-|=14\), it reduces MSE by 3–7 dB versus ScFGSS and GSSS. The paper also states its limitations plainly: the nuclear norm is only a surrogate for rank, and the active-vertex constraint is enforced through a DC penalty, so the number of active vertices may slightly exceed \(Z\) in practice [2509.14836].

## 4. Unknown spectral support and transform-domain extensions

When spectral support is unknown, vertex-wise flexibility becomes a question of where sampling is permitted rather than which known spectral coordinates need to be preserved. For finite time-vertex graph signals \(X\in\mathbb R^{N\times T}\), Sheng et al. propose subset random sampling: first select a subset \(I\subseteq V_G\) of rows and a subset \(J\subseteq V_T\) of columns, then sample entries only inside the submatrix \(X_{I,J}\). The model assumes low rank, smoothness, and coherence:
\[
\mathrm{rank}(X)=r<\min\{N,T\},
\]
together with bounded graph and temporal gradients and coherence bounds on the thin SVD factors. If
\[
|I| \ge (3r\mu_1/\epsilon^2)\ln(2r/\delta),\qquad
|J| \ge (3r\mu_2/\epsilon^2)\ln(2r/\delta),
\]
then rank is preserved in the selected row and row-column submatrices with high probability. A further sample-complexity bound on \(|S|\) inside \(I\times J\) gives exact reconstruction of the original \(X\) with high probability. The paper emphasizes the design trade-off: rows and time instants outside \(I\) or \(J\) are never sampled, so fewer sensors and fewer time windows are needed, but a larger sample budget inside the chosen submatrix is required to satisfy the rank, incoherence, and RIP-type conditions [2410.22731].

The corresponding reconstruction strategy has two stages. First, recover \(X_{I,J}\) by nuclear-norm minimization under the observed-entry constraint. Second, extend to the full \(X\) by exploiting smoothness, for example through total-variation inpainting with graph and temporal gradients. The same paper notes that the experiment instead uses a joint estimator combining a low-rank surrogate \(g(x)=\log(\sqrt{x}+1)\), graph- and time-domain sparsity penalties, and an error term [2410.22731].

A 2025 extension develops this latter idea into an explicit low-rank, sparsity, and smoothness prior (LSSP) framework. The sampling model selects row and column fractions \(\alpha_R,\alpha_C\), then samples a fraction \(\alpha_{\text{sub}}\) inside the resulting submatrix. Reconstruction minimizes the sum of the nonconvex low-rank surrogate, graph- and time-spectral \(\ell_1\) penalties, and temporal smoothness \(\|XD_2\|_F^2\), subject to graph/time spectral consistency and exact fitting on the observed set. The paper solves the problem by ADMM with reweighting. On synthetic data with \(N=T=128\), LSSP yields the lowest NRMSE among CCS-ICURC, nonconvex MC, ReLaSP, LIMC, and LRDS; at \(\alpha_{RC}=\alpha_{\text{sub}}=0.9\) it reports NRMSE \(0.112\), increasing to approximately \(0.233\) at \(0.6\). On METR-LA traffic with \(N=207\) and \(T=512\), it again reports the best average NRMSE, including \(0.125\) at \(\alpha_{RC}=\alpha_{\text{sub}}=0.9\), and identifies a “knee” near \(\alpha_{RC}=\alpha_{\text{sub}}\approx0.7\), corresponding to total sampling of approximately \(34\%\) [2508.21415].

A different extension changes the transform itself. JFRFT-based sampling replaces the classical joint Fourier transform by the joint time-vertex fractional Fourier transform,
\[
\mathrm{JFRFT}_{\alpha,\beta}\{X\}=F_{\mathcal G}^\beta X (F^\alpha)^H,
\]
or, in vectorized form, \(\hat x=\mathrm J^{\alpha,\beta}x\) with \(\mathrm J^{\alpha,\beta}=F_{\mathcal G}^\beta\otimes(F^\alpha)^H\). For \(\{\alpha,\beta\}\)-bandlimited signals with support \(\mathcal F_J\), perfect recovery from a sampling support \(\mathcal S_J\) is characterized by the restricted transform matrix and the pseudoinverse recovery operator
\[
R_J=\bigl(D_JB_J^{\alpha,\beta}\bigr)^\dagger
=\mathrm J^{-\alpha,-\beta}_{[:,\mathcal F_J]}
\bigl(\mathrm J^{-\alpha,-\beta}_{\mathcal S_J,\mathcal F_J}\bigr)^\dagger.
\]
Sampling-set design is posed through criteria such as MaxSigMin, A-optimal trace minimization, MinPinv, MaxSig, and MaxVol, and a localized operator \(T^{\alpha,\beta}\) enables lower-cost design based on much smaller dense matrices. The paper states that these criteria yield a sampling pattern that can differ from vertex to vertex and time to time according to the local spectral content measured by the fractional transform. In experiments on seasonal U.S. states, all optimized strategies outperform random sampling; on “dog-walking” meshes, MinPinv gives the best runtime/accuracy trade-off; and on spaceborne sea-clutter, sweeping \(\alpha,\beta\in[-4,4]\) finds \((\alpha\approx1.9,\beta\approx0.7)\), reducing NMSE from approximately \(0.12\) under JFT to approximately \(0.065\) under JFRFT with MinPinv [2506.00023].

## 5. Learning-oriented sampling in network embeddings and GNNs

In representation learning, vertex-wise flexible sampling generally refers to non-uniform stochastic generation of training pairs or neighborhoods. The “Vertex-Context Sampling” framework for weighted network embedding targets source-context pairs \((v_s,v_c)\) with conditional distribution
\[
P(v_c\mid v_s=i)=\frac{w_{i,c}}{W_i^+},\qquad
W_i^+=\sum_{u\in \mathcal N^+(i)} w_{i,u},
\]
rather than uniform choice over outgoing neighbors. The same weighted logic extends to higher-order walks by multiplying transition probabilities along the walk. To support this efficiently, the framework uses two levels of Walker alias tables: a global source-vertex table, per-vertex context tables stored back-to-back, and a global negative-sampling table. Preprocessing takes \(\mathcal O(|V|+|E|)\), each of VertexSampling, ContextSampling, and NegativeSampling runs in \(\mathcal O(1)\) time, and total memory is \(\Theta(|V|+|E|)\). The framework integrates with DeepWalk, LINE, Walklets, and HPE, and the paper reports improved performance on text9, MovieLens-latest, and KKBOX. On KKBOX, for example, DeepWalk attains \(\mathrm{mAP}\approx6.8\%\), uniform VCS-HPE \(\approx8.2\%\), and non-uniform VCS-HPE \(\approx14.5\%\) [1711.00227].

The GraphSAGE line of work addresses a related issue at the neighborhood-aggregation level. Standard GraphSAGE samples neighbors uniformly; the cited paper argues that this induces high variance in training and inference. Its alternative is a learned importance function for each pair \((v,u)\), \(u\in\mathcal N(v)\), using a one-layer perceptron on concatenated node attributes,
\[
\hat{\mathcal V}_{v,u}
=
-\exp\!\Bigl(\sigma\bigl(W[x_v\parallel x_u]+b\bigr)\Bigr).
\]
The target is an empirical return obtained from the negative classification loss accumulated across GraphSAGE hops, and the regressor is trained by an \(\ell_2\) value-function loss. The resulting scores induce a vertex-specific sampling distribution
\[
p(u\mid v)=
\frac{\exp(\mathcal G_\theta(v,u))}
{\sum_{u'\in\mathcal N(v)}\exp(\mathcal G_\theta(v,u'))},
\]
which is approximated in practice by block-wise argmax to preserve parallelism. On PPI, Reddit, and PubMed, this vertex-wise flexible sampling improves over uniform GraphSAGE: on PPI, 2-layer GraphSAGE with sample size \(30\) goes from Micro-F1 \(0.674\) to \(0.755\); on PPI 3-layer, from \(0.780\) to \(0.846\); on Reddit 2-layer, from \(0.950\) to \(0.954\); and on PubMed 3-layer, from \(0.877\) to \(0.888\). The reported overhead is about \(1.5\times\) to \(2\times\) the wall-clock time of uniform GraphSAGE [1904.12935].

These learning-oriented formulations differ fundamentally from the reconstruction-oriented literature. They do not seek full-rank recovery operators or critical densities. Instead, they alter the stochastic process that generates training data so that edge weights, neighborhood importance, or return estimates influence where the algorithm spends its sampling budget. The shared feature is still heterogeneity at the vertex level: \(P(v_c\mid v_s)\) depends on the source vertex in VCS, and \(p(u\mid v)\) depends on the target-neighbor pair in the GraphSAGE sampler [1711.00227], [1904.12935].

## 6. Vertex weighting, active annotation, and spatial designs

Another strand of the literature makes vertex-wise flexibility explicit through application-dependent weighting of the graph-signal space. In the Hilbert-space formulation of graph vertex sampling, the signal space \(\mathbb R^N\) is equipped with an inner product
\[
\langle x,y\rangle_M=x^TMy,
\]
where \(M\) is symmetric positive-definite. Choosing \(M\) changes both the graph Fourier basis, via the generalized eigenproblem \(\Delta u=\lambda M u\), and the sampling criterion. For a candidate set \(S\), the \(k\)-th order spectral-proxy cutoff is
\[
\Omega_k(S)=\min_{x\neq 0,\ x_S=0}\omega_k(x),
\qquad
\omega_k(x)=\left(\frac{\|\Delta^k x\|_M}{\|x\|_M}\right)^{1/k}.
\]
The practical design is a greedy maximization of \(\Omega_k(S)\). The framework is explicitly motivated by vertex importance: \(M=I\) recovers the classical case, \(M=\mathrm{diag}(\deg)\) emphasizes high-degree nodes, and on geometric graphs \(M\) may be the diagonal matrix of Voronoi-cell areas. In the reported experiment on \(N=100\) geometric graphs with \(k=3\), the Voronoi-area inner product consistently outperforms \(M=I\) and \(M=\mathrm{diag}(\deg)\) in both the smallest singular value bound and the reconstruction error over \(5{,}000\) trials [2002.11238].

Active few-shot vertex classification introduces a query-based variant of vertex-wise flexibility. The setting begins with an unlabeled graph \(G=(V,E)\), a total annotation budget \(B\), and iterative selection of vertices to be labeled by a human annotator. The paper studies three scenarios: “Balanced Sampling,” which assumes a class oracle and queries one vertex per true class per round; “Unbalanced Sampling,” which drops oracle labels and partitions current embeddings by \(k\)-medoids into \(K=|C|\) groups; and “Unknown Number of Classes,” which first estimates \(K\) by an elbow method on Deep Graph Infomax embeddings and then uses \(k\)-medoids. Within groups, four strategies are evaluated: Random, Entropy, PageRank, and Medoid sampling. The paper reports that prototypical models outperform discriminative models when fewer than \(20\) samples per class are available, that removing the class oracle reduces GCN performance by \(9\%\) but the prototypical network by only \(1\%\) on average, and that moving to the unknown-number-of-classes setting causes a further \(1\%\) average decrease for both models. It also reports early-round gains of \(+3\)–\(5\) percentage points from label propagation and identifies medoid sampling as the strongest active-learning heuristic [2504.18696].

Graph spatial sampling provides yet another interpretation. In the Lagged Metropolis–Hastings Walk framework, one specifies target stationary probabilities \(\pi_i\) on vertices of a simple undirected graph and chooses a jump rate \(r\), a backtrack weight \(w\), and a preference vector \(u_i\) satisfying
\[
u_i=\frac{\pi_i/(d_i+r)}{\sum_k \pi_k/(d_k+r)}.
\]
The proposal depends on both the current and previous vertex, and the Metropolis–Hastings acceptance ratio enforces the target stationary law. By designing the graph \(G\), one can forbid edges between spatially contiguous units and thereby enforce separation in the sampled set. The paper states that the resulting graph spatial sampling approach can be “more flexible for improving the design efficiency compared to the existing spatial sampling methods,” and reports relative efficiencies several times higher than LPM or GRTS when the graph is designed to match the anticipated spatial trend [2203.11734].

## 7. Constraints, limitations, and open directions

The literature consistently shows that flexibility is conditional rather than absolute. In critical time-vertex sampling, any qualified set must satisfy \(|S|\ge K\), \(|S_G|\ge K_G\), and \(|S_T|\ge K_T\), and later jointly bandlimited theory sharpens this into density or ratio bounds such as \(D(S)\ge B/N\). The same papers provide per-vertex lower bounds, so redistributing samples across vertices is allowed only if the spectral and rank conditions remain intact [1909.02692], [2508.21412].

Methods for unknown spectral support introduce different constraints. Subset random sampling relies on low rank, coherence, and smoothness, and compensates for never sampling outside \(I\times J\) by requiring a larger sample budget inside the chosen submatrix. Its guarantees are high-probability statements rather than deterministic exactness, and the reconstruction pipeline typically combines matrix completion with smoothness-based inpainting or with a joint low-rank/sparsity/smoothness program [2410.22731], [2508.21415].

Optimization-based generalized sampling offers a larger design space but inherits the usual trade-offs of nonconvex relaxations. The nuclear norm is only a surrogate for rank, and the active-vertex cardinality is approximated by a DC penalty rather than enforced exactly, so the number of active vertices may slightly exceed \(Z\). JFRFT-based designs introduce additional freedom through fractional orders \((\alpha,\beta)\), but their theory still assumes bandlimitedness and relies on idealized spectral constructions; the multi-band time-vertex theory likewise notes open questions on approximate filters, asynchronous sampling, noise robustness, and model mismatch [2509.14836], [2506.00023], [2508.21412].

Learning-based samplers are constrained less by algebraic recoverability than by model quality. Weighted vertex-context sampling presumes informative edge weights and alias-table preprocessing; RL-based GraphSAGE sampling depends on the learned regressor and current embeddings; active few-shot sampling depends on clustering quality, the quality of label propagation, and the availability of useful low-label inductive biases. The cited results also show that not all models respond equally to loss of oracle information: in the few-shot setting, the GCN is more sensitive than the prototypical model when the class oracle is removed [1711.00227], [1904.12935], [2504.18696].

Taken together, these works indicate that “vertex-wise flexible sampling” is best understood as an umbrella term for heterogeneous vertex-level allocation mechanisms rather than as a single theorem or algorithmic template. A plausible implication is that future unification will require combining the hard guarantees of sampling theory with the adaptive, data-dependent allocation rules used in contemporary graph learning.

Source: https://www.emergentmind.com/topics/vertex-wise-flexible-sampling