Papers
Topics
Authors
Recent
Search
2000 character limit reached

PolyGraph: Graph Generative Model Evaluation

Updated 14 July 2026
  • PolyGraph is a modular evaluation framework that transforms graph evaluation by using classifier-based, variational estimates of the Jensen–Shannon divergence to compare reference and generated graphs.
  • It employs a three-stage pipeline—descriptor extraction, classifier fitting, and metric computation—to produce the PolyGraph Discrepancy (PGD), a bounded quality metric ranging from 0 to 1.
  • The framework overcomes limitations of traditional MMD metrics by providing absolute, robust, and descriptor-comparable scores with improved performance at moderate sample sizes.

PolyGraph is a modular evaluation pipeline for graph generative models that replaces ad hoc, kernel-based metrics with classifier-based, variational estimates of the Jensen–Shannon distance between a reference graph distribution PP and a generated distribution QQ. Its central metric, PolyGraph Discrepancy (PGD), is designed to provide an absolute, bounded, and descriptor-comparable assessment of generative quality, in contrast to descriptor-dependent Maximum Mean Discrepancy (MMD) scores. The framework is publicly available as a benchmarking library for graph generative models and is organized around descriptor extraction, probabilistic discrimination, and held-out metric computation (Krimmel et al., 7 Oct 2025).

1. Problem setting and scope

PolyGraph is motivated by a limitation of prevailing evaluation practice in graph generation: existing methods rely primarily on MMD metrics based on graph descriptors. In that formulation, descriptor choice and kernel parametrization strongly influence numerical values, so scores can rank models within a fixed setup but do not define an absolute scale and are not comparable across descriptor families. PolyGraph addresses this by evaluating how well a probabilistic classifier can distinguish reference graphs from generated graphs after descriptor featurization, and by interpreting the resulting held-out log-likelihood as a variational lower bound on Jensen–Shannon divergence (Krimmel et al., 7 Oct 2025).

The framework operates on two multisets of graphs,

Pref={Gref(i)}i=1N,Qgen={Ggen(j)}j=1M,P_\mathrm{ref} = \{G^{(i)}_\mathrm{ref}\}_{i=1}^N, \qquad Q_\mathrm{gen} = \{G^{(j)}_\mathrm{gen}\}_{j=1}^M,

and a family of graph descriptors {dk}k=1K\{d_k\}_{k=1}^K. Each descriptor induces a marginal comparison problem in feature space rather than on raw combinatorial objects. This design makes evaluation modular: one may inspect descriptor-specific subscores while also deriving a single summary metric. A common misconception is that PolyGraph eliminates descriptor dependence entirely; the framework instead standardizes the scale and interpretation of descriptor-wise scores and then selects the tightest lower bound among the available descriptors.

2. Pipeline structure

The PolyGraph pipeline consists of three stages: descriptor extraction, classifier fitting, and metric computation. Descriptors convert graphs into fixed-dimensional vectors,

dk:G→Rnk,xi=dk(G(i)),d_k:\mathcal{G}\to\mathbb{R}^{n_k},\qquad x_i = d_k\bigl(G^{(i)}\bigr),

after which a probabilistic binary classifier Dk:Rnk→[0,1]D_k:\mathbb{R}^{n_k}\to[0,1] is trained for each descriptor, with label y=1y=1 for reference and y=0y=0 for generated graphs. The classifier is fit on disjoint “fit” subsets and evaluated on held-out “test” subsets of equal size (Krimmel et al., 7 Oct 2025).

The descriptor set described for the framework includes both classical statistics and learned embeddings.

Descriptor family Examples
Histogram/statistical descriptors Degree histograms, clustering coefficient histograms, Laplacian spectrum histograms
Motif-based descriptors Graphlet/orbit counts: Orb4\mathrm{Orb}_4, Orb5\mathrm{Orb}_5
Learned representations GIN embeddings

For classifier fitting, the framework uses TabPFN, described as a hyperparameter-free transformer approximate Bayesian classifier. The intended role of this choice is methodological rather than incidental: calibrated probabilistic outputs are required because the downstream quantity is not classification accuracy but data log-likelihood on held-out samples. This suggests that PolyGraph is less a single scalar metric than a benchmarking protocol in which classifier quality directly controls the tightness of the variational lower bound (Krimmel et al., 7 Oct 2025).

3. Variational formulation of PGD

The mathematical basis of PolyGraph is the Jensen–Shannon divergence

QQ0

and the associated Jensen–Shannon distance

QQ1

The key variational identity used in the framework is

QQ2

Accordingly, for any classifier QQ3,

QQ4

The held-out average log-likelihood for descriptor QQ5 is

QQ6

In the framework description, QQ7 is treated as a lower bound on the JS divergence between descriptor-marginal distributions, and each descriptor yields a bounded metric

QQ8

Descriptor-wise evaluation is then aggregated by a max-reduction: QQ9 The summary metric is explicitly presented as the maximally tight lower bound available from the candidate descriptor family. This does not imply that the maximum reconstructs the full graph-distribution distance; rather, it yields the strongest lower bound accessible through the chosen feature maps.

4. Relation to MMD-based evaluation

For comparison, the framework recalls the MMD functional

Pref={Gref(i)}i=1N,Qgen={Ggen(j)}j=1M,P_\mathrm{ref} = \{G^{(i)}_\mathrm{ref}\}_{i=1}^N, \qquad Q_\mathrm{gen} = \{G^{(j)}_\mathrm{gen}\}_{j=1}^M,0

or, in RKHS variational form,

Pref={Gref(i)}i=1N,Qgen={Ggen(j)}j=1M,P_\mathrm{ref} = \{G^{(i)}_\mathrm{ref}\}_{i=1}^N, \qquad Q_\mathrm{gen} = \{G^{(j)}_\mathrm{gen}\}_{j=1}^M,1

The reported drawbacks for graph generative model evaluation are threefold: absence of intrinsic scale, incomparability across descriptor and kernel choices, and high bias and variance at typical sample sizes. PolyGraph is positioned against these issues by producing scores in Pref={Gref(i)}i=1N,Qgen={Ggen(j)}j=1M,P_\mathrm{ref} = \{G^{(i)}_\mathrm{ref}\}_{i=1}^N, \qquad Q_\mathrm{gen} = \{G^{(j)}_\mathrm{gen}\}_{j=1}^M,2, making descriptor-wise values directly comparable, and exhibiting greater robustness at moderate sample sizes of Pref={Gref(i)}i=1N,Qgen={Ggen(j)}j=1M,P_\mathrm{ref} = \{G^{(i)}_\mathrm{ref}\}_{i=1}^N, \qquad Q_\mathrm{gen} = \{G^{(j)}_\mathrm{gen}\}_{j=1}^M,3–Pref={Gref(i)}i=1N,Qgen={Ggen(j)}j=1M,P_\mathrm{ref} = \{G^{(i)}_\mathrm{ref}\}_{i=1}^N, \qquad Q_\mathrm{gen} = \{G^{(j)}_\mathrm{gen}\}_{j=1}^M,4 graph pairs (Krimmel et al., 7 Oct 2025).

Two methodological cautions are built into the framework. First, no single descriptor suffices under all perturbations; the descriptor set should therefore be interpreted as a family of complementary probes rather than interchangeable surrogates. Second, descriptor selection must avoid test-set leakage. The framework addresses this by using cross-validated fit-split performance for descriptor selection and reserving the held-out test split exclusively for final reporting. A plausible implication is that PolyGraph treats evaluation itself as a statistical estimation problem, not merely as score computation.

5. Summary metric, cross-validation, and implementation

The framework’s summary metric is not computed by averaging descriptor scores. Instead, for each descriptor Pref={Gref(i)}i=1N,Qgen={Ggen(j)}j=1M,P_\mathrm{ref} = \{G^{(i)}_\mathrm{ref}\}_{i=1}^N, \qquad Q_\mathrm{gen} = \{G^{(j)}_\mathrm{gen}\}_{j=1}^M,5, one performs Pref={Gref(i)}i=1N,Qgen={Ggen(j)}j=1M,P_\mathrm{ref} = \{G^{(i)}_\mathrm{ref}\}_{i=1}^N, \qquad Q_\mathrm{gen} = \{G^{(j)}_\mathrm{gen}\}_{j=1}^M,6-fold stratified cross-validation on the fit split to estimate Pref={Gref(i)}i=1N,Qgen={Ggen(j)}j=1M,P_\mathrm{ref} = \{G^{(i)}_\mathrm{ref}\}_{i=1}^N, \qquad Q_\mathrm{gen} = \{G^{(j)}_\mathrm{gen}\}_{j=1}^M,7, selects

Pref={Gref(i)}i=1N,Qgen={Ggen(j)}j=1M,P_\mathrm{ref} = \{G^{(i)}_\mathrm{ref}\}_{i=1}^N, \qquad Q_\mathrm{gen} = \{G^{(j)}_\mathrm{gen}\}_{j=1}^M,8

then retrains on all fit data with Pref={Gref(i)}i=1N,Qgen={Ggen(j)}j=1M,P_\mathrm{ref} = \{G^{(i)}_\mathrm{ref}\}_{i=1}^N, \qquad Q_\mathrm{gen} = \{G^{(j)}_\mathrm{gen}\}_{j=1}^M,9 and evaluates on held-out test data to produce the final PGD. This selection rule is described as yielding the tightest lower bound on {dk}k=1K\{d_k\}_{k=1}^K0 subject to the candidate descriptors (Krimmel et al., 7 Oct 2025).

The implementation is distributed as the polygraph-benchmark library. Installation is specified for Python {dk}k=1K\{d_k\}_{k=1}^K1, and the API exposes functions such as evaluate_pgd and load_graphs. The default descriptor list in the usage example is ["degree","clustering","orbit4","orbit5","spectrum","gin"], with classifier="tabpfn", folds=4, and a random seed. The command-line interface supports polygraph eval --method pgd ..., and the reported JSON output includes best_descriptor, pgd, and descriptor-specific subscores. TabPFN is described as fast to train, requiring approximately seconds per fold for typical graph-descriptor vectors, and as empirically yielding tighter bounds than logistic regression in ablation (Krimmel et al., 7 Oct 2025).

6. Empirical behavior and interpretive use

The empirical validation reported for PolyGraph covers bias and variance studies, synthetic perturbation tests, training-dynamics analyses, and benchmarking across graph generative models. On procedural datasets such as SBM, Planar, and Lobster, standard RBF-MMD and Gaussian-TV MMD are reported to show high bias at small sample sizes of {dk}k=1K\{d_k\}_{k=1}^K2–{dk}k=1K\{d_k\}_{k=1}^K3 and large variance under subsampling, with relative fluctuations of {dk}k=1K\{d_k\}_{k=1}^K4–{dk}k=1K\{d_k\}_{k=1}^K5. Even the unbiased MMD estimator is described as unreliable without hundreds of samples (Krimmel et al., 7 Oct 2025).

Under five synthetic perturbation types—edge deletion/addition, rewiring, mixing with Erdős–Rényi graphs, and edge-swap—PGD tracks perturbation magnitude monotonically, with Spearman correlation approximately {dk}k=1K\{d_k\}_{k=1}^K6–{dk}k=1K\{d_k\}_{k=1}^K7. For the diffusion model DiGress on Planar-L, varying denoising steps from {dk}k=1K\{d_k\}_{k=1}^K8 to {dk}k=1K\{d_k\}_{k=1}^K9 yields validity increasing from dk:G→Rnk,xi=dk(G(i)),d_k:\mathcal{G}\to\mathbb{R}^{n_k},\qquad x_i = d_k\bigl(G^{(i)}\bigr),0 to dk:G→Rnk,xi=dk(G(i)),d_k:\mathcal{G}\to\mathbb{R}^{n_k},\qquad x_i = d_k\bigl(G^{(i)}\bigr),1; PGD is reported to have almost perfect linear correlation with validity, with Pearson dk:G→Rnk,xi=dk(G(i)),d_k:\mathcal{G}\to\mathbb{R}^{n_k},\qquad x_i = d_k\bigl(G^{(i)}\bigr),2, while MMD metrics fall in the range dk:G→Rnk,xi=dk(G(i)),d_k:\mathcal{G}\to\mathbb{R}^{n_k},\qquad x_i = d_k\bigl(G^{(i)}\bigr),3–dk:G→Rnk,xi=dk(G(i)),d_k:\mathcal{G}\to\mathbb{R}^{n_k},\qquad x_i = d_k\bigl(G^{(i)}\bigr),4. During training on SBM-L, Lobster-L, and Planar-L, PGD and validity increase monotonically, whereas MMD is described as erratic and sometimes negatively correlated (Krimmel et al., 7 Oct 2025).

Across models including AutoGraph, DiGress, ESGG, and GRAN, and across datasets including SBM-L, Lobster-L, Planar-L, Proteins, and molecule sets, PGD yields rankings aligned with validity, uniqueness, and novelty measures. Descriptor-specific subscores reveal which structural aspects a model captures. This suggests an interpretive division of labor within the framework: the scalar PGD supports model ranking on a common scale, while the subscore profile functions as a structural diagnostic.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to PolyGraph Framework.