---
title: 'PolyGraph: Graph Generative Model Evaluation'
url: https://www.emergentmind.com/topics/polygraph-framework
type: topic
---

# PolyGraph: Graph Generative Model Evaluation

PolyGraph is a modular evaluation pipeline for graph generative models that replaces ad hoc, kernel-based metrics with classifier-based, variational estimates of the Jensen–Shannon distance between a reference graph distribution \(P\) and a generated distribution \(Q\). Its central metric, PolyGraph Discrepancy (PGD), is designed to provide an absolute, bounded, and descriptor-comparable assessment of generative quality, in contrast to descriptor-dependent Maximum Mean Discrepancy (MMD) scores. The framework is publicly available as a benchmarking library for graph generative models and is organized around descriptor extraction, probabilistic discrimination, and held-out metric computation [2510.06122].

## 1. Problem setting and scope

PolyGraph is motivated by a limitation of prevailing evaluation practice in graph generation: existing methods rely primarily on MMD metrics based on graph descriptors. In that formulation, descriptor choice and kernel parametrization strongly influence numerical values, so scores can rank models within a fixed setup but do not define an absolute scale and are not comparable across descriptor families. PolyGraph addresses this by evaluating how well a probabilistic classifier can distinguish reference graphs from generated graphs after descriptor featurization, and by interpreting the resulting held-out log-likelihood as a variational lower bound on Jensen–Shannon divergence [2510.06122].

The framework operates on two multisets of graphs,
\[
P_\mathrm{ref} = \{G^{(i)}_\mathrm{ref}\}_{i=1}^N,
\qquad
Q_\mathrm{gen} = \{G^{(j)}_\mathrm{gen}\}_{j=1}^M,
\]
and a family of graph descriptors \(\{d_k\}_{k=1}^K\). Each descriptor induces a marginal comparison problem in feature space rather than on raw combinatorial objects. This design makes evaluation modular: one may inspect descriptor-specific subscores while also deriving a single summary metric. A common misconception is that PolyGraph eliminates descriptor dependence entirely; the framework instead standardizes the *scale and interpretation* of descriptor-wise scores and then selects the tightest lower bound among the available descriptors.

## 2. Pipeline structure

The PolyGraph pipeline consists of three stages: descriptor extraction, classifier fitting, and metric computation. Descriptors convert graphs into fixed-dimensional vectors,
\[
d_k:\mathcal{G}\to\mathbb{R}^{n_k},\qquad x_i = d_k\bigl(G^{(i)}\bigr),
\]
after which a probabilistic binary classifier \(D_k:\mathbb{R}^{n_k}\to[0,1]\) is trained for each descriptor, with label \(y=1\) for reference and \(y=0\) for generated graphs. The classifier is fit on disjoint “fit” subsets and evaluated on held-out “test” subsets of equal size [2510.06122].

The descriptor set described for the framework includes both classical statistics and learned embeddings.

| Descriptor family | Examples |
|---|---|
| Histogram/statistical descriptors | Degree histograms, clustering coefficient histograms, Laplacian spectrum histograms |
| Motif-based descriptors | Graphlet/orbit counts: \(\mathrm{Orb}_4\), \(\mathrm{Orb}_5\) |
| Learned representations | GIN embeddings |

For classifier fitting, the framework uses TabPFN, described as a hyperparameter-free transformer approximate Bayesian classifier. The intended role of this choice is methodological rather than incidental: calibrated probabilistic outputs are required because the downstream quantity is not classification accuracy but data log-likelihood on held-out samples. This suggests that PolyGraph is less a single scalar metric than a benchmarking protocol in which classifier quality directly controls the tightness of the variational lower bound [2510.06122].

## 3. Variational formulation of PGD

The mathematical basis of PolyGraph is the Jensen–Shannon divergence
\[
\mathrm{JS}(P\Vert Q)
=\tfrac12\,\mathrm{KL}(P\Vert M)+\tfrac12\,\mathrm{KL}(Q\Vert M),
\qquad
M=\tfrac12(P+Q),
\]
and the associated Jensen–Shannon distance
\[
d_{\mathrm{JS}}(P,Q)=\sqrt{\mathrm{JS}(P\Vert Q)},\qquad d_{\mathrm{JS}}\in[0,1].
\]

The key variational identity used in the framework is
\[
\mathrm{JS}(P\Vert Q)
=
\sup_{D:\mathcal{X}\to[0,1]}
\left[
\tfrac12\,E_{x\sim P}[\log_2 D(x)]
+\tfrac12\,E_{x\sim Q}[\log_2(1-D(x))]
\right]
+1.
\]
Accordingly, for any classifier \(D\),
\[
\mathcal{L}(D)
:=
\tfrac12\,E_{x\sim P}[\log_2 D(x)]
+\tfrac12\,E_{x\sim Q}[\log_2(1-D(x))]
\le \mathrm{JS}(P\Vert Q)-1.
\]
The held-out average log-likelihood for descriptor \(d_k\) is
\[
\mathcal{L}_k
=
\tfrac12\,E_{x\sim P_\mathrm{test}}[\log_2 D_k(x)]
+\tfrac12\,E_{x\sim Q_\mathrm{test}}[\log_2(1-D_k(x))].
\]
In the framework description, \(\mathcal{L}_k+1\) is treated as a lower bound on the JS divergence between descriptor-marginal distributions, and each descriptor yields a bounded metric
\[
\mathrm{PGD}_k=\sqrt{\max(\mathcal{L}_k,0)}\in[0,1].
\]

Descriptor-wise evaluation is then aggregated by a max-reduction:
\[
\mathrm{PGD}=\max_{k=1,\ldots,K}\mathrm{PGD}_k.
\]
The summary metric is explicitly presented as the *maximally tight* lower bound available from the candidate descriptor family. This does not imply that the maximum reconstructs the full graph-distribution distance; rather, it yields the strongest lower bound accessible through the chosen feature maps.

## 4. Relation to MMD-based evaluation

For comparison, the framework recalls the MMD functional
\[
\mathrm{MMD}^2(P,Q;k)
=
E_{x,x'\sim P}[k(x,x')]
-2E_{x\sim P,y\sim Q}[k(x,y)]
+E_{y,y'\sim Q}[k(y,y')],
\]
or, in RKHS variational form,
\[
\mathrm{MMD}(P,Q;k)
=
\sup_{\|f\|_{\mathcal H}\le 1}
\bigl(E_P[f(x)]-E_Q[f(x)]\bigr).
\]
The reported drawbacks for graph generative model evaluation are threefold: absence of intrinsic scale, incomparability across descriptor and kernel choices, and high bias and variance at typical sample sizes. PolyGraph is positioned against these issues by producing scores in \([0,1]\), making descriptor-wise values directly comparable, and exhibiting greater robustness at moderate sample sizes of \(250\)–\(1\,000\) graph pairs [2510.06122].

Two methodological cautions are built into the framework. First, no single descriptor suffices under all perturbations; the descriptor set should therefore be interpreted as a family of complementary probes rather than interchangeable surrogates. Second, descriptor selection must avoid test-set leakage. The framework addresses this by using cross-validated fit-split performance for descriptor selection and reserving the held-out test split exclusively for final reporting. A plausible implication is that PolyGraph treats evaluation itself as a statistical estimation problem, not merely as score computation.

## 5. Summary metric, cross-validation, and implementation

The framework’s summary metric is not computed by averaging descriptor scores. Instead, for each descriptor \(d_k\), one performs \(K\)-fold stratified cross-validation on the fit split to estimate \(\mathrm{PGD}_k^\mathrm{CV}\), selects
\[
k^*=\arg\max_k \mathrm{PGD}_k^\mathrm{CV},
\]
then retrains on all fit data with \(d_{k^*}\) and evaluates on held-out test data to produce the final PGD. This selection rule is described as yielding the tightest lower bound on \(d_{\mathrm{JS}}\) subject to the candidate descriptors [2510.06122].

The implementation is distributed as the `polygraph-benchmark` library. Installation is specified for Python \(\ge 3.8\), and the API exposes functions such as `evaluate_pgd` and `load_graphs`. The default descriptor list in the usage example is `["degree","clustering","orbit4","orbit5","spectrum","gin"]`, with `classifier="tabpfn"`, `folds=4`, and a random seed. The command-line interface supports `polygraph eval --method pgd ...`, and the reported JSON output includes `best_descriptor`, `pgd`, and descriptor-specific `subscores`. TabPFN is described as fast to train, requiring approximately seconds per fold for typical graph-descriptor vectors, and as empirically yielding tighter bounds than logistic regression in ablation [2510.06122].

## 6. Empirical behavior and interpretive use

The empirical validation reported for PolyGraph covers bias and variance studies, synthetic perturbation tests, training-dynamics analyses, and benchmarking across graph generative models. On procedural datasets such as SBM, Planar, and Lobster, standard RBF-MMD and Gaussian-TV MMD are reported to show high bias at small sample sizes of \(20\)–\(40\) and large variance under subsampling, with relative fluctuations of \(\pm 50\)–\(100\%\). Even the unbiased MMD estimator is described as unreliable without hundreds of samples [2510.06122].

Under five synthetic perturbation types—edge deletion/addition, rewiring, mixing with Erdős–Rényi graphs, and edge-swap—PGD tracks perturbation magnitude monotonically, with Spearman correlation approximately \(0.8\)–\(0.95\). For the diffusion model DiGress on Planar-L, varying denoising steps from \(15\) to \(90\) yields validity increasing from \(0\) to \(51\%\); PGD is reported to have almost perfect linear correlation with validity, with Pearson \(r=0.995\), while MMD metrics fall in the range \(r\approx 0.70\)–\(0.85\). During training on SBM-L, Lobster-L, and Planar-L, PGD and validity increase monotonically, whereas MMD is described as erratic and sometimes negatively correlated [2510.06122].

Across models including AutoGraph, DiGress, ESGG, and GRAN, and across datasets including SBM-L, Lobster-L, Planar-L, Proteins, and molecule sets, PGD yields rankings aligned with validity, uniqueness, and novelty measures. Descriptor-specific subscores reveal which structural aspects a model captures. This suggests an interpretive division of labor within the framework: the scalar PGD supports model ranking on a common scale, while the subscore profile functions as a structural diagnostic.

Source: https://www.emergentmind.com/topics/polygraph-framework