PolyGraph: Graph Generative Model Evaluation
- PolyGraph is a modular evaluation framework that transforms graph evaluation by using classifier-based, variational estimates of the Jensen–Shannon divergence to compare reference and generated graphs.
- It employs a three-stage pipeline—descriptor extraction, classifier fitting, and metric computation—to produce the PolyGraph Discrepancy (PGD), a bounded quality metric ranging from 0 to 1.
- The framework overcomes limitations of traditional MMD metrics by providing absolute, robust, and descriptor-comparable scores with improved performance at moderate sample sizes.
PolyGraph is a modular evaluation pipeline for graph generative models that replaces ad hoc, kernel-based metrics with classifier-based, variational estimates of the Jensen–Shannon distance between a reference graph distribution and a generated distribution . Its central metric, PolyGraph Discrepancy (PGD), is designed to provide an absolute, bounded, and descriptor-comparable assessment of generative quality, in contrast to descriptor-dependent Maximum Mean Discrepancy (MMD) scores. The framework is publicly available as a benchmarking library for graph generative models and is organized around descriptor extraction, probabilistic discrimination, and held-out metric computation (Krimmel et al., 7 Oct 2025).
1. Problem setting and scope
PolyGraph is motivated by a limitation of prevailing evaluation practice in graph generation: existing methods rely primarily on MMD metrics based on graph descriptors. In that formulation, descriptor choice and kernel parametrization strongly influence numerical values, so scores can rank models within a fixed setup but do not define an absolute scale and are not comparable across descriptor families. PolyGraph addresses this by evaluating how well a probabilistic classifier can distinguish reference graphs from generated graphs after descriptor featurization, and by interpreting the resulting held-out log-likelihood as a variational lower bound on Jensen–Shannon divergence (Krimmel et al., 7 Oct 2025).
The framework operates on two multisets of graphs,
and a family of graph descriptors . Each descriptor induces a marginal comparison problem in feature space rather than on raw combinatorial objects. This design makes evaluation modular: one may inspect descriptor-specific subscores while also deriving a single summary metric. A common misconception is that PolyGraph eliminates descriptor dependence entirely; the framework instead standardizes the scale and interpretation of descriptor-wise scores and then selects the tightest lower bound among the available descriptors.
2. Pipeline structure
The PolyGraph pipeline consists of three stages: descriptor extraction, classifier fitting, and metric computation. Descriptors convert graphs into fixed-dimensional vectors,
after which a probabilistic binary classifier is trained for each descriptor, with label for reference and for generated graphs. The classifier is fit on disjoint “fit” subsets and evaluated on held-out “test” subsets of equal size (Krimmel et al., 7 Oct 2025).
The descriptor set described for the framework includes both classical statistics and learned embeddings.
| Descriptor family | Examples |
|---|---|
| Histogram/statistical descriptors | Degree histograms, clustering coefficient histograms, Laplacian spectrum histograms |
| Motif-based descriptors | Graphlet/orbit counts: , |
| Learned representations | GIN embeddings |
For classifier fitting, the framework uses TabPFN, described as a hyperparameter-free transformer approximate Bayesian classifier. The intended role of this choice is methodological rather than incidental: calibrated probabilistic outputs are required because the downstream quantity is not classification accuracy but data log-likelihood on held-out samples. This suggests that PolyGraph is less a single scalar metric than a benchmarking protocol in which classifier quality directly controls the tightness of the variational lower bound (Krimmel et al., 7 Oct 2025).
3. Variational formulation of PGD
The mathematical basis of PolyGraph is the Jensen–Shannon divergence
0
and the associated Jensen–Shannon distance
1
The key variational identity used in the framework is
2
Accordingly, for any classifier 3,
4
The held-out average log-likelihood for descriptor 5 is
6
In the framework description, 7 is treated as a lower bound on the JS divergence between descriptor-marginal distributions, and each descriptor yields a bounded metric
8
Descriptor-wise evaluation is then aggregated by a max-reduction: 9 The summary metric is explicitly presented as the maximally tight lower bound available from the candidate descriptor family. This does not imply that the maximum reconstructs the full graph-distribution distance; rather, it yields the strongest lower bound accessible through the chosen feature maps.
4. Relation to MMD-based evaluation
For comparison, the framework recalls the MMD functional
0
or, in RKHS variational form,
1
The reported drawbacks for graph generative model evaluation are threefold: absence of intrinsic scale, incomparability across descriptor and kernel choices, and high bias and variance at typical sample sizes. PolyGraph is positioned against these issues by producing scores in 2, making descriptor-wise values directly comparable, and exhibiting greater robustness at moderate sample sizes of 3–4 graph pairs (Krimmel et al., 7 Oct 2025).
Two methodological cautions are built into the framework. First, no single descriptor suffices under all perturbations; the descriptor set should therefore be interpreted as a family of complementary probes rather than interchangeable surrogates. Second, descriptor selection must avoid test-set leakage. The framework addresses this by using cross-validated fit-split performance for descriptor selection and reserving the held-out test split exclusively for final reporting. A plausible implication is that PolyGraph treats evaluation itself as a statistical estimation problem, not merely as score computation.
5. Summary metric, cross-validation, and implementation
The framework’s summary metric is not computed by averaging descriptor scores. Instead, for each descriptor 5, one performs 6-fold stratified cross-validation on the fit split to estimate 7, selects
8
then retrains on all fit data with 9 and evaluates on held-out test data to produce the final PGD. This selection rule is described as yielding the tightest lower bound on 0 subject to the candidate descriptors (Krimmel et al., 7 Oct 2025).
The implementation is distributed as the polygraph-benchmark library. Installation is specified for Python 1, and the API exposes functions such as evaluate_pgd and load_graphs. The default descriptor list in the usage example is ["degree","clustering","orbit4","orbit5","spectrum","gin"], with classifier="tabpfn", folds=4, and a random seed. The command-line interface supports polygraph eval --method pgd ..., and the reported JSON output includes best_descriptor, pgd, and descriptor-specific subscores. TabPFN is described as fast to train, requiring approximately seconds per fold for typical graph-descriptor vectors, and as empirically yielding tighter bounds than logistic regression in ablation (Krimmel et al., 7 Oct 2025).
6. Empirical behavior and interpretive use
The empirical validation reported for PolyGraph covers bias and variance studies, synthetic perturbation tests, training-dynamics analyses, and benchmarking across graph generative models. On procedural datasets such as SBM, Planar, and Lobster, standard RBF-MMD and Gaussian-TV MMD are reported to show high bias at small sample sizes of 2–3 and large variance under subsampling, with relative fluctuations of 4–5. Even the unbiased MMD estimator is described as unreliable without hundreds of samples (Krimmel et al., 7 Oct 2025).
Under five synthetic perturbation types—edge deletion/addition, rewiring, mixing with Erdős–Rényi graphs, and edge-swap—PGD tracks perturbation magnitude monotonically, with Spearman correlation approximately 6–7. For the diffusion model DiGress on Planar-L, varying denoising steps from 8 to 9 yields validity increasing from 0 to 1; PGD is reported to have almost perfect linear correlation with validity, with Pearson 2, while MMD metrics fall in the range 3–4. During training on SBM-L, Lobster-L, and Planar-L, PGD and validity increase monotonically, whereas MMD is described as erratic and sometimes negatively correlated (Krimmel et al., 7 Oct 2025).
Across models including AutoGraph, DiGress, ESGG, and GRAN, and across datasets including SBM-L, Lobster-L, Planar-L, Proteins, and molecule sets, PGD yields rankings aligned with validity, uniqueness, and novelty measures. Descriptor-specific subscores reveal which structural aspects a model captures. This suggests an interpretive division of labor within the framework: the scalar PGD supports model ranking on a common scale, while the subscore profile functions as a structural diagnostic.