Papers
Topics
Authors
Recent
Search
2000 character limit reached

Likelihood Inference for Latent Network Models under Snowball Sampling

Published 19 Jun 2026 in stat.ME | (2606.21466v1)

Abstract: Snowball sampling is a widely used design for collecting network data from large or hard-to-reach populations, yet naive inference that ignores the sampling mechanism produces systematically biased parameter estimates. We derive the exact likelihood of a multi-wave snowball sample for the class of continuous latent space (CLS) models, in which edges form independently conditional on latent vertex-level quantities, and show that conditional edge independence reduces the marginalization over unobserved network configurations to a closed-form expression portable across the entire CLS class. We develop a stochastic Expectation-Maximization algorithm for the Euclidean latent distance model as a concrete implementation, and apply the framework to the large-scale co-inventor network of German semiconductor patent applicants by drawing multiple snowball samples. We find that the naive procedure severely underestimates latent space variance, produces networks with nearly twice the observed edge count, and achieves a spectral goodness-of-fit nine times worse than the corrected model, which directly affects the quantitative interpretation of covariate effects.

Summary

  • The paper develops an exact closed-form likelihood that adjusts for snowball sampling bias in continuous latent space network models.
  • It introduces a stochastic EM estimator that employs quasi-Monte Carlo integration for scalable latent position inference in large networks.
  • Empirical results on co-inventor networks validate the approach by matching observed edge counts and capturing accurate clustering and degree patterns.

Exact Likelihood Inference for Latent Network Models under Snowball Sampling

Overview

The paper develops a principled likelihood framework for inference in network models fitted to multi-wave snowball samples, specifically targeting continuous latent space (CLS) models. The central result is an exact, closed-form likelihood representation for the snowball sampling design over the CLS class, which solves a longstanding computational and inferential bottleneck: sampling-aware estimation that avoids explicit enumeration of unobserved network configurations and generalizes across widely used network models (e.g., latent space, stochastic block, random dot product, and graphon models).

The technical novelty relies on exploiting conditional edge independence given latent vertex-level quantities, reducing intractable marginalizations to computationally practical expressions. As a concrete implementation, the paper introduces a stochastic EM (sEM) estimator for the Euclidean latent distance model, demonstrates its scalability via quasi-Monte Carlo integration, and validates its performance through simulations and a substantive large-scale application in co-inventor networks.

Likelihood Formulation for Snowball Samples

Snowball sampling adaptively recruits network neighbors in waves, resulting in samples systematically biased toward highly connected subgraphs. Naive inference ignoring the sampling mechanism inflates edge density and compresses estimated latent space variance. Prior literature offered either design-based estimators for specific graph statistics or model-based approaches confined to marginalization over unobserved configurations (tractable only with restrictive or model-specific assumptions).

This work rectifies those limitations for CLS models. For any network model where edges form independently across vertex pairs conditional on latent variables, the likelihood of observing a snowball sample factorizes over dyads:

  • In the ErdÅ‘s-Rényi (ER) case, the marginalizations yield closed-form likelihoods for edge probability, properly adjusting for over-representation of dense clusters and high-degree vertices.
  • For general CLS models, the likelihood depends on latent positions of sampled and unsampled vertices only through exclusion probabilities—the chance an unsampled vertex remains disconnected from sampled waves.

The mathematical development demonstrates identifiability restrictions: sampled vertex positions are estimable, while unsampled positions are constrained only via exclusion probabilities, preserving equivalence classes in latent configurations.

Scaling and Estimation: Stochastic EM for Latent Distance Models

Direct latent position inference scales poorly with network size due to quadratic computational cost. The paper's sEM procedure for the latent distance model proceeds as follows:

  • Integrates out unsampled latent positions, approximating their contribution via QMC, thus reducing dimensionality and enabling tractable inference for large networks.
  • Samples latent positions of observed nodes using an elliptical slice within Gibbs MCMC, optimizing parameters (aa, latent variance yy) via L-BFGS on the sEM Q-function.
  • Simulation studies establish that the snowball-corrected estimator remains unbiased across realistic sampling fractions, while the naive estimator systematically underestimates latent space variance unless the sample covers nearly the full network.

Empirical Validation: The Co-Inventor Network Application

A large-scale German semiconductor co-inventor network (N=5979N=5979; $21,073$ edges) is analyzed, exhibiting pronounced small-world characteristics and heavy-tailed degree distributions. The naive latent distance model grossly misstates edge probabilities and latent space dispersion, yielding networks with nearly double the observed edge count and spectral goodness-of-fit (SGOF) nine times worse than the corrected model.

Key findings:

  • The snowball-corrected model generates latent positions distributed across a much more diffuse space, aligns mean edge count with observed topology, and explains degree variability and clustering more accurately.
  • Aggregated parameter estimates across hundreds of snowball samples are robustly derived via modal meta-analysis (ZEMBA), further mitigating effects of sample heterogeneity.
  • Corrected models yield substantially lower and more variable edge probabilities, faithfully capturing gender homophily and spatial proximity effects, and amplifying the geographic distance penalty compared to the naive model.

Theoretical and Practical Implications

The main theoretical advance is portable likelihood-based estimation for CLS models under snowball sampling, applicable to latent space, stochastic block, and graphon models. For models with dyad dependence (e.g., ERGM), the framework does not directly apply; further likelihood factorization is not possible, and the conditional approach remains necessary for those cases.

Non-ignorable ego selection remains an open challenge: if seed vertices are chosen by network attributes, proper modeling is computationally costly and unsolved for general CLS models.

Empirical results expose structural limitations of Euclidean latent space models for networks with small-world and heavy-tailed degree distributions. Even snowball-corrected inference cannot fully capture clustering and degree heterogeneity. Extending the framework to hyperbolic latent space or graphon models is suggested, as is incorporation of covariates and degree-correcting random effects. Bayesian implementation is natural and computationally feasible given the QMC-based exclusion probability computation.

Conclusion

The derivation of exact closed-form likelihoods for snowball samples in CLS network models enables statistically valid, scalable inference in large or partially observed networks, overcoming fundamental challenges in sampling bias correction. The developed tools provide a foundation for principled network model fitting in settings where full observation is infeasible but detailed latent structure estimation remains crucial. Future directions include extension to richer latent geometries, degree heterogeneity, and robust treatment of non-ignorable selection mechanisms.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.