---
title: General Path Models Overview
url: https://www.emergentmind.com/topics/general-path-models
type: topic
---

# General Path Models Overview

“General path models” is not a single, uniform formalism. In the cited literature, the expression designates several families of models in which a path, pathway parameter, or path-dependent transition rule organizes the construction of distributions, closures, counterfactuals, predictions, or efficiency scores. The term covers, among other objects, parametric paths through distribution families, admissible path systems in graphs, path loss laws in wireless networks, path diagrams in SEM, path-specific effects in causal mediation, higher-order models for path data, AST-path representations of programs, branching paths through model space, and path-based DEA projections [1409.2768] [1702.06112] [1611.04704] [2011.06436] [1801.06069] [2501.08302] [1803.09544] [2103.03462] [2311.16382] [2008.10706].

## 1. Scope and recurrent structure

Across these literatures, the central modeling move is to replace an unstructured description by an explicitly path-indexed one. In some cases the path is a continuous parametric route in distribution space; in others it is a family of admissible graph paths, an observed sequence through a network, a causal pathway in a DAG, a syntactic path in an AST, or a nested sequence of models. Taken together, these works suggest a recurring structure: a path object, a rule for traversing or selecting it, and an induced operator such as a closure, projection, likelihood, or counterfactual distribution.

| Domain | Path object | Representative formal device |
|---|---|---|
| Distribution theory | pathway parameter \(\alpha\) or \(q_1,q_2\) | generalized type-1 beta, generalized type-2 beta, generalized gamma |
| Graph convexity | family \(\mathcal{P}\) of admissible graph paths | interval operator \(I_{\mathcal P}\); matrix path convexity |
| Network path data | observed trigrams and state nodes | convex NMF factorization of regularized second-order transitions |
| Causal and social-science models | path diagrams and causal paths \(\pi\) | reflexive SEMs, \(Y(\pi,a,a')\), PDSEMs |
| Computation and model search | AST paths and forward model paths | path-contexts \(\langle x_s,p,x_f\rangle\); forward stability selection |
| DEA | projection path \(\phi_o(\theta)\) | GS path-based model with ID, MO, and PR |

A second recurrent feature is that “path” often mediates between local and global structure. A local rule—such as chord constraints in a graph path, a transition kernel conditioned on the previous node, or a structural equation tied to a particular regime—induces a global object such as a convexity space, a higher-order Markov chain, or a counterfactual trajectory.

## 2. Distributional pathways and propagation laws

In stochastic modeling, the pathway model is a univariate parametric family that moves continuously between three distribution classes by means of a pathway parameter \(\alpha\): the generalized type-1 beta family for \(\alpha<1\), the generalized gamma family in the limit \(\alpha\to 1\), and the generalized type-2 beta family for \(\alpha>1\) [1409.2768]. Its core scalar forms are
\[
f_1^{*}(x)=c_1^{*}|x|^{\gamma}\left[1-a(1-\alpha)|x|^{\delta}\right]^{\frac{\eta}{1-\alpha}}, \qquad \alpha<1,
\]
\[
f_2^{*}(x)=c_2^{*}|x|^{\gamma}\left[1+a(\alpha-1)|x|^{\delta}\right]^{-\frac{\eta}{\alpha-1}}, \qquad \alpha>1,
\]
and
\[
f_3^{*}(x)=c_3^{*}|x|^{\gamma}\exp\!\left(-a\eta|x|^{\delta}\right), \qquad \alpha\to 1.
\]
The interpretation is explicit: \(\alpha<1\) yields finite support, \(\alpha=1\) yields exponential-type tails, and \(\alpha>1\) yields power-law tails. The pathway parameter therefore determines a continuous route from compactly supported densities through generalized gamma laws to heavy-tailed type-2 beta laws [1409.2768].

The same paper embeds this pathway idea in input-output and reaction-diffusion settings. With \(x\) as input or production, \(y\) as output or destruction, and \(u=x-y\) as the residual, Laplace-type and gamma-type residual laws arise from simple assumptions on \(x\) and \(y\). Section 4.1 then replaces exponential kernels in reaction-rate integrals by pathway analogues. The generalized integral
\[
I_p=\int_0^\infty x^\gamma
\left[1+a(q_1-1)x^\delta\right]^{-1/(q_1-1)}
\left[1+b(q_2-1)x^\rho\right]^{-1/(q_2-1)}\,dx
\]
reduces to classical exponential-kernel integrals as \(q_1,q_2\to 1\), but otherwise introduces type-2 beta heavy-tailed factors. The same framework is used to thicken or thin gamma tails through appended Mittag–Leffler and Bessel factors, with \(a>0\) producing thicker tails and \(a<0\) thinner ones [1409.2768].

Within this distributional pathway formalism, specific well-known models appear as special cases. For \(x>0,\gamma=0,a=1,\delta=1,\eta=1\), the \(\alpha<1\) form becomes Tsallis statistics; for \(a=1,\delta=1,\eta=1\), the \(\alpha>1\) form becomes superstatistics. The paper also states that Gaussian, Maxwell–Boltzmann, and standard gamma distributions arise under suitable choices of \(\delta\), \(\gamma\), and scale parameters [1409.2768].

A different use of “path model” occurs in wireless networks, where the path object is the propagation law \(\ell(x)\). Two general path loss classes are analyzed: the singular model
\[
\ell(x)=\|x\|^{-\alpha}, \qquad \alpha>2,
\]
and the bounded model
\[
\ell(x)=\left(\epsilon+\|x\|^\alpha\right)^{-1}, \qquad \alpha>2,\ \epsilon>0,
\]
with \(\delta=2/\alpha\in(0,1)\) as the main asymptotic parameter [1611.04704]. In this setting, lower SIR tails decay polynomially and upper tails are either exponential-type or power-law, depending on the network class and whether path loss is singular or bounded. The paper’s summary gives, for example, \(1-P_s(\theta)=\Theta(\theta^\delta)\) for ad hoc networks with singular path loss and \(P_s(\theta)=\Theta(\theta^{-\delta})\) for simple cellular networks with singular path loss, while bounded path loss yields \(P_s(\theta)=e^{-\Theta(\theta^\delta)}\) in the upper tail [1611.04704].

This juxtaposition is instructive because it shows two distinct meanings of path modeling in probability. In the pathway model, a scalar parameter traces a route through distribution families. In SIR asymptotics, the path model is a path loss law whose near-field boundedness or singularity governs tail behavior. The commonality is not semantic identity but the use of a path-specifying rule to determine asymptotic or distributional regime.

## 3. Path-defined graph and network constructions

In graph theory, a path model is a family \(\mathcal P\) of graph paths from which one derives an interval operator and a convexity space [1702.06112]. For a connected simple graph \(G\), the interval function is
\[
I_{\mathcal P}(S)
=
S\cup\{z\in V(G)\mid \exists u,v\in S\text{ such that } z \text{ lies on a }uv\text{-path }P\in\mathcal P\},
\]
and the convex sets are exactly the fixed points
\[
\mathcal C_{\mathcal P}=\{S\subseteq V(G)\mid I_{\mathcal P}(S)=S\}.
\]
The framework’s main generalization is matrix path convexity. Four symmetric length matrices \(A,B,C,D\) specify, for each endpoint pair \((i,j)\), lower and upper bounds on path length and chord length. A path belongs to \(\mathcal P(A,B,C,D)\) iff it satisfies \(|P|\ge A(i,j)\), \(|P|\le B(i,j)\), and every chord has length between \(C(i,j)\) and \(D(i,j)\). The resulting convexity space \((V(G),\mathcal C_{\mathcal P})\) unifies geodesic, monophonic, \(P_3\), triangle-path, detour, and all-path convexities [1702.06112].

The constant-matrix specialization \((a,b,c,d)\)-path convexity isolates fixed path models in which admissible paths have length between \(a\) and \(b\) and all chords have length between \(c\) and \(d\). This formulation is algorithmically significant. When \(A,B,C,D\) are part of the input, MATRIX CONVEX SET is CoNP-complete, MATRIX INTERVAL DETERMINATION and MATRIX CONVEX HULL DETERMINATION are NP-complete, and the associated optimization problems are NP-hard or NP-complete. By contrast, for fixed constants \(a,b,c,d\), \((a,b,c,d)\)-INTERVAL DETERMINATION is in \(\mathrm P\), and for bounded treewidth graphs all six standard problems become linearly solvable via CMSOL\(_2\) formulations and Courcelle-type meta-theorems [1702.06112]. The framework is therefore both descriptive and generative: it systematizes known path convexities and supports the definition of new ones such as \(k\)-\(g\), \(m_k\), \((k,\ell)\)-path, \(k\)-path, and Hamiltonian convexities [1702.06112].

A second graph- and network-based notion of general path models appears in higher-order models for path data. Here the primitive observations are sequences
\[
(x_1\to x_2\to \cdots \to x_\ell),
\]
not aggregated edges, and first-order Markov assumptions are often inadequate [2501.08302]. For each physical node \(j\), the paper constructs a trigram count matrix \(A_{ki}=\#\{i\to j\to k\}\), a second-order MLE \(M2_{ki}\), and a first-order successor distribution \(M1_k\). A Dirichlet prior centered on \(M1\) yields the regularized second-order estimate
\[
X_{ki}=\frac{A_{ki}+\mu M1_k}{\sum_{k'}A_{k'i}+\mu},
\]
with \(\mu\) chosen by leave-one-out cross-validation [2501.08302].

The central compression step factorizes \(X\) by convex NMF into a small number of latent state nodes:
\[
\hat X_{ki}=\sum_\alpha \hat P(i\to \alpha)\hat P(\alpha\to k), \qquad
\hat{\mathbf X}=\hat{\mathbf X}_{\text{out}}\hat{\mathbf X}_{\text{in}}^\top.
\]
When \(r=1\), the model collapses to first-order behavior at node \(j\); when \(r\) is large enough to match predecessor-specific columns, it approaches a full second-order model. The trade-off between complexity and accuracy is quantified by flow overlap,
\[
\text{flow overlap}(p,q)=\sum_i \min(p(i),q(i)),
\]
and, for matrices, by a column-weighted mean overlap [2501.08302]. In air-travel data the method identified interpretable state nodes for hubs such as Denver, Atlanta, Dallas–Fort Worth, and Chicago, while in synthetic information flow on the Lazega law firm network it recovered overlapping communities that were not visible in the first-order graph [2501.08302].

These two bodies of work share an important formal trait: a path family is not merely descriptive. It defines an operator. In matrix path convexity, the operator is \(I_{\mathcal P}\) and its iterates. In concise higher-order networks, it is the state-node Markov chain induced by \(\hat{\mathbf X}_{\text{out}}\) and \(\hat{\mathbf X}_{\text{in}}\). In both cases, changing the admissible path class changes the induced global geometry.

## 4. Path diagrams, path-specific causation, and path-dependent structure

In the social sciences, path analysis is the use of path diagrams and linear structural equations to represent relations among observed and latent variables [2011.06436]. In the reflexive model treated there, latent constructs \(\xi\) and \(\eta\) are linked to indicator vectors \(X\) and \(Y\) through measurement equations
\[
X=\mu_X+\beta_{X|\xi}(\xi-\mu_\xi)+\varepsilon_{X|\xi}, \qquad
Y=\mu_Y+\beta_{Y|\eta}(\eta-\mu_\eta)+\varepsilon_{Y|\eta},
\]
with jointly normal latent variables and mutually independent error blocks [2011.06436]. The observable cross-covariance has rank one,
\[
\Sigma_{YX}=\beta_{Y|\eta}\sigma_{\xi\eta}\beta_{X|\xi}^\top,
\]
so the regression of \(Y\) on \(X\) is a reduced-rank regression. A central identification point is that the latent correlation \(cor(\xi,\eta)\) is not generally identifiable without further assumptions, but
\[
\rho^*:=cor\{E(\xi\mid X),E(\eta\mid Y)\}
\]
is identifiable, and \(|\rho^*|\) equals the first canonical correlation between \(X\) and \(Y\) [2011.06436]. This distinction underlies the paper’s comparison of SEM, PLS, envelope models, and SERR.

Causal mediation theory extends the path idea from diagrammatic representation to path-specific counterfactuals [1801.06069]. The total effect \(E\{Y(a)-Y(a')\}\), the natural direct effect \(E\{Y(a,M(a'))-Y(a',M(a'))\}\), and the natural indirect effect \(E\{Y(a,M(a))-Y(a,M(a'))\}\) are special cases of path-specific effects along chosen sets of directed paths \(\pi\). The general object is \(Y(\pi,a,a')\), which sets \(A\) to \(a\) along paths in \(\pi\) and to \(a'\) along paths in \(\overline{\pi}\) [1801.06069]. Identification is harder than for total effects because adjustment for confounding is not sufficient; path-specific effects depend on cross-world counterfactual structure. The paper develops graphical criteria based on recanting witnesses and, with hidden variables, recanting districts. In hidden-variable settings, marginal path-specific effects are identifiable iff there is no recanting district for \(\pi\) and the total effect is identifiable; in that case the identification functional is an edge-specific truncated district factorization [1801.06069].

Path Dependent Structural Equation Models generalize longitudinal SEMs by allowing the qualitative causal structure itself to depend on the realized path through a discrete state space [2008.10706]. A PDSEM uses a finite state space \({\bf s}=\{s^1,\dots,s^m,s^*\}\), state-specific DAGs for the initial state, transition-specific CDAGs for allowed transitions \((i,j)\in{\cal T}\), and state-determining variables \(S_1,S_{ij}\). The observed-data distribution is
\[
p_\infty({\bf V})=p_1({\bf V}_1)\prod_{t=1}^\infty
\Bigg[
\prod_{(i,j)\in{\cal T}}
\big(p_{ij}({\bf V}_{ij}\mid{\bf V}_i)\big)^{\mathbb I(s_t^i,s_{t+1}^j)}
\Bigg]1^{\mathbb I(s_t^*)},
\]
and interventions may change not only variable values but also subsequent state trajectories [2008.10706]. For fully observed PDSEMs, a generalized g-formula identifies counterfactual distributions; with hidden variables, identification is characterized using nested Markov factorizations on the induced ADMGs and CADMGs [2008.10706].

A common misunderstanding is to treat all path diagrams as fixed-structure models. The cited work shows three increasingly rich levels: reflexive SEMs with a single latent-variable diagram, mediation models with path-specific counterfactual semantics on a fixed DAG, and PDSEMs where the very graph applied at time \(t+1\) depends on the realized path. This suggests that “path model” in causal work ranges from graphical depiction to full path-dependent data-generating mechanism.

## 5. Computational path representations and model-space paths

In program analysis, a general path-based representation is built from paths in abstract syntax trees rather than from flat token sequences [1803.09544]. An AST is formalized as \(\langle N,T,X,s,\delta,val\rangle\), with parent function \(\pi\), and an AST path of length \(k\) is a sequence
\[
n_1 d_1 n_2 d_2 \dots d_k n_{k+1},
\]
where \(d_i\in\{\uparrow,\downarrow\}\) indicates traversal to parent or child. A path-context is the triplet
\[
\langle x_s,p,x_f\rangle,
\]
with \(x_s=val(start(p))\) and \(x_f=val(end(p))\) [1803.09544]. Abstraction functions \(\alpha\) yield abstract path-contexts \(\langle x_s,\alpha(p),x_f\rangle\), and hyperparameters \(max\_len\) and \(max\_w\) constrain path length and width.

This representation is explicitly task- and model-agnostic. The same path-contexts drive CRF-based and word2vec-based learning for variable naming, method naming, and full type prediction in JavaScript, Java, Python, and C\# [1803.09544]. The reported exact-match results for CRF models are 67.3% on JavaScript variable naming, 58.2% on Java, 56.7% on Python, and 56.1% on C\#; for method naming, 53.1% on JavaScript, 47.3% exact with F1 49.9 on Java, and 51.1% on Python; for Java full type prediction, 69.1% [1803.09544]. In the word2vec setting on JavaScript variable naming, AST paths yield 40.4% accuracy, compared with 23.2% for AST neighbors without paths and 20.6% for linear token-stream contexts [1803.09544]. The representation is therefore “general” in the strong sense of being usable across tasks, algorithms, and languages.

A different computational meaning of “path model” arises in model path selection. Here the path is a sequence of nested models produced by forward stability selection under data perturbations [2103.03462]. At forward step \(j\), with current model \(\hat f_{(j-1)}\), one draws \(B\) subsamples, fits all one-variable extensions, records the winning covariate \(k_i^*\) on each subsample, and computes selection proportions
\[
\hat\theta_k=\frac{1}{B}\sum_{i=1}^B C_{k,i},
\]
where \(C_{k,i}=\mathbf 1\{k_i^*=k\}\) [2103.03462]. The hypotheses
\[
H_{0,k}^{(j)}:\theta_k\ge \theta_q \text{ for all } X_q\in \mathbf X_{j+}
\]
formalize forward stability. The practical algorithm stops when one multinomial cell count reaches \(r\), computes the smallest \(D\) guaranteeing \(\mathbb P(C\mid r,D)\ge P^*\), and retains all covariates with count exceeding \(r-D\) [2103.03462].

The output is not a single forward path but a branching tree of plausible model paths. On the diabetes data, the method found 67 distinct paths; 24 of 67 models (36%) outperformed the lasso model chosen by CV, and 40 of 67 models (60%) outperformed the CV-tuned forward stepwise model [2103.03462]. On the breast-cancer data, MPS selected 79 out of 84 possible 3-variable logistic-regression models, but only 1 regression-tree path, illustrating that path multiplicity can itself diagnose model misspecification or structural instability [2103.03462].

These two computational uses of paths differ sharply in ontology. AST paths are syntactic objects extracted from a structured input. Model paths are trajectories through subset space. Yet both provide a compact alternative to unstructured search: the former by encoding code semantics through tree relations, the latter by organizing the Rashomon set of plausible models into a nested, visualizable family.

## 6. Path-based efficiency analysis and general implications

In DEA, a path-based model is defined over the VRS technology
\[
T=\left\{(x,y)\in\mathbb{R}^m\times\mathbb{R}^s \mid X\lambda\le x,\ Y\lambda\ge y,\ \lambda\ge 0,\ e^\top\lambda=1\right\},
\]
together with a direction \(g_o=(g_o^x,g_o^y)\) and two scalar functions \(\psi^x,\psi^y\) [2311.16382]. The general GS model is
\[
\min_{\theta,\lambda}\ \theta
\]
subject to
\[
X\lambda \le x_o+(\psi^x(\theta)-1)g_o^x,\qquad
Y\lambda \ge y_o+(\psi^y(\theta)-1)g_o^y,
\]
plus the VRS constraints [2311.16382]. This induces the path
\[
\phi_o^x(\theta)=x_o+(\psi^x(\theta)-1)g_o^x,\qquad
\phi_o^y(\theta)=y_o+(\psi^y(\theta)-1)g_o^y,
\]
with efficiency score \(\theta_o^*\le 1\) and projection \(\phi_o(\theta_o^*)\in\partial T\) [2311.16382]. BCC radial models, DDF-g, HDF, and GDF are all special cases of this GS construction.

The paper studies three properties: indication of strong efficiency,
\[
\theta_o^*=1 \Rightarrow (x_o,y_o)\in \partial^S T,
\]
strict monotonicity,
\[
(x_o,y_o)\succnsim (x_p,y_p)\Rightarrow \theta_o^*<\theta_p^*,
\]
and strong efficiency of projections,
\[
\phi_o(\theta_o^*)\in\partial^S T.
\]
Under mild assumptions, strict monotonicity and strong efficiency of projections are equivalent, and both imply indication [2311.16382]. The paper also characterizes the narrow class in which all three hold: GS models with GS range directions over ideal technology sets. In general VRS technologies, however, standard path-based models typically fail all three properties, unlike slacks-based graph models, which typically satisfy them [2311.16382].

This final comparison clarifies a broad point about general path models. The choice of path is rarely innocuous. In distribution theory it determines tail class and support; in graph convexity it determines closure and complexity; in higher-order network models it determines the effective memory order; in mediation it determines which counterfactual effect is even identifiable; in program analysis it determines what structure a learner can exploit; in model path selection it determines which alternatives remain visible; and in DEA it determines whether projections land on strong or weak frontiers. The literature therefore does not support a unitary theory of “general path models.” It supports a general modeling strategy in which a carefully specified path object mediates between local structure and global inference.

Source: https://www.emergentmind.com/topics/general-path-models