Papers
Topics
Authors
Recent
Search
2000 character limit reached

PST Hierarchy in Multi-Domain Research

Updated 8 July 2026
  • PST Hierarchy is a term describing various ordered frameworks that structure inference, computation, and classification across diverse fields.
  • It encompasses methodologies from progressive segmented training in continual learning to subgroup characterizations in finite group theory, rooted-tree constructions in topological data analysis, semidefinite hierarchies in quantum information, and layered approaches in causal inference.
  • This multi-domain concept offers practical insights by aligning specific ordering principles to boost algorithmic efficiency, theoretical clarity, and experimental robustness.

Across the cited literature, “PST hierarchy” does not denote a single canonical object. It names several unrelated hierarchical constructions: a strict parameter-age segmentation of a single network in continual learning (Du et al., 2019), the subgroup-property framework that characterizes soluble PST-groups in finite group theory (Yi et al., 2013), a rooted-tree enrichment of zero-dimensional persistence pairs (Rieck et al., 2019), an intermediate semidefinite-programming hierarchy for entanglement detection (Pena et al., 7 Aug 2025), and Pearl’s three-layer observational–interventional–counterfactual hierarchy as formalized in potential-outcomes notation (Wu et al., 28 Jan 2026). What unifies these usages is not a shared formalism but the role of hierarchy itself: each framework imposes an order relation that structures inference, computation, or classification.

1. Terminological scope

The abbreviation PST is polysemous. In the materials considered here, it refers to hierarchies over network parameters, subgroup properties, persistence pairs, semidefinite relaxations, and causal estimands. The resulting objects differ in algebraic type, but each organizes a family of entities by a monotone relation.

Domain PST hierarchy Ordering principle
Continual learning Progressive segmented training Older task segments are frozen; newer tasks use remaining capacity
Finite group theory PST-group characterizations Stronger and weaker subgroup properties; Hall-type criteria
Topological data analysis Rooted tree on persistence pairs Ancestry induced by nested branch-merges
Quantum information Partial-PPT extension hierarchy EXTkPSTkDPSkEXT_k \supseteq PST_k \supseteq DPS_k
Causal inference P–S–T hierarchy Observational \to interventional \to counterfactual

A common source of confusion is to treat these as variants of one theory. They are not. The continual-learning usage concerns parameter allocation in a neural network; the group-theoretic usage concerns transitivity of SS-permutability; the topological usage augments persistence diagrams with branch structure; the quantum-information usage defines a cone between EXT and DPS; and the causal-inference usage stratifies estimands by the type of potential-outcomes object required.

2. Progressive segmented training as a hierarchy of network capacity

In continual learning, PST denotes “Progressive Segmented Training,” a single-network method that builds a strict hierarchy of non-overlapping parameter segments across tasks (Du et al., 2019). Just before learning task TiT_i, the parameter set is decomposed as

Θ=(Θfixed,Θfree),\Theta = (\Theta_{\text{fixed}}, \Theta_{\text{free}}),

where Θfixed\Theta_{\text{fixed}} is frozen to preserve previously learned tasks and Θfree\Theta_{\text{free}} is available for adaptation to TiT_i. Training uses new data X(i)X^{(i)} together with a fixed-size memory \to0 containing up to \to1 examples from previous tasks. The training loss is a memory-assisted classification loss on a balanced batch drawn equally from \to2 and \to3, with only \to4 updated.

After training the free parameters on task \to5, PST scores each elementary unit by a first-order Taylor criterion. For a convolutional filter,

\to6

and for a fully connected neuron,

\to7

Within each layer, units are ranked by score, and the top \to8 are selected as \to9; the remainder is \to0.

The hierarchy is produced by the segmented update rules. PST re-initializes the selected important weights, retrains only those weights for a small number of epochs, freezes the reinforced result by moving it into \to1, and releases \to2 as the new free pool for the next task. Over time this yields

\to3

with \to4 containing all parameters important to tasks \to5. The supplied description states that this produces a strict hierarchy by task age: older tasks sit deeper in the frozen portion of the network, while newer tasks progressively claim a small slice of the capacity.

The fixed-size memory buffer is explicitly part of the stability–plasticity mechanism. Memory is uniformly sampled over previously seen classes, and replay is injected in three stages: initial epochs, periodic replay, and final classifier fine-tuning. This design is used without additional regularization. Empirically, the method achieves state-of-the-art accuracy in the single-head evaluation on CIFAR-10 and CIFAR-100, and the shrinking free pool yields a computational advantage: on CIFAR-100, PST reports a \to6 reduction in weight-update FLOPs and an \to7 reduction in total train-time FLOPs by the final tasks. The important conceptual point is that the hierarchy is not architectural growth; it is progressive freezing inside a single network.

3. The PST-group hierarchy in finite group theory

In finite group theory, a PST-group is a finite group in which \to8-permutability is transitive (Yi et al., 2013). Formally, if \to9, with SS0 SS1-permutable in SS2 and SS3 SS4-permutable in SS5, then SS6 is SS7-permutable in SS8. The classical characterization cited in the source states that a finite soluble group SS9 is a PST-group if and only if

TiT_i0

where TiT_i1 is an abelian Hall subgroup of TiT_i2, and every TiT_i3 induces on TiT_i4 only power automorphisms. In particular, TiT_i5 and TiT_i6 splits over TiT_i7.

The 2013 paper refines this theory via quasipermutability and TiT_i8-quasipermutability. If TiT_i9, then Θ=(Θfixed,Θfree),\Theta = (\Theta_{\text{fixed}}, \Theta_{\text{free}}),0 is quasipermutable if it permutes with Θ=(Θfixed,Θfree),\Theta = (\Theta_{\text{fixed}}, \Theta_{\text{free}}),1 and with every subgroup Θ=(Θfixed,Θfree),\Theta = (\Theta_{\text{fixed}}, \Theta_{\text{free}}),2 such that Θ=(Θfixed,Θfree),\Theta = (\Theta_{\text{fixed}}, \Theta_{\text{free}}),3. It is Θ=(Θfixed,Θfree),\Theta = (\Theta_{\text{fixed}}, \Theta_{\text{free}}),4-quasipermutable if it permutes with Θ=(Θfixed,Θfree),\Theta = (\Theta_{\text{fixed}}, \Theta_{\text{free}}),5 and with every Sylow subgroup Θ=(Θfixed,Θfree),\Theta = (\Theta_{\text{fixed}}, \Theta_{\text{free}}),6 such that Θ=(Θfixed,Θfree),\Theta = (\Theta_{\text{fixed}}, \Theta_{\text{free}}),7. These notions sit strictly below classical permutability conditions: requiring permutation with all subgroups recovers quasinormality, while requiring permutation with all Sylow subgroups recovers Θ=(Θfixed,Θfree),\Theta = (\Theta_{\text{fixed}}, \Theta_{\text{free}}),8-permutability.

The principal Hall-type characterization is Theorem A. Let Θ=(Θfixed,Θfree),\Theta = (\Theta_{\text{fixed}}, \Theta_{\text{free}}),9 and Θfixed\Theta_{\text{fixed}}0. Then the following are equivalent: Θfixed\Theta_{\text{fixed}}1 is a Hall subgroup and every Hall subgroup of Θfixed\Theta_{\text{fixed}}2 is quasipermutable; Θfixed\Theta_{\text{fixed}}3 is a soluble PST-group; every subgroup of Θfixed\Theta_{\text{fixed}}4 is quasipermutable; and every Θfixed\Theta_{\text{fixed}}5-subgroup together with some minimal supplement of Θfixed\Theta_{\text{fixed}}6 is quasipermutable. This equivalence is the key sense in which hierarchy enters the theory: a Hall-level criterion already determines the global PST property.

The paper also places PST-groups inside a broader hierarchy of subgroup properties:

Θfixed\Theta_{\text{fixed}}7

with the parallel chain

Θfixed\Theta_{\text{fixed}}8

On the global scale, the text records

Θfixed\Theta_{\text{fixed}}9

A common misconception is to conflate quasipermutability with quasinormality or Θfree\Theta_{\text{free}}0-permutability. The source explicitly distinguishes them: the “quasi” condition is relative to a complement Θfree\Theta_{\text{free}}1 of Θfree\Theta_{\text{free}}2 and only concerns coprime-order subgroups or Sylow subgroups of Θfree\Theta_{\text{free}}3.

4. The rooted-tree hierarchy on persistence pairs

In topological data analysis, the PST hierarchy is a rooted-tree structure on zero-dimensional persistence pairs, introduced to capture spatial relations that ordinary persistence diagrams do not encode (Rieck et al., 2019). The paper itself calls the construction the Interlevel-Set Persistence Hierarchy, and the supplied exposition identifies it as a zero-dimensional Persistence-Space-Tree hierarchy. Its motivation is the absence of spatial relationships between features in persistence diagrams, which limits expressive power.

Let Θfree\Theta_{\text{free}}4 be a Morse-type function with distinct critical values, and let

Θfree\Theta_{\text{free}}5

A zero-dimensional persistence pair is a pair of critical points Θfree\Theta_{\text{free}}6 where Θfree\Theta_{\text{free}}7 is a local minimum with birth time Θfree\Theta_{\text{free}}8 and Θfree\Theta_{\text{free}}9 is the critical point at which that component merges into an older one, with death time TiT_i0. The persistence is

TiT_i1

The global minimum yields a root pair TiT_i2.

The hierarchy places one node for each pair TiT_i3 and defines parent–child relations through branch-merges. The supplied definition uses the condition that two birth points remain in distinct connected components of an interlevel set

TiT_i4

which distinguishes a genuine branch-merge from a mere extension of an existing branch. The induced order

TiT_i5

makes the hierarchy a rooted tree with exactly one root and no node with more than one parent.

Algorithmically, the construction specializes union-find to persistence computation. Critical events are processed in ascending order, components are labeled by their highest minima, and when components merge one tests whether the corresponding minima remain separated in the interlevel set. If they do, a child edge is added in the hierarchy. This augments the usual persistence pairing with branching information.

Two derived quantities make the structure useful beyond visualization. First, the rank of a node TiT_i6 is the number of descendants:

TiT_i7

Second, the supplied exposition defines edge and node stability measures based on the birth and death coordinates of adjacent pairs. These quantities are intended to measure how stable the pairing and its branch structure are under perturbation. The conceptual significance is that the hierarchy records which features are nested inside others, thereby encoding branching structure that plain persistence diagrams omit.

5. The partial-PPT extension hierarchy for entanglement detection

In quantum information, PST denotes a new semidefinite-programming hierarchy for entanglement detection, positioned strictly between the EXT and DPS hierarchies (Pena et al., 7 Aug 2025). If TiT_i8 denotes the cone of separable operators on a bipartite space TiT_i9, the hierarchy is defined through symmetric extensions:

X(i)X^{(i)}0

X(i)X^{(i)}1

The intermediate cone is

X(i)X^{(i)}2

Hence, for each X(i)X^{(i)}3,

X(i)X^{(i)}4

and in the limit,

X(i)X^{(i)}5

A central contribution is a polynomial-size description of EXT and PST via a partition matrix X(i)X^{(i)}6 onto the symmetric subspace and a lifting-down operator X(i)X^{(i)}7. The action of X(i)X^{(i)}8 and its adjoint can be implemented in X(i)X^{(i)}9 flops rather than in dimension \to00, which is the key compression step. The paper states that these compact descriptions satisfy the Slater condition for both primal and dual SDPs, ensuring strong duality and the absence of duality gaps.

The hierarchy is accompanied by tailored algorithms. Three first-order methods are developed from least-squares formulations—Frank–Wolfe, projected gradient, and fast projected gradient—and a custom primal–dual interior-point method is derived from a conic formulation. The complexity statement is explicit. If \to01 and \to02, each primal–dual interior-point iteration for EXT\to03 or PST\to04 costs

\to05

flops and converges in

\to06

iterations for accuracy \to07. Each first-order iteration costs

\to08

flops and yields a duality gap of \to09 or \to10 after \to11 iterations.

The numerical position of PST is also specific. The paper reports that EXT\to12 or EXT\to13 with first-order methods quickly detects strongly entangled states, while PST\to14 plus first-order methods extends detection to weaker entanglement. The implementations scale to levels \to15 in minutes, compared with previous limits \to16–\to17, and the custom interior-point method is faster than off-the-shelf SDP solvers such as MOSEK via PICOS. The essential correction to a possible misunderstanding is that PST is neither a replacement for EXT nor for DPS; it is an intermediate cone that is tighter than EXT and cheaper than DPS.

6. The P–S–T causal hierarchy

In causal inference, PST refers to Pearl’s three-layer causal hierarchy, recast in potential-outcomes notation as P, S, and T (Wu et al., 28 Jan 2026). The layers are defined by the class of objects required to express a query. Let \to18 denote treatment, \to19 pre-treatment covariates, and \to20 the outcome, with potential outcomes \to21 and consistency \to22.

Layer P is observational or “seeing.” It contains only the joint and conditional distributions of observed variables:

\to23

Queries at this level remain within the observed \to24-algebra, such as \to25.

Layer S is interventional or “doing.” It contains marginal distributions of potential outcomes under single interventions:

\to26

and

\to27

Typical estimands include the average treatment effect \to28, the conditional ATE, quantile treatment effects, dose–response functions, and the controlled direct effect \to29.

Layer T is counterfactual or “imagining.” It contains joint or cross-world features of potential outcomes, nested counterfactuals, and individual-level contrasts:

\to30

Representative estimands include the probability of necessary causation, probability of sufficient causation, treatment benefit and harm rates, persuasion rate, natural direct and indirect effects, principal causal effects, and the individual treatment effect. An important technical distinction in the source is that controlled direct effects lie in layer S, whereas natural direct and indirect effects lie in layer T because they involve nested counterfactuals.

The hierarchy is also a hierarchy of identifiability difficulty. For layer P, no causal assumptions are needed. For layer S, the principal obstacle is confounding, and the standard assumptions are ignorability, consistency, and overlap:

\to31

Under these assumptions one obtains the g-formula,

\to32

For layer T, marginal laws are no longer sufficient because the dependence between \to33 and \to34 is not observed. The source therefore lists stronger assumptions or structures such as cross-world independence, monotonicity, copula models, rank preservation, a fully specified structural causal model, and sequential ignorability for mediation.

Identification strategies mirror the hierarchy. Back-door adjustment, front-door adjustment, instrumental variables, negative controls, proximal inference, data fusion, and sensitivity analysis are listed for layer S. For layer T, the paper emphasizes strong copula or independence assumptions, monotonicity, specification of an association parameter between \to35 and \to36, data fusion across multiple experiments, and partial-identification bounds. For individual-level counterfactuals, it highlights abduction–action–prediction, rank-preservation or quantile matching, and conformal inference. The stated overarching principle is that higher layers correspond to progressively richer features of the potential outcomes distribution and therefore require stronger assumptions for identification.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to PST Hierarchy.