Papers
Topics
Authors
Recent
Search
2000 character limit reached

SPSS: Sequential Permutation for Selecting Sparsity

Updated 14 July 2026
  • The paper introduces SPSS as a sequential permutation test mechanism to select optimal sparsity levels in sparse conditional cross-covariance reduction for reliable variable selection.
  • SPSS employs two-way iterative thresholding on cross-covariance differences to derive sparse singular vectors that effectively capture sex-specific associations in TMJ studies.
  • The method outperforms alternatives by selecting biologically meaningful skull and muscle variables with high true positive rates and minimal false positives, validated through resampling.

Searching arXiv for the primary paper and closely related references. {"5query5 (Park et al., 30 Sep 2025) Sequential Permutation for Selecting Sparsity conditional cross-covariance reduction5", "5max_results5 5} {"5query5 Reduction for Characterizing Sexual Dimorphism in Biomechanics of the Temporomandibular Joint5\5 "5max_results5 5arXiv (Park et al., 30 Sep 2025) Sequential Permutation for Selecting Sparsity conditional cross-covariance reduction5query5} {"5query5 "5max_results5 5arXiv (Park et al., 30 Sep 2025) Sequential Permutation for Selecting Sparsity conditional cross-covariance reduction5query5} {"5query5 sparse filter estimation permutation problem", "5max_results5 5} Sequential Permutation for Selecting Sparsity (SPSS) is a data-driven procedure introduced within the conditional cross-covariance reduction (CCR) model to choose the sparsity levels in a sparse singular-vector representation of a cross-covariance contrast that varies with a conditioning variable (&&&5query5&&&). In the temporomandibular joint (TMJ) study where the conditioning variable is binary sex, SPSS is used to decide how many skull variables and how many muscle-attachment variables should be retained so as to capture the part of the skull–muscle association whose dependence structure differs between males and females. The method was proposed because standard approaches such as cross-validation or penalization are not ideal for the cadaver TMJ dataset, where the sample size is extremely small and variable selection is unstable (&&&5query5&&&).

CCR is designed to examine the dynamic association between two sets of random variables conditioned on a third variable. Its central object is the conditional cross-covariance matrix

PRESERVED_PLACEHOLDER_5query5^

which is assumed to lie in low-dimensional subspaces,

PRESERVED_PLACEHOLDER_5arXiv (Park et al., 30 Sep 2025) Sequential Permutation for Selecting Sparsity conditional cross-covariance reduction5^

for semi-orthogonal basis matrices PRESERVED_PLACEHOLDER_5max_results5^ and PRESERVED_PLACEHOLDER_5query5, and a latent matrix-valued function PRESERVED_PLACEHOLDER_5\5^ (&&&5query5&&&).

When ZZ is binary, the entire dynamic association is summarized by

ΣXY(1)ΣXY(2),\Sigma_{XY}(1)-\Sigma_{XY}(2),

and the model reduces to finding its singular vectors. If

ΣXY(1)ΣXY(2)=UDV,\Sigma_{XY}(1)-\Sigma_{XY}(2)=UDV^{\top},

then UU and VV define the two subspaces that best capture sex-specific differences in covariance. The sparse version of this decomposition is the interpretable target, because the nonzero rows of the estimated singular vectors determine which variables in PRESERVED_PLACEHOLDER_5arXiv (Park et al., 30 Sep 2025) Sequential Permutation for Selecting Sparsity conditional cross-covariance reduction5query5^ and PRESERVED_PLACEHOLDER_5arXiv (Park et al., 30 Sep 2025) Sequential Permutation for Selecting Sparsity conditional cross-covariance reduction5arXiv (Park et al., 30 Sep 2025) Sequential Permutation for Selecting Sparsity conditional cross-covariance reduction5^ participate in the sex-differential association (&&&5query5&&&).

In this framework, SPSS does not construct the CCR model itself. Rather, it provides the practical rule for choosing the sparsity levels PRESERVED_PLACEHOLDER_5arXiv (Park et al., 30 Sep 2025) Sequential Permutation for Selecting Sparsity conditional cross-covariance reduction5max_results5^ and PRESERVED_PLACEHOLDER_5arXiv (Park et al., 30 Sep 2025) Sequential Permutation for Selecting Sparsity conditional cross-covariance reduction5query5^ required by sparse CCR. A plausible implication is that SPSS is the tuning mechanism that turns sparse CCR from a low-rank contrastive covariance model into an operational variable-selection procedure for very small samples.

5max_results5. Sparse CCR formulation and the role of sparsity

The estimation procedure first computes the sample cross-covariance difference

PRESERVED_PLACEHOLDER_5arXiv (Park et al., 30 Sep 2025) Sequential Permutation for Selecting Sparsity conditional cross-covariance reduction5\5^

after centering within sex group (&&&5query5&&&). The sparse subspaces are then obtained by maximizing

PRESERVED_PLACEHOLDER_5arXiv (Park et al., 30 Sep 2025) Sequential Permutation for Selecting Sparsity conditional cross-covariance reduction55^

which is equivalently written as

PRESERVED_PLACEHOLDER_5arXiv (Park et al., 30 Sep 2025) Sequential Permutation for Selecting Sparsity conditional cross-covariance reduction56

Without sparsity, the solution is given by the left and right singular vectors of PRESERVED_PLACEHOLDER_5arXiv (Park et al., 30 Sep 2025) Sequential Permutation for Selecting Sparsity conditional cross-covariance reduction57. The paper instead introduces a sparse SVD-like algorithm because interpretability and small-sample stability are crucial (&&&5query5&&&). Hard thresholding is used instead of shrinkage penalties, which helps reduce the bias often induced by penalized estimators. In effect, sparse SVD makes the low-rank contrastive covariance problem directly interpretable: the selected rows identify the skull and muscle variables entering the sex-differential association.

The paper also defines covariance and correlation contrasts along the extracted directions. For rank PRESERVED_PLACEHOLDER_5arXiv (Park et al., 30 Sep 2025) Sequential Permutation for Selecting Sparsity conditional cross-covariance reduction58,

PRESERVED_PLACEHOLDER_5arXiv (Park et al., 30 Sep 2025) Sequential Permutation for Selecting Sparsity conditional cross-covariance reduction59

The covariance difference PRESERVED_PLACEHOLDER_5max_results5query5^ is the quantity optimized by the CCR model, while PRESERVED_PLACEHOLDER_5max_results5arXiv (Park et al., 30 Sep 2025) Sequential Permutation for Selecting Sparsity conditional cross-covariance reduction5^ is used to interpret the resulting subspaces in correlation scale. The authors stress that the model is built to maximize covariance differences, not directly correlation differences, because optimizing correlation contrast is much harder (&&&5query5&&&).

5query5. Two-way iterative thresholding and the SPSS procedure

Sparse estimation is implemented by two-way iterative thresholding. The algorithm starts from the top-PRESERVED_PLACEHOLDER_5max_results5max_results5^ singular vectors of PRESERVED_PLACEHOLDER_5max_results5query5, then alternates between left and right updates (&&&5query5&&&). At iteration PRESERVED_PLACEHOLDER_5max_results5\5, the left update computes

PRESERVED_PLACEHOLDER_5max_results55^

then keeps only the PRESERVED_PLACEHOLDER_5max_results56 rows of PRESERVED_PLACEHOLDER_5max_results57 with largest row norms and sets the others to zero. After left orthonormalization via QR, the right update computes

PRESERVED_PLACEHOLDER_5max_results58

then keeps only the PRESERVED_PLACEHOLDER_5max_results59 rows with largest row norms, followed by right orthonormalization via QR. Iteration continues until convergence, measured by the change in the projection matrices,

PRESERVED_PLACEHOLDER_5query5query5^

The output is the sparse pair PRESERVED_PLACEHOLDER_5query5arXiv (Park et al., 30 Sep 2025) Sequential Permutation for Selecting Sparsity conditional cross-covariance reduction5, together with

PRESERVED_PLACEHOLDER_5query5max_results5^

SPSS is the rule that selects the optimal sparsity levels PRESERVED_PLACEHOLDER_5query5query5^ and PRESERVED_PLACEHOLDER_5query5\5^ sequentially. Its logic is to start with a small sparsity, compare it to the next larger sparsity, and increase the sparsity only if the additional variables produce a statistically significant improvement in the estimated contrast (&&&5query5&&&). For PRESERVED_PLACEHOLDER_5query55, the hypotheses are

PRESERVED_PLACEHOLDER_5query56

for PRESERVED_PLACEHOLDER_5query57 and fixed PRESERVED_PLACEHOLDER_5query58. The same idea is then used symmetrically to select PRESERVED_PLACEHOLDER_5query59.

The test is implemented with a leave-two-out (LTO) resampling scheme. One observation from each sex group is omitted, the CCR model is refit on the remaining PRESERVED_PLACEHOLDER_5\5query5^ samples, and the difference in the estimated leading contrast is recorded across all PRESERVED_PLACEHOLDER_5\5arXiv (Park et al., 30 Sep 2025) Sequential Permutation for Selecting Sparsity conditional cross-covariance reduction5^ LTO splits: PRESERVED_PLACEHOLDER_5\5max_results5^ The observed mean difference is

PRESERVED_PLACEHOLDER_5\5query5^

To generate the null distribution, the signs of the PRESERVED_PLACEHOLDER_5\5\5^ values are randomly flipped many times; the TMJ application uses 5arXiv (Park et al., 30 Sep 2025) Sequential Permutation for Selecting Sparsity conditional cross-covariance reduction5query5query5,5query5query5query5^ permutations (&&&5query5&&&). The p-value PRESERVED_PLACEHOLDER_5\55^ is the proportion of permuted mean differences at least as large as the observed mean difference.

The sequential stopping rule is explicit. One increases PRESERVED_PLACEHOLDER_5\56 from PRESERVED_PLACEHOLDER_5\57 to PRESERVED_PLACEHOLDER_5\58 as long as all tested PRESERVED_PLACEHOLDER_5\59 are below 5query5.5query55 and chooses ZZ5query5^ as the smallest ZZ5arXiv (Park et al., 30 Sep 2025) Sequential Permutation for Selecting Sparsity conditional cross-covariance reduction5^ such that at least one ZZ5max_results5. An analogous rule determines ZZ5query5. In practice, SPSS is therefore a stepwise permutation test on increments in explained covariance contrast (&&&5query5&&&).

5\5. TMJ cadaver application

The method is applied to cadaver data with

ZZ5\5^

after centering each variable within sex (&&&5query5&&&). The study uses data from 5arXiv (Park et al., 30 Sep 2025) Sequential Permutation for Selecting Sparsity conditional cross-covariance reduction5query5^ male and 5arXiv (Park et al., 30 Sep 2025) Sequential Permutation for Selecting Sparsity conditional cross-covariance reduction5arXiv (Park et al., 30 Sep 2025) Sequential Permutation for Selecting Sparsity conditional cross-covariance reduction5^ female cadaver heads to investigate sex-specific relationships between craniofacial skeletal morphology and TMJ-related masticatory muscle attachments.

SPSS produces boxplots of the p-values over all ZZ5 comparisons. From these, the authors conclude that the increment from 5 to 6 variables is not significant, so they select

ZZ6

(&&&5query5&&&). Using these sparsity levels, the CCR model identifies the linear combinations

ZZ7

and

ZZ8

The maximal covariance difference is reported as ZZ9, and the associated correlation difference is ΣXY(1)ΣXY(2),\Sigma_{XY}(1)-\Sigma_{XY}(2),5query5^ (&&&5query5&&&). The resulting scatterplot shows that the sex-specific linear combinations have different correlation patterns by group. The combined score ΣXY(1)ΣXY(2),\Sigma_{XY}(1)-\Sigma_{XY}(2),5arXiv (Park et al., 30 Sep 2025) Sequential Permutation for Selecting Sparsity conditional cross-covariance reduction5^ is reported as positively correlated for males and negatively correlated for females, with correlations roughly ΣXY(1)ΣXY(2),\Sigma_{XY}(1)-\Sigma_{XY}(2),5max_results5^ for males and ΣXY(1)ΣXY(2),\Sigma_{XY}(1)-\Sigma_{XY}(2),5query5^ for females.

The selected skull variables include features such as bicondylar width (ΣXY(1)ΣXY(2),\Sigma_{XY}(1)-\Sigma_{XY}(2),5\5), bigonial width (ΣXY(1)ΣXY(2),\Sigma_{XY}(1)-\Sigma_{XY}(2),5), and mandibular length, which align with prior anthropological evidence about sex differences in mandibular geometry. The selected TO variables are all related to muscle attachment size. The authors argue that the CCR model reveals an association hidden in raw pairwise plots: within females, shorter mandibular length can correspond to larger temporalis origin attachment size when other selected skull variables are held fixed (&&&5query5&&&).

5. Empirical behavior, comparative results, and interpretation

The main empirical finding is that SPSS successfully identifies the true sparsity level in simulations when signal strength is moderate to strong (&&&5query5&&&). In the same simulation settings, CCR with sparse SVD recovers the correct variables and subspaces with very high true positive rates and essentially zero false positives. Compared with competing methods, CCR achieves the smallest subspace distances and no false positives in the simulation benchmark.

The comparison reported in the paper distinguishes several alternatives. GLAA also captures the dynamic association but is penalized and therefore more biased. BCCA performs less favorably in this setting and does not force exact zeros in the loadings. RGCCA fails to clearly separate the group differences (&&&5query5&&&). These comparisons situate SPSS as part of a broader sparse multivariate estimation problem rather than as an isolated testing device.

In the TMJ data, the selected sparsity ΣXY(1)ΣXY(2),\Sigma_{XY}(1)-\Sigma_{XY}(2),6 yields an interpretable sex-differential association with biologically plausible variables (&&&5query5&&&). The biomechanical interpretation is linked to joint reaction force (JRF): shorter mandibular length can increase JRF, suggesting potentially higher TMJ loading and a possible risk subgroup for temporomandibular disorder. This suggests that SPSS contributes not only to variable selection, but also to the extraction of a clinically interpretable covariance pattern whose sign differs by sex.

6. Limitations, open problems, and terminological distinction

A key advantage of SPSS is that it is well suited to tiny samples because it uses resampling and sequential testing instead of unstable tuning-based selection (&&&5query5&&&). At the same time, the paper notes a limitation: SPSS delivers a single decision about sparsity rather than a full uncertainty quantification. Future work could add permutation-based confidence intervals or other resampling uncertainty measures. The authors also note that the rank ΣXY(1)ΣXY(2),\Sigma_{XY}(1)-\Sigma_{XY}(2),7 is treated as pre-specified and not estimated, which remains an open problem.

The term “permutation” in SPSS refers to the sign-permutation test used to assess whether increasing sparsity significantly improves the leading covariance-difference statistic. This should be distinguished from the “permutation problem” in sparse filter estimation, where permutations refer to frequency-wise source permutations in convolutive blind source separation (&&&5max_results5&&&). That earlier line of work also uses sparsity as a criterion, but it addresses recovery of the correct frequency permutations of estimated filters rather than sparsity selection in a sparse singular-vector decomposition. A plausible implication is that the shared vocabulary reflects a common reliance on sparsity and permutation-based reasoning, while the underlying statistical objects, optimization targets, and inferential goals are different.

Within CCR, SPSS is the paper’s practical solution for choosing the sparsity levels required by the model: it uses leave-two-out perturbations and sign-permutation tests to decide when adding variables no longer significantly increases the sex-specific covariance contrast, and in the TMJ study this leads to a sparse, interpretable, and biomechanically meaningful sex-dimorphism finding (&&&5query5&&&).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Sequential Permutation for Selecting Sparsity (SPSS).