---
title: 'SPO-codes: Contextual Variants & Applications'
url: https://www.emergentmind.com/topics/spo-codes
type: topic
---

# SPO-codes: Contextual Variants & Applications

“SPO-codes” is a context-dependent label rather than a single uniformly standardized object. In the most explicit usage, it denotes **suffix–prefix-codes with overlap**, a class of codes with overlapping code words introduced in symbolic dynamics [2509.04122]. In other recent literatures, the same label is used for implementation-oriented descriptions of **Standardized Preference Optimization** in x-to-audio alignment [2601.02900], for combinatorial encodings associated with the orthosymplectic Lie superalgebra \(\mathfrak{spo}(4|1)\) [2604.19511], for explicit rule systems governing \(\mathfrak{spo}(2|2)\)-equivariant quantization on the supercircle [1302.3727], for SpO\(_2\)-related signal encodings in biomedical prediction [2407.20989, 2607.07996], and for **Smart Predict-then-Optimize** mechanisms in decision-focused portfolio learning [2605.01176]. The term therefore requires disambiguation by field, mathematical object, and intended implementation.

## 1. Terminological scope and disambiguation

The following usages are distinct.

| Context | Meaning of “SPO-codes” | Primary source |
|---|---|---|
| Symbolic dynamics | Suffix–prefix-codes with overlap | [2509.04122] |
| X-to-audio evaluation | Standardized Preference Optimization code path for CLAPScore training | [2601.02900] |
| \(\mathfrak{spo}(4|1)\) representation theory | Integer- and tableau-based encodings of Verma basis vectors | [2604.19511] |
| \(\mathfrak{spo}(2|2)\) quantization | Explicit recursive correction rules for equivariant quantization | [1302.3727] |
| Biomedical learning | SpO\(_2\)-related continuous or threshold encodings, and predictor-guided reconstruction | [2407.20989], [2607.07996] |
| Portfolio optimization | SPO\(^+\)-based Smart Predict-then-Optimize decision signals | [2605.01176] |

A recurring source of confusion is that the same three letters expand differently across these settings. In symbolic dynamics, “SPO” is **suffix–prefix–overlap**. In x-to-audio alignment, “SPO” is **Standardized Preference Optimization**. In portfolio optimization, “SPO” is **Smart Predict-then-Optimize** through the SPO\(^+\) surrogate. In the orthosymplectic papers, “spo” refers to the Lie superalgebras \(\mathfrak{spo}(4|1)\) and \(\mathfrak{spo}(2|2)\), and the phrase “SPO-codes” is best understood as a convenient description of the associated combinatorial or recursive encodings rather than as a named code family.

A common misconception is that “SPO-codes” always refers to the symbolic-dynamics construction. The literature does not support that reading. Only the symbolic-dynamics paper formally introduces a class of codes with that exact name, whereas the other usages are contextual and tied to specific implementation or representation schemes.

## 2. Suffix–prefix-codes with overlap in symbolic dynamics

In symbolic dynamics, an SPO-code is defined over a finite alphabet \(\Sigma\) relative to a bifix code \(\mathcal F \subset \Sigma^+\). For each word \(c \in \mathcal C_{\mathcal F}\), there is a proper prefix \(f^-(c) \in \mathcal F\) and a proper suffix \(f^+(c) \in \mathcal F\). Writing
\[
c = \mathring{c}\, f^+(c),
\]
the concatenation with overlap is
\[
a \circledast b = \mathring{a}\, b.
\]
This operation identifies the distinguished suffix of \(a\) with the distinguished prefix of \(b\), so the overlap appears only once. The paper defines an SPO-code as a code \(\mathcal C\) contained in \(\mathcal C_{\mathcal F}\) for some bifix code \(\mathcal F\), and shows that \(a \circledast b \in \mathcal C_{\mathcal F}\) with
\[
f^-(a \circledast b)=f^-(a), \qquad f^+(a \circledast b)=f^+(b)
\]
[2509.04122].

The associated concatenation set \(C(\mathcal C)\) consists of bi-infinite sequences \(x \in \Sigma^{\mathbb Z}\) admitting a bi-infinite sequence of indices \((j_k,i_k)\) such that
\[
i_{k-1}<j_k\le i_k,
\]
and
\[
x_{(i_{k-1},i_k]} \in \mathcal C, \qquad
x_{(j_k,i_k]} = f^+\bigl(x_{(i_{k-1},i_k]}\bigr)=f^-\bigl(x_{(i_k,i_{k+1}]}\bigr).
\]
The generated subshift is
\[
X(\mathcal C)=\overline{C(\mathcal C)}.
\]
An SPO-code is **unambiguous** if the normalized index sequence is unique for each \(x \in C(\mathcal C)\). This is the overlap analogue of the usual uniqueness condition for coded systems.

The paper relates SPO-codes to Keller’s Markov codes by constructing, from a suitable unambiguous SPO-code \(\mathcal C\), a modified code \(\widehat{\mathcal C}\) and a Markov code \(\mathcal D_{\mathcal C}\) whose concatenation set is Borel conjugate to that of \(\widehat{\mathcal C}\). The construction uses
\[
\mathcal C^\bullet_{\mathcal F}
=
\{c \in \mathcal C_{\mathcal F} : \ell(c)\le \ell(f^-(c))+\ell(f^+(c))\},
\]
then groups overlap products until a word in \(\mathcal C^\bullet_{\mathcal F}\) is reached. This yields a countable-state topological Markov shift that models the SPO-coded system.

The ergodic consequence is an intrinsic ergodicity criterion. If \(\mathcal C\) has irreducible transition matrix, satisfies \(\mathcal C \cap \mathcal C^\bullet_{\mathcal F}\neq \emptyset\), obeys
\[
h(C(\mathcal C)) > h(X(\mathcal C)\setminus C(\mathcal C)),
\]
and \(X(\mathcal C)\) has a measure of maximal entropy of full support, then \(X(\mathcal C)\) is intrinsically ergodic [2509.04122]. This is the main route by which SPO-codes connect overlap combinatorics to thermodynamic properties.

The same paper shows that every synchronized subshift admits a canonical SPO-code. For a synchronized subshift \(X\), the code \(\mathcal C(X)\) is built from segments between successive changes of the last synchronizing left endpoint. Its concatenation set is \(B(X)\), and the construction is unambiguous and invariant under topological conjugacy. A conjugacy-invariant condition
\[
\sup_{c\in\mathcal C(X)}
\{
\ell(c)-\ell(f^-(c))-\ell(f^+(c))
\}
=\infty
\]
appears as Condition (H), and together with
\[
h(B(X)) > h(X\setminus B(X))
\]
and existence of a full-support measure of maximal entropy, it implies intrinsic ergodicity for topologically transitive synchronized subshifts.

The symbolic-dynamics paper also uses SPO-codes as a construction tool. It produces synchronized subshifts whose Markov boundary consists of finitely many orbits or countably many orbits, and it constructs SPO-coded systems that are not semisynchronized. In this literature, therefore, SPO-codes are neither merely a coding trick nor a notational convenience; they are the central structural object.

## 3. Standardized Preference Optimization as “SPO-codes” in x-to-audio alignment

In the XACLE Challenge submission “SPO-CLAPScore,” “SPO-codes” refers to implementation-oriented documentation for **Standardized Preference Optimization** in a CLAPScore-based x-to-audio alignment predictor [2601.02900]. The model predicts an alignment score from CLAP-style embeddings by cosine similarity, scaled to \(0\)–\(10\):
\[
\hat{x} = 10 \cdot \cos(e^{\textsf{audio}}, e^{\textsf{text}}),
\]
with the audio encoder given by M2D-CLAP 2025 and the text encoder by BERT base. The text encoder is frozen, while the audio encoder is fine-tuned.

The SPO step standardizes each listener’s raw score \(x\) into a listener-wise z-score
\[
x_{\text{spo}} = \frac{x-\mu_{\text{listener}}}{\sigma_{\text{listener}}},
\]
where \(\mu_{\text{listener}}\) and \(\sigma_{\text{listener}}\) are computed over all items rated by that listener. The central motivation is that each audio–text pair is rated by only four listeners, each listener uses an idiosyncratic numeric scale, and direct regression on raw MOS encourages fitting listener-dependent bias. After standardization, \(x_{\text{spo}}>0\) indicates a score above that listener’s personal mean, and \(x_{\text{spo}}<0\) indicates a score below it.

Predictions are also standardized, but with **global** training statistics:
\[
\hat{z}=\frac{\hat{x}-\mu_{\text{train}}}{\sigma_{\text{train}}},
\]
and training minimizes an MSE regression loss plus an optional contrastive loss from UTMOS:
\[
L
=
L_{\text{reg}}(x_{\text{spo}},\hat{z})
+
\lambda L_{\text{con}}(x_{\text{spo}},\hat{z}).
\]
The reported setting uses \(\lambda=0.5\) when the contrastive term is enabled. The method is therefore not pairwise ranking in the usual sense and not DPO-style preference optimization; it is per-listener z-scoring combined with supervised learning on the standardized targets.

The data processing pipeline includes **listener screening**. For a given raw score \(x\), if none of the other scores for that item lies in \([x-\tau,x+\tau]\), the score is marked as an NG-Score. A listener is removed if the proportion of NG-Scores exceeds \(r\). The paper uses \(\tau=5\) and \(r=0.2\), reducing training scores from \(30{,}000\) to \(29{,}308\) and validation scores from \(12{,}000\) to \(11{,}747\). Three settings are trained: A without listener screening and without contrastive loss; B with screening and with contrastive loss; and C with screening but without contrastive loss. Each setting is trained with and without warm-up, with three random seeds per warm-up option, for \(18\) models total, ensembled by averaging predictions.

Experimentally, the ensemble achieved \(6\)th place in the challenge with SRCC \(0.6142\), against an official baseline SRCC of \(0.3345\). On validation, a single model with SPO reached SRCC \(0.6367\), LCC \(0.6572\), KTAU \(0.4606\), and MSE \(3.256\), whereas the same architecture without SPO yielded SRCC \(0.5408\), LCC \(0.5514\), KTAU \(0.3837\), and MSE \(6.709\) [2601.02900]. In this usage, “SPO-codes” denotes the code path implementing listener screening, per-listener standardization, global standardization of predictions, and the standardized loss.

## 4. Orthosymplectic encodings: Verma bases and equivariant quantization

In the representation-theoretic paper on \(\mathfrak{spo}(4|1)\), “SPO-codes” is best understood as the combinatorial encoding of basis vectors in finite-dimensional irreducible modules by integer exponent patterns and Kashiwara–Nakashima tableaux [2604.19511]. The highest weight is written
\[
\lambda
=
m_1\omega_1 + 2m_2\omega_2
=
(m_1+m_2)\epsilon_1 + m_2\epsilon_2,
\qquad
m_1,m_2\in\mathbb Z_{\ge 0},
\]
with corresponding irreducible module \(L(\lambda)\). Using the negative simple root vectors
\[
f_1 = E_{21}-E_{34}, \qquad f_2 = E_{52}-E_{45},
\]
the Verma vectors are monomials
\[
f_1^{b_4} f_2^{b_3} f_1^{b_2} f_2^{b_1} v_\lambda
\]
subject to
\[
0 \le b_1 \le 2m_2,\qquad
0 \le b_2 \le m_1+b_1,\qquad
0 \le b_3 \le \min\{b_2+m_1,2b_2\},\qquad
0 \le b_4 \le \min\{m_1,\tfrac12 b_3\}.
\]
These inequalities define the Verma vector system \(H\).

The same basis vectors are encoded by KN tableaux of shape \(\lambda=(m_1+m_2,m_2)\) over the alphabet
\[
\mathcal N = \{1,2,0,\overline{2},\overline{1}\},
\qquad
1<2<0<\overline{2}<\overline{1}.
\]
Each tableau \(T\) gives a four-integer code \((b_1,b_2,b_3,b_4)\) by counting specific occurrences of \(0,\overline{2},\overline{1}\) in the two rows. The map
\[
\psi:\mathrm{KN}_\lambda(4|1)\to H,\qquad
T\mapsto f_1^{b_4} f_2^{b_3} f_1^{b_2} f_2^{b_1} v_\lambda
\]
is a bijection, and the weight of the tableau matches the weight of the corresponding Verma vector:
\[
\mathrm{wt}
=
(m_1+m_2-b_2-b_4)\epsilon_1
+
(m_2-b_1+b_2-b_3+b_4)\epsilon_2.
\]
Because the Verma vectors expand triangularly in a tableau-indexed monomial basis of
\[
W = V^{\otimes m_1}\otimes (\wedge^2 V)^{\otimes m_2},
\]
they are linearly independent, and since \(|H|=|\mathrm{KN}_\lambda(4|1)|=\dim L(\lambda)\), they form a basis. In this setting, the “code” is the equivalence among exponent tuples, tableaux, and basis vectors.

A related but different orthosymplectic usage appears in the paper on \(\mathfrak{spo}(2|2)\)-equivariant quantization on the supercircle \(S^{1|2}\) [1302.3727]. There the core objects are weighted densities \(\mathcal F_\lambda\), differential operators \(\mathcal D_{\lambda\mu}\), and the associated graded symbol space
\[
\mathcal S_\delta = \bigoplus_{k\in \frac{\mathbb N}{2}} \mathcal S_\delta^k,
\qquad
\delta=\mu-\lambda.
\]
The supercircle carries the standard contact structure generated by
\[
\bar D_1=\partial_{\theta_1}-\theta_1\partial_x,
\qquad
\bar D_2=\partial_{\theta_2}-\theta_2\partial_x,
\]
which induces the contact filtration in half-integer order. The central result is the existence and uniqueness, for non-critical \(\delta\), of an \(\mathfrak{spo}(2|2)\)-equivariant quantization
\[
Q:\mathcal S_\delta \to \mathcal D_{\lambda\mu}
\]
preserving principal symbols.

The construction is Casimir-based. On \(\mathcal S_\delta^k\), the Casimir eigenvalue is
\[
\alpha_k=
\begin{cases}
(-k+\delta)^2, & k\in\mathbb N,\\[2mm]
(-k+\delta)^2-\frac14, & k\in \frac12+\mathbb N.
\end{cases}
\]
Non-critical values are those for which \(\alpha_k\neq \alpha_l\) for all \(l<k\). The explicit formulas for the lower-order corrections \(S_{k-l}\) involve iterated \(\partial_x\), \(\bar D_1\), \(\bar D_2\), the coefficients \(A_k=-k\), \(B_k=-(k+2\lambda)\), and denominators built from \(\alpha_k-\alpha_{k-i}\). In this literature, “SPO-codes” denotes the explicit recursive rule system by which lower-order symbol components are corrected so that the resulting differential operator respects \(\mathfrak{spo}(2|2)\)-symmetry.

## 5. SpO\(_2\)-related encodings in biomedical learning

In biomedical applications, “SPO-codes” refers to SpO\(_2\)-related labels or reconstructions rather than symbolic codes. One study treats SpO\(_2\) as a continuous scalar target \(y\in[0,100]\) and as a thresholded binary label derived from the rule
\[
y=1 \text{ if } \mathrm{SpO}_2 \le 92\%, \qquad
y=0 \text{ if } \mathrm{SpO}_2 > 92\%
\]
[2407.20989]. The models are pretrained audio neural networks—CNN6, CNN10, CNN14, and Audio-MAE—operating on 4-second speech segments resampled to \(16\) kHz and represented as spectrograms or log-mel spectrograms. SpO\(_2\) regression is trained with MSE and evaluated by RMSE, MAE, \(R^2\), and Pearson correlation. The reported RMSE values are \(4.4\pm0.8\) for Audio-MAE, \(4.8\pm1.0\) for CNN6, \(4.5\pm0.7\) for CNN10, and \(4.7\pm0.8\) for CNN14, all exceeding the accepted clinical range of \(3.5\%\). Pearson \(r\) never exceeds \(0.3\). Binary SpO\(_2\) threshold classification reaches F1-scores of \(0.621\pm0.052\) for Audio-MAE, \(0.641\pm0.041\) for CNN6, \(0.615\pm0.067\) for CNN10, and \(0.643\pm0.045\) for CNN14, while direct respiratory-insufficiency detection with the same architectures achieves near-perfect accuracy. The paper interprets this as a separation of domains: speech is highly informative for global RI status but weak for exact SpO\(_2\).

A more direct SpO\(_2\)-estimation paper uses “SPO-codes” as implementation details for a **SpO\(_2\) predictor-guided stage-wise time-frequency reconstruction** framework applied to low-quality dual-wavelength PPG [2607.07996]. Inputs are AC/DC-normalized red and infrared PPG segments of length \(10\) s with \(1\)-s stride, sampled at \(100\) Hz, so each segment is \(\mathbf x_i\in\mathbb R^{1000\times 2}\). The system has two learned components: a Bi-LSTM plus attention predictor \(P(\cdot)\) producing a scalar SpO\(_2\) estimate, and a 4-layer Transformer-encoder reconstructor \(R(\cdot)\) operating on masked dual-channel PPG.

Training is divided into four stages. Stage 1 pretrains \(P\) on high-quality segments selected by Orphanidou-style signal quality assessment implemented in NeuroKit2, requiring all 1-second averaged quality scores on the red channel to be at least \(0.6\). Stage 2 freezes \(P\) and trains \(R\) with random contiguous masks of length \(1\)–\(5\) s using a joint loss
\[
\mathcal L_{\mathrm{recon}}
=
\mathcal L_{\mathrm{time}}
+
\lambda_{\mathrm{freq}}\mathcal L_{\mathrm{freq}}
+
\lambda_{\mathrm{SpO}_2}\mathcal L_{\mathrm{SpO}_2}^{\mathrm{guide}}.
\]
Here \(\mathcal L_{\mathrm{time}}\) is masked-region MSE in the time domain, \(\mathcal L_{\mathrm{freq}}\) is an STFT-domain MSE using FFT size \(200\), window length \(200\), and hop \(20\), and \(\mathcal L_{\mathrm{SpO}_2}^{\mathrm{guide}}\) feeds a merged original-plus-reconstructed segment into the frozen predictor \(P\). Stage 3 freezes \(R\), masks the worst \(3\)-s quality region in each segment, reconstructs it, and refines \(P\) on the reconstructed inputs. Stage 4 freezes the refined \(P\) and retrains \(R\) again on high-quality segments with the same joint loss.

The reported subject-level MAE on OpenOximetry is \(2.882\%\), improving over calibration (\(3.460\)), NormWear (\(3.390\)), and a direct Bi-LSTM+attention baseline (\(3.063\)). On a private wearable dataset, the framework achieves subject-level MAE \(2.359\%\) and RMSE \(2.973\%\). Ablations show that the full loss \(\mathcal L_{\mathrm{time}}+\mathcal L_{\mathrm{freq}}+\mathcal L_{\mathrm{SpO}_2}^{\mathrm{guide}}\) outperforms variants missing either the frequency or the predictor-guided term. In this biomedical setting, “SPO-codes” denotes continuous SpO\(_2\) labels, threshold codes, or predictor-guided reconstruction pathways rather than a discrete code family.

## 6. SPO in decision-focused portfolio optimization

In portfolio optimization, “SPO” denotes **Smart Predict-then-Optimize** through the SPO\(^+\) surrogate, not a code in the symbolic-dynamics sense [2605.01176]. The setting is decision-focused learning for portfolio allocation. A predictor outputs asset returns
\[
\hat{\bm r}_t = \bm\Theta \bm x_t + \bm b,
\]
and these predictions are passed to a mean–variance optimizer with transaction cost:
\[
\bm w_t^*
\in
\arg\max_{\bm w}
\left\{
\hat{\bm r}_t^\top \bm w
-
\lambda \bm w^\top \bm\Sigma_t \bm w
-
\kappa \|\bm w-\bm w_{t-1}\|_1
\right\},
\]
subject to
\[
\bm 1^\top \bm w = 1,
\qquad
\bm w \ge 0.
\]
Standard predict-then-optimize would train the predictor with an error such as
\[
\ell_{\mathrm{PtO}}(\hat{\bm c},\bm c)=\|\hat{\bm c}-\bm c\|_2^2,
\]
whereas decision-focused learning minimizes a decision loss tied to the downstream optimizer. Since exact regret is non-convex and non-differentiable, SPO\(^+\) provides the surrogate used in training.

The paper’s main theoretical point is KKT-based. For the mean–variance formulation, stationarity gives
\[
\hat r_i - 2\lambda (\bm\Sigma \bm w)_i - \kappa s_i - \nu + \mu_i = 0,
\]
with \(s_i \in \partial |w_i-w_{t-1,i}|\). Defining the risk- and cost-adjusted marginal score
\[
m_i = \hat r_i - 2\lambda (\bm\Sigma \bm w)_i - \kappa s_i,
\]
active assets satisfy \(m_i=\nu\) and inactive assets satisfy \(m_i\le \nu\). The optimizer therefore behaves like a ranking system over adjusted marginal scores. This explains why SPO-trained predictors can inflate return magnitudes: the learning objective is not calibration of \(\hat{\bm r}_t\), but production of scores that induce better downstream allocations.

Empirically, the paper reports prediction inflation and excessive turnover. Standard SPO-trained portfolios show average monthly turnover around \(80\%\)–\(95\%\) across DOW, ETF\_A, and ETF\_B, and increasing the risk-aversion parameter \(\lambda\) from \(0.1\) to \(50\) hardly reduces turnover. To stabilize the system, the paper evaluates three mechanisms. **Prediction clipping** replaces \(\hat{\bm r}_t\) by
\[
\tilde{\bm r}_t = \mathrm{clip}(\hat{\bm r}_t,-\gamma,\gamma),
\]
with \(\gamma=0.1\). **Min–max rescaling** maps predictions to \([-c,c]\) with \(c=0.1\), preserving cross-sectional ranking. **Partial portfolio adjustment** updates
\[
\bm w_t
=
\bm w_{t-1}
+
\delta(\bm w_t^*-\bm w_{t-1}),
\]
with \(\delta=0.1\), starting from the equally weighted portfolio. Adjustment alone reduces turnover to about \(9\%\)–\(10\%\) monthly, while Clip+Adj yields the lowest turnover and volatility, and Rescale+Adj often gives the highest returns and Sharpe ratios among the SPO variants [2605.01176].

Taken together, these results show that in portfolio learning the relevant “SPO-code” is the optimizer-facing signal produced by SPO\(^+\)-trained prediction models. Its defining property is not discrete combinatorics but decision sensitivity: output values matter insofar as they alter the optimizer’s ranking and threshold structure.

Source: https://www.emergentmind.com/topics/spo-codes