---
title: Combined Alignment Score (CAS) Overview
url: https://www.emergentmind.com/topics/combined-alignment-score-cas
type: topic
---

# Combined Alignment Score (CAS) Overview

Combined Alignment Score (CAS) denotes, in the broadest sense, a scalar intended to summarize alignment quality by combining multiple alignment signals, error channels, or normalization factors. Across the cited arXiv literature, however, the term is not used uniformly. Several alignment papers do not define CAS at all and instead evaluate alignment with task-specific quantities such as note-level \(F_{\text{align}}\), Sum-of-Pairs (SP), beat-accuracy thresholds, or human-calibrated scalar ratings [2507.12175] [1708.01508] [2601.02900]. Other papers introduce explicit combined alignment-style metrics, such as the Pose Alignment Score (PAS) or the Clinical Alignment Score, while still not using CAS to mean “Combined Alignment Score” [2407.20391] [2605.12650]. This suggests that CAS is best understood as an umbrella notion for combined alignment evaluation rather than as a single standardized metric.

## 1. Terminological status in the literature

In several recent alignment papers, CAS is explicitly absent. The music-analysis framework RUMAA does not define or report CAS; its closest alignment metric is repeat-aware note-level \(F_{\text{align}}\) with a \(\pm 50\) ms onset tolerance and an adaptation in which repeated notes are counted independently [2507.12175]. The multiple-sequence method PoMSA likewise does not define CAS; its reported “alignment score” is SP score on BAliBASE, OXBench, and SMART [1708.01508]. The audio–text system SPO-CLAPScore predicts a scalar alignment score calibrated to human ratings, but does not introduce a metric called CAS [2601.02900]. The diffusion post-training method AGSM also does not define CAS, instead using an intrinsic alignment reward, a Plackett–Luce alignment probability, and several separate evaluation metrics such as ImageReward, CLIP, PickScore, HPSv2, and GenEval submetrics [2605.30038].

The acronym itself is also overloaded. In CAS-IQA, “CAS” stands for Contrast-free Angiography Synthesis rather than any form of combined alignment score, and the paper predicts three separate quality targets—Vessel Morphology Consistency (VMC), Vessel Branch Detection (VBD), and Overall Quality (OQ) [2505.17619]. In CRAFT, by contrast, CAS is an explicit metric name, but it means Clinical Alignment Score rather than Combined Alignment Score [2605.12650]. This corpus-level pattern suggests that one must distinguish carefully between a genuine combined metric, a task-specific scalar alignment score, and an acronym collision.

## 2. Aggregated alignment evidence in neural word alignment

One of the clearest CAS-like constructions appears in neural word alignment through score aggregation. For a sentence pair \((\mathbf e,\mathbf f)\), the model defines a local source–target score
\[
s(i,j)=\mathrm{net}_e([\mathbf e]_i^{d^e_{\mathrm{win}}})\cdot \mathrm{net}_f([\mathbf f]_j^{d^f_{\mathrm{win}}}),
\]
then combines all source-side scores for target word \(e_i\) into an aggregated matching score
\[
s_{\mathrm{aggr}}(i,\mathbf f)=\mathrm{aggr}_{j=1}^{|\mathbf f|} s(i,j).
\]
The paper studies three aggregation operators—Sum, Max, and LogSumExp (LSE)—with
\[
s_{\mathrm{aggr}}(i,\mathbf f)=\frac{1}{r}\log\left(\sum_{j=1}^{|\mathbf f|} e^{r\,s(i,j)}\right),
\]
and reports that LSE performs best on all datasets, using \(r=1\) in experiments [1606.09560].

This aggregated score is used only during unsupervised training. The soft-margin objective is
\[
{\cal L}(\mathbf e^+,\mathbf e^-,\mathbf f)
=
\sum_{i^+=1}^{|\mathbf e^+|}
\log\!\left(1+e^{-s_{\mathrm{aggr}}(i^+,\mathbf f)}\right)
+
\sum_{i^-=1}^{|\mathbf e^-|}
\log\!\left(1+e^{+s_{\mathrm{aggr}}(i^-,\mathbf f)}\right).
\]
At inference time, aggregation is removed; decoding reverts to the pairwise score \(s(i,j)\), followed by thresholding
\[
s(i,j) > \mu^{-}(e_i)+\alpha\,\sigma^{-}(e_i),
\]
and bidirectional symmetrization. In CAS terms, this is a combined training-side score rather than a final evaluation metric.

The paper’s empirical role for aggregation is explicit. LSE substantially outperforms Max and Sum in AER on all tested language pairs, and the resulting model improves over Fast Align by \(7\) AER on English–Czech, \(6\) AER on Romanian–English, and \(1.7\) AER on English–French [1606.09560]. A plausible implication is that, when CAS is interpreted as a combined latent-alignment score, the most relevant design choice is not merely what is aggregated, but how softly evidence from multiple candidate alignments is fused.

## 3. Explicit combined metrics for pose alignment

The most explicit evaluation-side analogue of CAS in the cited corpus is the pose-alignment metric family composed of Translation Alignment Score (TAS), Rotation Alignment Score (RAS), and Pose Alignment Score (PAS). TAS is computed by first defining a geometric scale
\[
d=\underset{i\in\mathcal C}{\mathrm{Q3}}
\left(
\min_{j\in\mathcal C,\ j\neq i}
\|\mathbf c_i-\mathbf c_j\|
\right),
\]
then robustly aligning estimated camera centers to ground truth, computing cumulative frequencies over thresholds \(\{0.01d,0.02d,\dots,d\}\), and normalizing
\[
\mathrm{TAS}=\frac{1}{100n}\left(\sum_{k=1}^{100} f_k\right).
\]
RAS is defined analogously for angular thresholds \(\{0.1^\circ,0.2^\circ,\dots,10^\circ\}\),
\[
\mathrm{RAS}=\frac{1}{100n}\left(\sum_{k=1}^{100} f_k\right),
\]
and PAS is the arithmetic mean
\[
\mathrm{PAS}=\frac{\mathrm{TAS}+\mathrm{RAS}}{2}.
\]
The authors also note a weighted-average interpretation \(\alpha\,\mathrm{TAS}+(1-\alpha)\,\mathrm{RAS}\) with \(\alpha=0.5\) for PAS [2407.20391].

The construction is notable because translation and rotation are evaluated separately and only then fused. PAS is therefore not a joint \(SE(3)\) path cost, but a post hoc combined score. The paper argues that TAS and RAS are robust to outliers because they are based on robust alignment procedures and on normalized sums of cumulative frequencies rather than mean or RMS residuals. It also claims that TAS handles collinear motion better than mAA because it measures aligned pointwise distance errors rather than angular errors between camera pairs [2407.20391].

Within the cited literature, PAS is the cleanest example of a metric that functions as a combined alignment score without using that exact name. Its structure—separate primitive alignment scores, normalized to \([0,1]\), followed by equal-weight averaging—is a recurring template later echoed in clinical alignment scoring.

## 4. Normalized and structural composite scores in sequence alignment

In multiple sequence alignment, the paper on normalized MSA defines three explicit ratio criteria that are directly relevant to any formal CAS design:
\[
\gamma_1[A]=
\begin{cases}
0,& |A|=0\\
\gamma[A]/|A|,& \text{otherwise}
\end{cases},
\qquad
\gamma_2[A]=\sum_{h=1}^{k-1}\sum_{i=h+1}^{k}\frac{\gamma[A_{\{h,i\}]}{|A_{\{h,i\}|},
\]
and
\[
\gamma_3[A]=
\begin{cases}
0,& |A|=0\\
\gamma[A]\Big/\left(\sum_{h=1}^{k-1}\sum_{i=h+1}^{k}|A_{\{h,i\}|\right),& \text{otherwise}.
\end{cases}
\]
These criteria differ in the denominator: global alignment length for \(\gamma_1\), per-pair induced alignment lengths for \(\gamma_2\), and total induced pairwise alignment mass for \(\gamma_3\) [2107.01607].

The paper proves that \(NMSA\text{-}z\) is NP-complete for each \(z\in\{1,2,3\}\), gives exact dynamic-programming algorithms for all three, and provides a 12-approximation for \(NMSA\text{-}2\) under a restricted scoring-matrix class [2107.01607]. It also shows by example that the optimal alignment under one normalized criterion need not be optimal under another. This makes denominator choice substantive rather than cosmetic. If CAS is understood as a combined similarity–length objective, \(\gamma_2\) is the most direct sum-of-normalized-pairs instantiation; if it is understood as a single global ratio, \(\gamma_1\) and \(\gamma_3\) are simpler alternatives. This is an inference from the formal definitions rather than an explicit endorsement by the paper.

A different composite scheme appears in RNA secondary-structure alignment with coaxial helical stacking. CHSalign aligns ordered labeled trees whose nodes are helices, junctions, and hairpin loops. For junction nodes, the local score is
\[
\gamma(t_1[i],t_2[j])=
\begin{cases}
s+\frac{w}{2},& \text{if branches match and both junctions have no CHS}\\[4pt]
s+w,& \text{if branches match and the nonzero CHS motif matches}\\[4pt]
-\infty,& \text{otherwise},
\end{cases}
\]
with \(w=100\) fixed in experiments and values \(w>50\) reported to work well [1609.01987]. This score then enters a tree-structured dynamic program over subtrees and forests. CHSalign therefore supplies a raw composite alignment score built from local structural similarity, topology compatibility, and CHS agreement, and it is intended to yield a high score for RNAs with similar CHS motifs or helical arrangement patterns and a low score otherwise.

By contrast, PoMSA is an important counterexample. Despite repeated references to “higher alignment score,” the paper’s actual evaluation metric is only SP score, and it does not define CAS, TC score, Total Column score, or any benchmark-specific combined score [1708.01508]. This contrast underscores that “alignment score” in the literature is often generic unless a paper specifies the composition explicitly.

## 5. Music and audio alignment metrics closest to CAS

In score–performance alignment, RUMAA is explicit that CAS is absent. The model instead produces a strict one-to-one alignment through three synchronized streams: \(T1\) for score-aligned performance transcription, \(T2\) for performance-aligned score conversion, and \(T3\) for edit operation tagging, with Insert, Delete, Match, Repeat, and a skip token “-”. When a note exists in only one modality, the model emits the exclusive placeholder token \(\langle-\rangle\) in the opposite stream so that both streams remain perfectly aligned event by event [2507.12175]. This aligned notewise sequence is the model’s alignment substrate.

Evaluation is performed with note-level \(F_{\text{align}}\), defined in prose as counting matched note pairs and inserted/deleted notes as true positives, unmatched predicted notes as false positives, and missing ground-truth notes as false negatives, under a \(\pm 50\) ms onset tolerance. Because the original metric was designed for symbolic alignment tasks without repetitions, RUMAA redefines repeated notes to be counted independently [2507.12175]. On the revised Vienna dataset, RUMAA reports \(98.4\) \(F_{\text{align}}\) both without repeats and with repeats, whereas baselines on repeated scores collapse to the range \(12.7\)–\(36.4\). In CAS terms, the closest equivalent here is therefore repeat-aware note-level \(F_{\text{align}}\), not any path-based combined score.

A related structure-aware audio-to-score alignment paper likewise does not define CAS. It separates structural-difference detection from final alignment by predicting inflection points in a cross-similarity matrix and then using an extended DTW recurrence. Final performance is reported only as beat alignment accuracy under \(25\), \(50\), \(100\), and \(200\) ms tolerances [2102.00382]. On the Tido subset with structural differences, the best progressively dilated model, DCNN\(_{2+3}\), reaches \(73.9\), \(81.3\), \(85.6\), and \(92.8\) percent at those four thresholds. The paper is therefore informative for CAS design mainly because it distinguishes structural correctness from local timing correctness, even though it does not combine them into one published scalar.

A newer audio-to-score method that directly bridges audio-like and symbol-level features also omits CAS and instead reports mean error, median error, and thresholded success rates. Its proposed method achieves mean \(86\) ms, median \(21\) ms, \(83.7\%\) within \(50\) ms, \(91.7\%\) within \(100\) ms, \(95.2\%\) within \(200\) ms, and \(97.9\%\) within \(500\) ms, compared with an audio-to-audio baseline at mean \(135\) ms, median \(49\) ms, \(53.2\%\) within \(50\) ms, and \(91.7\%\) within \(500\) ms [2605.20014]. The paper also notes that several audio-to-audio alignments had to be excluded because of obviously spurious alignment, whereas the proposed method worked robustly across the full dataset. This suggests that any CAS for audio-to-score work would need to combine precision with robustness rather than rely on a single threshold alone.

## 6. Domain-specific CAS acronyms and clinically motivated composite scores

The sharpest acronym collision occurs in medical image quality assessment. CAS-IQA defines CAS as Contrast-free Angiography Synthesis, not Combined Alignment Score. The framework predicts three separate targets—VMC, VBD, and OQ—each annotated on a continuous \(0\!\sim\!100\) scale. Subject ratings are normalized by
\[
Z_{ij}=\frac{S_{ij}-\mu_i}{\sigma_i},
\]
rescaled to \([0,100]\), and averaged as
\[
MOS_j=\frac{1}{N}\sum_{i=1}^{N}\hat Z_{ij}.
\]
OQ is described as a holistic metric that balances VMC and VBD while also accounting for visual artifacts, but it is separately annotated and separately predicted rather than defined as a deterministic algebraic combination of the other two scores [2505.17619].

CRAFT, by contrast, introduces CAS explicitly as the Clinical Alignment Score. CAS is defined as the macro-average of four primitive scores—Visual Description Consistency (VDC), Clinical Criteria Satisfaction (CCS), Diagnostic Discriminability (DD), and Semantic Feature Similarity (SFS)—computed in the SigLIP evaluator embedding space [2605.12650]. The reward-side primitives are
\[
r_{\mathrm{vdc}}=\cos(E_I(\hat x_0),E_T(t_{ep})),\qquad
r_{\mathrm{ccs}}=\cos(E_I(\hat x_0),E_T(t_c)),
\]
\[
r_{\mathrm{dd}}=\log \frac{\exp(\phi(E_I(\hat x_0))_y)}{\sum_{k=1}^{C}\exp(\phi(E_I(\hat x_0))_k)},
\qquad
r_{\mathrm{sfs}}=\cos(E_I(\hat x_0),E_I(x_r)),
\]
and are averaged during training as
\[
r_{\mathrm{cam}}=\frac{1}{M}\sum_{m=1}^{M}\frac{1}{4}\left(r_{\mathrm{vdc}}^{(m)}+r_{\mathrm{ccs}}^{(m)}+r_{\mathrm{dd}}^{(m)}+r_{\mathrm{sfs}}^{(m)}\right).
\]
This is a fully explicit composite metric, but its name is Clinical Alignment Score, not Combined Alignment Score. The paper further analyzes the low-alignment tail using a dataset-specific threshold \(\tau\) equal to the 25th percentile of the real-image CAS distribution and reports \(5.5\)–\(34.7\) percentage point reductions relative to the strongest baseline, corresponding to a \(20.4\%\) average relative reduction across datasets [2605.12650].

Outside medical imaging, multimodal alignment papers often prefer scalar predictors or reward functions rather than named CAS metrics. SPO-CLAPScore predicts
\[
\hat x=\frac{e^{\mathrm{audio}}\cdot e^{\mathrm{text}}}{\|e^{\mathrm{audio}}\|\,\|e^{\mathrm{text}}\|}\times 10
\]
and trains against the listener-wise standardized target
\[
x_{\mathrm{spo}}=\frac{x-\mu_{\mathrm{listener}}}{\sigma_{\mathrm{listener}}},
\]
combining MSE with a contrastive term [2601.02900]. AGSM defines an internal alignment probability
\[
p(z=1\mid x_t,c)=\frac{\exp(r(x_t,c))}{\sum_i \exp(r(x_t,c^i))}
\]
and a modified score-matching target
\[
\nabla\log \tilde p_t(x_t\mid c,z)=\nabla\log p_t(x_t\mid c)+\gamma_z \nabla\log p(z=1\mid x_t,c),
\]
but evaluates alignment through multiple separate benchmarks rather than a single universal score [2605.30038]. Taken together, these works reinforce a general pattern: a combined alignment score becomes meaningful only when the paper specifies exactly which primitives are combined, how they are normalized, and whether the resulting scalar is used for training, evaluation, or both.

Source: https://www.emergentmind.com/topics/combined-alignment-score-cas