---
title: 'UniSino: CT Sinogram Standardization'
url: https://www.emergentmind.com/topics/unisino
type: topic
---

# UniSino: CT Sinogram Standardization

Searching arXiv for the specified paper and closely related context.
UniSino is a physics-driven foundational model for universal CT sinogram standardization that operates directly in the CT projection domain rather than the reconstructed image domain [2508.17816]. It is designed to transform a degraded sinogram $x_{\mathrm{corr}}(s,\theta)$ into a clean, physically consistent sinogram $x_{\mathrm{std}}(s,\theta)$ that matches the distribution and physics of fully sampled, noise-free projections, with the goal of mitigating heterogeneous degradations such as undersampling and noise before they are amplified by reconstruction [2508.17816]. In the formulation reported for UniSino, this projection-domain strategy is motivated by the observation that sinograms exhibit more uniform distributions across defect types than reconstructed images, and that correction at the raw-data stage improves downstream reconstruction quality in both single and mixed undersampling scenarios [2508.17816].

## 1. Concept and problem formulation

UniSino addresses degradation in CT raw data arising from non-standard scanning protocols and non-ideal reconstruction conditions [2508.17816]. The degradation modes listed for the framework include sparse-view, limited-angle, low-dose, detector downsampling, ring artifacts from detector channels, geometric miscalibration, truncation, metal, and motion [2508.17816]. In the problem setting described for the model, these degradations act on the sinogram and subsequently propagate through the nonlinear reconstruction pipeline, producing severe artifacts such as streaks, rings, structural collapse, and high-frequency distortions in reconstructed images [2508.17816].

The term “sinogram standardization” is used in a specific sense: transforming degraded raw projection data into a physically consistent form aligned with fully sampled, noise-free projections [2508.17816]. This distinguishes UniSino from conventional correction procedures, which are described as artifact-specific, calibration-dependent, and limited in portability, and from image-domain foundational models, which must disentangle artifact-specific clusters in a highly nonuniform image-space distribution [2508.17816]. The reported rationale is that a universal projection-domain model can generalize across heterogeneous degradations, including mixed artifacts, because the source-space statistics are more amenable to unified modeling [2508.17816].

A common misconception is to treat UniSino as merely a denoising or sparse-view completion model. The description instead positions it as a universal standardization framework spanning multiple subtasks and mixed degradation settings, with training explicitly structured to handle heterogeneous artifact types rather than a single isolated corruption mode [2508.17816].

## 2. Physical basis and mathematical structure

UniSino is grounded in standard CT forward physics, beginning with the Beer–Lambert law for x-ray transmission [2508.17816]:

$$
I = I_0 \exp\!\left(-\int_{\mathcal{L}} \mu(l)\,dl\right),
$$

where $\mu(l)$ is the linear attenuation coefficient along path $\mathcal{L}$, $I$ is the detected intensity, and $I_0$ is the incident intensity [2508.17816]. After logarithmic transformation, the projection value becomes a line integral,

$$
p = -\ln(I/I_0) = \int_{\mathcal{L}} \mu(l)\,dl,
$$

which is the basis of the sinogram representation used by the model [2508.17816].

The framework further adopts the Radon-transform view of projection data [2508.17816]:

$$
Rf(s,\theta) = \int_{\mathbb{R}^2} f(\mathbf{x})\,\delta\!\big(s - \mathbf{x}\cdot\mathbf{n}_\theta\big)\,d\mathbf{x},
$$

with unit normal $\mathbf{n}_\theta = (\cos\theta, \sin\theta)$ [2508.17816]. Under low-dose acquisition, photon counts are modeled as Poisson random variables. If $I(s,\theta) = I_0 \exp(-x_0(s,\theta))$ are expected counts, then the measured counts satisfy

$$
Y(s,\theta) \sim \operatorname{Poisson}(I(s,\theta)),
$$

and the log-domain measurement is

$$
\hat{x}(s,\theta) = -\ln\!\left(\frac{Y(s,\theta)}{I_0}\right),
$$

yielding heteroscedastic noise in the projection domain [2508.17816]. In the reported training pipeline, low-dose artifacts are simulated precisely by sampling $Y$ from $\operatorname{Poisson}(I_0 e^{-x_0})$ and then applying the log transform [2508.17816].

Rather than imposing the full Helgason–Ludwig moment conditions, UniSino uses two practical projection-domain consistency constraints through SinoLoss: cross-view mean consistency and angular visibility constraints [2508.17816]. Their concrete forms are given as

$$
L_{\mathrm{mean}} = \sum_{\theta}\Big|\frac{1}{S}\sum_{s} \hat{x}(s,\theta) - \bar{m}\Big|,
$$

and

$$
L_{\mathrm{vis}} = \sum_{\theta}\|(1-M_\theta)\odot \hat{x}(\cdot,\theta)\|_1,
$$

where $\bar{m}$ is the cross-angle mean derived from observed angles and $M_\theta$ is a zero–one visibility mask inferred through intersection, backprojection, and forward projection [2508.17816]. SinoLoss is described as combining these constraints with bounded-variation behavior to discourage nonphysical oscillations [2508.17816].

This suggests that UniSino’s “physics-driven” characterization is not limited to using a CT forward model in data simulation; it also embeds physically motivated regularity conditions directly into the training objective.

## 3. Architecture: SinoVAE and latent refinement diffusion

UniSino comprises two tightly coupled modules that both operate in the projection domain: SinoVAE and Latent Refinement Diffusion (LRD) [2508.17816]. SinoVAE is described as a Sinogram Variation Autoencoder with perceptual compression, dual pathways, and physics-driven constraints [2508.17816]. Its encoder $E$ produces the mean and variance of a latent Gaussian,

$$
(\mu,\sigma) = E(x), \qquad
z = \mu + \sigma \odot \varepsilon,\quad \varepsilon\sim\mathcal{N}(0,I),
$$

with KL regularization used to structure the latent space [2508.17816].

The decoder stage is explicitly dual-path [2508.17816]. Decoder 1 reconstructs full-frequency content, while Decoder N emphasizes artifact-sensitive high-frequency structure such as streaks, rings, and undersampling edges, with high-frequency guidance computed using Sobel gradients:

$$
G_x =
\begin{bmatrix}
1&0&-1\\
2&0&-2\\
1&0&-1
\end{bmatrix}*x,\qquad
G_y =
\begin{bmatrix}
1&2&1\\
0&0&0\\
-1&-2&-1
\end{bmatrix}*x,\qquad
G = \sqrt{G_x^2 + G_y^2}.
$$

Discriminator-based adversarial heads are attached to both decoders to sharpen realism and suppress artifacts [2508.17816]. The latent embedding is therefore decomposed into a full-frequency pathway preserving global structure and intensity, and a high-frequency pathway encoding artifact-sensitive details [2508.17816].

The second module, LRD, is a conditional diffusion model operating in latent space to refine SinoVAE latents of undersampled sinograms [2508.17816]. Its forward process over $T$ steps is

$$
q(z_t|z_{t-1}) = \mathcal{N}\big(z_t;\sqrt{1-\beta_t}\,z_{t-1},\,\beta_t I\big),\quad
q(z_t|z_0) = \mathcal{N}\big(z_t;\sqrt{\bar{\alpha}_t}\,z_0,\,(1-\bar{\alpha}_t)I\big),
$$

with $\alpha_t = 1 - \beta_t$ and $\bar{\alpha}_t=\prod_{i=1}^{t}\alpha_i$ [2508.17816]. The reverse denoising process conditioned on the corrupted-sinogram latent is

$$
z_{t-1} =
\frac{1}{\sqrt{\alpha_t}}\Big(z_t -
\frac{1-\alpha_t}{\sqrt{1-\bar{\alpha}_t}}\,\varepsilon_\phi(z_t,t,c)\Big)
+ \sigma_t \xi,\quad \xi\sim\mathcal{N}(0,I),
$$

where $\varepsilon_\phi$ predicts the added noise and $c$ is the conditioning latent [2508.17816]. Conditioning uses both full-frequency and high-frequency latent channels so that restoration remains faithful to the corrupted input while suppressing artifacts [2508.17816].

A plausible implication is that the division of labor between SinoVAE and LRD is central to the model’s reported efficiency: compression and feature structuring occur before diffusion, rather than performing unconditional generation directly in full-resolution sinogram space.

## 4. Training procedure and artifact simulation

UniSino is trained in two stages using self-supervised data generation from clean sinograms [2508.17816]. In Stage I, SinoVAE is trained on standardized or clean sinograms with simulated degraded counterparts [2508.17816]. The objective is

$$
L_{\mathrm{SinoVAE}} =
\underbrace{\|D(z)-S\|_2^2}_{L_{\mathrm{rec}}}
+ \alpha\,\underbrace{\mathrm{LPIPS}(D(z),S)}_{L_{\mathrm{per}}}
+ \beta\,\underbrace{D_{\mathrm{KL}}(q(z|S)\,\|\,p(z))}_{L_{\mathrm{KL}}}
+ \lambda_{\mathrm{adv}}\,\underbrace{L_{\mathrm{adv}}(D(z))}_{\text{adversarial}}
+ \lambda_{\mathrm{phys}}\,\underbrace{(L_{\mathrm{mean}}+L_{\mathrm{vis}})}_{\text{SinoLoss}}.
$$

Here $S$ denotes the ground-truth sinogram and $D$ the decoder [2508.17816]. In Stage II, LRD is trained in latent space with conditioning from the undersampled-sinogram latent $Z_{\mathrm{cond}} = E(S_{\mathrm{corr}})$ using

$$
L_{\mathrm{diffusion}} =
\mathbb{E}_{t,z_t,\varepsilon}\big[\|\varepsilon - \varepsilon_\phi(z_t,t,Z_{\mathrm{cond}})\|_2^2\big].
$$

After $T$ diffusion steps, the refined latent $z_0$ is decoded by the SinoVAE global decoder [2508.17816].

The degradation simulation protocol covers multiple artifact modes [2508.17816]. Sparse-view corruption is modeled by removing angle subsets, limited-angle by restricting the angular range, truncation by cropping the detector axis and setting values outside the field of view to nominal values, downsampling by sub-sampling detector channels, ring artifacts by adding channel-specific offsets, geometry artifacts by angle-dependent detector shifts, metal artifacts by modifying projected high-attenuation regions, and motion artifacts by forward projecting a deformed object [2508.17816]. Mixed degradation is created by random degradation mixing during training, which is explicitly reported as a mechanism for robustness to co-occurring artifacts [2508.17816].

Training was conducted primarily on the NLST lung dataset with 300 patients and 800,112 slices, of which 790,112 were used for training and 10,000 for testing; sinogram size was $384\times384$ [2508.17816]. The implementation used PyTorch with Adam optimizer on two NVIDIA RTX 4090 GPUs, with learning rates of $10^{-6}$ for SinoVAE and $10^{-4}$ for LRD [2508.17816]. SinoVAE and LRD were trained sequentially, with the encoder frozen during diffusion training [2508.17816].

## 5. Datasets, tasks, and empirical performance

Although UniSino is described as being validated across eight CT datasets, the core benchmarking is reported on four representative datasets: NLST, CQ500, LIDC-IDRI, and KiTS19 [2508.17816]. Additional generalization tests use CHAOS, QIN LUNG, LiTS, and MSD Colon [2508.17816]. The subtasks include sparse-view completion, limited-angle restoration, low-dose denoising, downsample restoration, ring artifact correction, truncation correction, geometry correction, and mixed-artifact restoration [2508.17816]. Evaluation uses projection-domain PSNR, SSIM, and NRMSE on standardized sinograms, with image-domain metrics obtained after applying FBP to the standardized sinograms [2508.17816].

The overall projection-domain comparison reported in Table I gives UniSino PSNR 45.522, SSIM 96.272, and NRMSE 0.00607 [2508.17816]. The corresponding values reported for CycleGAN are 27.398, 77.046, and 0.08221; for ViT, 27.881, 70.363, and 0.06632; for U-Net, 39.764, 91.087, and 0.01820; and for DDPM, 39.613, 95.432, and 0.01385 [2508.17816].

Per-task projection-domain results are also reported [2508.17816]:

| Subtask | PSNR | SSIM |
|---|---:|---:|
| SV | 48.938 | 97.331 |
| LA | 39.794 | 93.237 |
| LD | 49.420 | 97.305 |
| DS | 47.567 | 95.091 |
| RI | 49.266 | 97.290 |
| TR | 42.867 | 96.391 |
| GE | 40.800 | 97.262 |

The corresponding NRMSE values are reported as 0.00353 for SV, 0.01141 for LA, 0.00335 for LD, 0.00373 for DS, 0.00341 for RI, 0.00794 for TR, and 0.00912 for GE [2508.17816]. Cross-dataset generalization on NLST evaluation is reported as PSNR 46.19, SSIM 96.731, and NRMSE 0.00331, exceeding the reported results for CycleGAN, ViT, U-Net, and DDPM [2508.17816].

For mixed undersampling examples combining LD+RI+SV and TR+GE+LA, the reported PSNR values are 42.641 and 37.152, respectively [2508.17816]. The text attributes these outcomes to random mixing during training and projection-domain consistency, which together yield reliable normalization under compounded artifacts [2508.17816].

## 6. Ablation, deployment characteristics, and limitations

Ablation studies identify the physics-driven components and projection-domain design as critical to performance [2508.17816]. Replacing SinoVAE with a standard encoder–decoder reduces projection-domain performance to PSNR 43.647, SSIM 94.483, and NRMSE 0.00791 [2508.17816]. Removing SinoLoss yields PSNR 44., SSIM 95.612, and NRMSE 0.00654 [2508.17816]. Full UniSino is reported at PSNR 45.522, SSIM 96.272, and NRMSE 0.00607 [2508.17816]. The associated interpretation in the source is that explicit high-frequency preservation and physics-guided losses reduce artifact propagation after backprojection and enhance restoration fidelity [2508.17816].

The model is described as efficient at inference because SinoVAE compresses sinograms into compact latents and the diffusion stage operates conditionally in latent space rather than as unconditional full-resolution generation [2508.17816]. Exact throughput numbers are not reported, but the model is described as having substantially shorter reconstruction times than patch-based baselines and unconditional DDPM [2508.17816]. Standardized sinograms can be fed into standard FBP or iterative reconstruction, and no specialized post-processing is required [2508.17816].

The paper also introduces the SimSinoCT dataset for benchmarking sinogram standardization and releases code at the stated repository [2508.17816]. This reinforces the model’s role not only as a restoration architecture but as part of a broader evaluation framework for projection-domain standardization.

Reported failure modes include extreme 2D truncation outside inferred visibility masks and very large miscalibrations where simple geometric shifts are insufficient [2508.17816]. The text suggests potential benefits from explicit scanner geometry modeling in future work [2508.17816]. Other future directions listed are extending to 3D sinograms for volumetric CT, cross-modality adaptation to MRI k-space and PET sinograms, and multi-domain strategies combining standardized sinograms with image-domain constraints [2508.17816].

## 7. Position within CT reconstruction research

UniSino is positioned as a projection-domain foundational model, in contrast to foundational models that operate in image space [2508.17816]. The distinction is methodological as well as conceptual. Image-domain systems must address artifacts after reconstruction, whereas UniSino attempts to standardize the raw measurement representation before backprojection or iterative inversion [2508.17816]. In the reported framing, this avoids amplification of source-space inconsistencies and exploits the more uniform distribution of sinograms across degradation types [2508.17816].

Its broader significance lies in harmonizing raw CT data across protocols and scanners while reducing dependence on vendor-specific, artifact-specific preprocessing [2508.17816]. This suggests a shift in medical imaging foundation-model design from image-domain semantics toward raw-data modeling with embedded acquisition physics. A plausible implication is that UniSino belongs to a class of models that treat the measurement domain itself as the principal locus of generalization, rather than as a preliminary step to be discarded once reconstruction begins.

Within that framing, UniSino can be understood as a universal standardization system for CT sinograms whose core contributions are a physics-constrained projection-domain representation, a dual-path variational latent model, and latent diffusion refinement trained under mixed degradation regimes [2508.17816]. Its reported empirical performance, cross-dataset generalization, and ablation results collectively support the claim that projection-domain foundational modeling is a viable strategy for robust CT raw-data enhancement [2508.17816].

Source: https://www.emergentmind.com/topics/unisino