---
title: Omnidirectional Image Quality Assessment
url: https://www.emergentmind.com/topics/omnidirectional-image-quality-assessment-oiqa
type: topic
---

# Omnidirectional Image Quality Assessment

Omnidirectional Image Quality Assessment (OIQA) concerns the objective prediction of the perceptual quality of omnidirectional, or \(360^\circ\), images as they are experienced in immersive viewing systems rather than as ordinary planar pictures. The problem departs from conventional image quality assessment because omnidirectional content is usually stored in equirectangular projection (ERP), which introduces severe nonuniform distortions, especially near the poles, while human observers inspect only a viewport at any moment through a head-mounted display and shift attention through head and eye movements [2207.02674]. As a result, OIQA has developed around three coupled issues: spherical geometry, viewport-dependent perception, and the aggregation of local quality variations into a global quality judgment. The field now includes full-reference, reduced-reference, and especially no-reference formulations, with extensions to non-uniform distortions, AI-generated panoramas, stitching artifacts, super-resolution, and stereoscopic depth quality [2303.06907].

## 1. Perceptual and geometric foundations

OIQA is shaped by two constraints that are largely absent from conventional 2D IQA. First, ERP stretches pixels near latitude \(\pm 90^\circ\), so pixel-space fidelity does not correspond uniformly to angular fidelity on the sphere. Second, a viewer in VR observes only a limited field of view at any instant, and the perceived quality of an omnidirectional image depends on which regions are visited, in what order, and for how long [2207.02674].

Psychophysical work has shown that visual attention and retinal eccentricity are structurally important. In the retina-related zoning study, the visual field is partitioned into five concentric zones centered at the foveation point: fovea \(Z_1\), parafovea \(Z_2\), perifovea \(Z_3\), near periphery \(Z_4\), and far periphery \(Z_5\). The reported subjective results indicate that the fovea and parafovea are extremely important for quality perception, while the impacts of perifovea and periphery are small; fitted zone weights gave \(w_1 \sim 0.40\!-\!0.95\) and \(w_2 \sim 0.02\!-\!0.40\), with the remaining zones contributing less than \(25\%\) in total [1908.06239]. This establishes a formal basis for foveated weighting, saliency-guided sampling, and viewport prioritization.

Viewing conditions also alter perceived quality. In a psychophysical study that varied starting point and exploration time, starting point, distortion type, and their interaction with exploration time were highly significant, whereas exploration time alone was not; the same work reported a recency effect for localized stitching distortions [2005.10547]. Later work on generative scanpath representation made this dependence explicit by parameterizing viewing condition as \(\Omega=\{(y_1,x_1),T\}\), with starting point and exploration time determining a distribution of plausible scanpaths [2309.03472]. The combined evidence identifies OIQA as a spatiotemporal perceptual inference problem even for static panoramas.

## 2. Databases and subjective methodology

The empirical basis of OIQA is a sequence of increasingly specialized databases, beginning with globally distorted natural panoramas and expanding to non-uniform, AI-generated, and task-specific content.

| Database | Content and scale | Notable protocol or annotation |
|---|---|---|
| OIQA | 16 source images, 320 distorted images | VR subjective study with head and eye movement data [2207.02674] |
| CVIQD / CVIQ | 16 reference images, 528 distorted ERP images | Compression-focused benchmark with JPEG, H.264, H.265 [2303.06907] |
| JUFE-10K | 430 reference OIs, 10,320 non-uniformly distorted OIs | Psychophysical experiment under free exploration with eye/head data [2501.11511] |
| OIQ-10K | 10,000 omnidirectional images with homogeneous and heterogeneous distortions | MOS, distortion spatial distribution, and head/eye movements [2502.15271] |
| AIGCOIQA2024 | 300 AI-generated panoramas | Triple ratings for quality, comfortability, and correspondence [2404.01024] |
| OHF2024 | 600 AI-generated omnidirectional images | Triple MOS plus distortion-aware saliency annotation [2506.21925] |

The OIQA database established a canonical early protocol: 16 pristine panoramas degraded by JPEG, JPEG2000, Gaussian blur, and white Gaussian noise at five levels, yielding 320 distorted images, with subjective quality collected in a VR environment using HTC Vive and an eye tracker [2207.02674]. The accompanying head-only and head-eye saliency maps were intended to support saliency-weighted objective assessment.

Compression-oriented evaluation was standardized by CVIQD, described as 16 reference images with JPEG, H.264, and H.265 distortions generating 528 distorted ERP images [2303.06907]. These two datasets became the primary benchmarks for early no-reference OIQA, including VGCN, MFILGN, S\(^2\), ST360IQ, Assessor360, and related models [2002.09140].

Later databases shifted the field toward spatial heterogeneity. JUFE-10K contains 10,320 non-uniformly distorted omnidirectional images generated from 430 references by applying Gaussian noise, Gaussian blur, brightness discontinuity, and stitching distortion to one or two fisheye lenses before stitching [2501.11511]. OIQ-10K broadened the design to four distortion situations—no perceptible distortion, one distorted region, two non-adjacent distorted regions, and global distortion—producing 10,000 images and recording quality values, distortion spatial distributions, and head and eye movements [2502.15271]. These datasets shifted evaluation away from globally uniform degradations and toward local quality variation, distortion range, and viewing-order effects.

AI-generated panoramas introduced additional annotation dimensions. AIGCOIQA2024 built a 300-image ERP dataset from 25 prompts and five AIGC engines, and collected ratings for quality, comfortability, and correspondence under ITU-R BT.500-14 in a Unity-based head-mounted display environment [2404.01024]. OHF2024 extended this line to 600 AI-generated omnidirectional images and added distortion-aware saliency maps obtained by letting subjects click distorted salient regions while voicing distortion descriptions [2506.21925].

## 3. Full-reference and projection-aware OIQA

Early OIQA relied heavily on adapting planar full-reference metrics to spherical imagery. In the OIQA database study, nine state-of-the-art FR models—PSNR, SSIM, MS-SSIM, IW-SSIM, VIF, FSIM, GMSD, GSI, and VSI—were evaluated after five-parameter logistic fitting. FSIM, GSI, and VSI were the top three performers on that database, with FSIM reaching PLCC \(=0.9171\), SRCC \(=0.9110\), and RMSE \(=0.8221\), while PSNR and SSIM performed poorly [2207.02674]. The same study emphasized two corrective mechanisms for \(360^\circ\) content: spherical angular weighting by \(\cos(\theta)\) and attentional weighting from head/eye saliency.

A separate line of work replaced global ERP comparison with viewport-domain evaluation. In the moving-camera formulation, a static panorama is transformed into videos by following user scanpaths and extracting tangent-plane viewports over time; standard 2D FR models such as PSNR, SSIM, VIF, NLPD, and DISTS are then applied frame by frame, followed by temporal hysteresis pooling and averaging across users [2005.10547]. On the authors’ database, the resulting O-DISTS achieved PLCC/SRCC \(=0.660/0.613\), exceeding projection-based and viewport-based baselines, which the paper interpreted as evidence that viewing conditions and browsing trajectories materially affect perceived quality.

Tangential projection has also been used to avoid ERP distortion more directly. For super-resolved omnidirectional images, tangential views are produced by gnomonic projection on a once-subdivided icosahedron, giving \(N_t=80\) distortion-free local views. Any 2D FR metric \(Q(\cdot,\cdot)\) can then be extended by averaging over views,
\[
t\text{-}Q(\tilde y,y)=\frac{1}{N_t}\sum_{i=1}^{N_t}Q(\tilde t_i,t_i).
\]
In that framework, most objective metrics favored CNN-based super-resolution, whereas subjective tests favored GAN-based architectures; among eleven tangential metrics, NLPD and GMSD aligned best with human preferences [2101.10396].

A task-specific FR formulation was developed for omnidirectional stitching. The cross-reference stitching dataset captures four fisheye frames at headings \(0^\circ\), \(90^\circ\), \(180^\circ\), and \(270^\circ\), allowing the orthogonal fisheye pair to serve as cross-reference for seam regions of the stitched pair. FR metrics such as MSE, PSNR, and SSIM were then computed only on the cross-reference mask, while BRISQUE, NIQE, PIQE, and CNN-IQA were likewise restricted to stitching regions [1904.04960]. The design reflects a broader OIQA principle: quality often depends on localized, geometrically structured regions rather than whole-frame fidelity.

## 4. No-reference OIQA: from NSS and SVR to transformers and graphs

No-reference OIQA developed first through handcrafted features and support vector regression. MFILGN decomposes the ERP image by discrete Haar wavelet transform, uses entropy intensities of low- and high-frequency subbands for multi-frequency information, extracts global natural scene statistics from the ERP, extracts local NSS from sampled viewports, concatenates these features, and regresses quality with SVR [2102.11393]. On OIQA and CVIQD, MFILGN reported SRCC/PLCC \(=0.9614/0.9695\) and \(0.9670/0.9751\), respectively. S\(^2\) took a complementary route by combining local statistics from sampled viewports with global semantic features from the full ERP image, again using SVR fusion; on CVIQD it reported SROCC \(=0.971\), PLCC \(=0.978\), and RMSE \(=2.894\) [2302.12393].

Graph-based methods introduced explicit modeling of inter-viewport dependency. VGCN selects \(N=20\) viewports using SURF keypoints and a Gaussian heatmap, extracts ResNet-18 features for each viewport, connects viewports whose angular distance is at most \(45^\circ\), propagates information through five GCN layers, and fuses the local graph score with a global ERP quality estimate from a DB-CNN branch [2002.09140]. On OIQA it reported PLCC \(=0.9584\), SRCC \(=0.9515\), and RMSE \(=0.5967\); on CVIQD it reported PLCC \(=0.9651\), SRCC \(=0.9639\), and RMSE \(=3.6573\). A later hierarchical graph attention model sampled viewports uniformly by the Fibonacci-sphere method, used a Swin backbone, a local GAT, and a graph transformer for long-range interactions, and reported PLCC/SRCC/RMSE \(=0.840/0.840/0.328\) on JUFE-10K and \(0.837/0.833/0.251\) on OIQ-10K [2508.09843].

Transformer-based methods substantially redefined the field. ST360IQ extracts tangent viewports from the salient parts of the ERP image using ATSal saliency, mean-shift clustering, and top-\(10\%\) saliency regions, processes each viewport by a ResNet-50 front-end and a ViT-Tiny-scale spherical vision transformer with positional, geometric, and source embeddings, and averages viewport scores to produce the final quality prediction [2303.06907]. On CVIQ, ST360IQ reported PLCC \(=0.98\), SRCC \(=0.98\), RMSE \(=2.98\); on OIQA, PLCC \(=0.96\), SRCC \(=0.97\), RMSE \(=0.57\). Ablations showed drops of \(2\!-\!3\) points in PLCC/SRCC without saliency sampling and a further \(2\!-\!4\) point drop without tangent projection.

Other models focused on browsing-process simulation. Assessor360 constructs multiple pseudo-viewport sequences using Recursive Probability Sampling, combines distorted and semantic features through a Multi-scale Feature Aggregation module with a Distortion-aware Block, and models viewport transitions with a GRU-based Temporal Modeling Module [2305.10983]. GSR instead generates multiple scanpaths under a specified viewing condition, extracts spherical-tangent foveal patches, assembles them into a unique generative scanpath representation, and evaluates quality with a video transformer backbone; on JUFE it reported SRCC \(\simeq 0.818\) and PLCC \(\simeq 0.830\), while reducing computational complexity by \(3\!-\!4\) orders of magnitude compared with viewport-based methods on \(8\!-\!11\)K inputs [2309.03472].

Large-scale non-uniform datasets led to specialized architectures. OIQAND uses eight equatorial viewports, multi-scale Swin features, distortion-adaptive viewport and channel attention, and a multi-head self-attention quality head, reporting overall PLCC \(=0.800\), SROCC \(=0.800\), and RMSE \(=0.362\) on JUFE-10K [2501.11511]. MTAOIQA introduced auxiliary tasks for distortion range, type, and degree, achieving PLCC/SRCC/RMSE \(=0.822/0.821/0.344\) on JUFE-10K and \(0.829/0.824/0.256\) on OIQ-10K [2501.11512]. Max360IQ used a MaxViT backbone, multi-scale feature integration, deep semantic guidance, and GRU-based recency-aware regression, and reported gains over Assessor360 on JUFE, OIQA, and CVIQ [2502.19046]. A different response to the same scalability problem was the viewport-unaware paradigm VU-BOIQA, which discarded viewport generation entirely, sampled ERP patches via an adaptive prior-equator scheme, fused deformation-immune features with DCNv3 and attention, and achieved competitive performance with 30.2M parameters and 40.8G FLOPs across CVIQ, OIQA, JUFE-10K, and OIQ-10K [2503.06129].

## 5. Non-uniform distortion and browsing behavior

A central development in OIQA has been the recognition that locally non-uniform distortion is not a minor extension of global distortion but a qualitatively different regime. JUFE-10K was explicitly designed to study this regime by perturbing one or two camera lenses before stitching, thereby generating non-uniform Gaussian noise, Gaussian blur, brightness discontinuity, and stitching distortion [2501.11511]. The associated subjective analysis reported that distortion level strongly correlates monotonically with MOS, Gaussian noise tends to receive higher MOS than brightness discontinuity, Gaussian blur, and stitching distortion, and images with two disturbed lenses receive lower MOS than single-lens distortion. The same study found that viewing-initial-viewport has minimal effect on final MOS under 15-second free exploration, because subjects can compensate by later exploration [2501.11511].

That finding coexists with earlier evidence that starting point and its interaction with exploration time can be highly significant under other designs, especially when localized distortion is encountered early [2005.10547]. The two results are not identical experiments: one uses fixed condition groups and voice-prompt scoring at 5 s and 15 s, while the other uses free exploration over 15 s on a different dataset. Taken together, they indicate that browsing effects are condition-dependent rather than negligible.

This realization led to models that explicitly simulate or infer viewport trajectories. Assessor360 models multiple pseudo-assessor sequences from a common starting point [2305.10983]; GSR aggregates multi-hypothesis scanpaths under a specified viewing condition [2309.03472]; Max360IQ uses scanpath-derived viewport sequences for nonuniformly distorted images [2502.19046]. By contrast, OIQAND and IQCaption360 report that simple equatorial sampling can be nearly as effective as more complex sampling schemes on their large non-uniform benchmarks [2501.11511; 2502.15271]. This suggests that the optimal balance between behavior realism and computational efficiency remains an open methodological question rather than a settled design rule.

A related misconception, contradicted by several datasets, is that global ERP quality is sufficient if the regressor is strong enough. On JUFE-10K and OIQ-10K, 2D-IQA methods and several earlier OIQA models degrade markedly under local stitching or localized region distortions, while models with explicit local distortion modeling, multitask supervision, or adaptive aggregation perform better [2501.11511; 2501.11512].

## 6. Adjacent domains: AI-generated panoramas, stitching, super-resolution, and stereoscopic depth

OIQA has expanded beyond traditional photographic degradations. AI-generated omnidirectional images exhibit low-level artifacts such as blur, noisy texture, uneven illumination, and mild geometric stretching, but high-level artifacts become dominant: unrealism in surface detail, implausible spatial layout or object composition, and text-to-image mismatches including missing objects and hallucinated elements [2404.01024]. AIGCOIQA2024 therefore evaluates three perceptual dimensions—quality, comfortability, and correspondence—and found that no single no-reference IQA model excels simultaneously on all three; TReS and HyperIQA were strongest for quality, while MANIQA led comfortability and correspondence [2404.01024].

OHF2024 extended this direction by adding distortion-aware saliency to AI-generated omnidirectional images. BLIP2OIQA uses six viewports and a shared BLIP-2 encoder to regress quality, comfortability, and correspondence, while BLIP2OISal predicts a distortion-aware saliency map from the full ERP image and prompt [2506.21925]. On the test split, BLIP2OIQA reported quality SRCC/PLCC \(=0.9074/0.8954\), comfort \(=0.8763/0.9011\), and correspondence \(=0.8354/0.8600\); BLIP2OISal improved over the best baseline in AUC, NSS, CC, SIM, and KLD [2506.21925]. This establishes a multimodal variant of OIQA in which semantic alignment is part of the quality target.

Stitching quality forms another specialized branch. The cross-reference stitching dataset used dual-fisheye captures at four headings to provide near distortion-free reference content specifically for seam regions, revealing that off-the-shelf IQA indices do not explicitly model ghosting, seam displacement, or chromatic misalignment [1904.04960]. Super-resolution quality assessment constitutes a further branch, where tangential projection enables the reuse of standard 2D FR metrics but subjective tests reveal a divergence between structural fidelity metrics and human preference for GAN-generated textures [2101.10396].

Stereoscopic omnidirectional quality introduces depth as an explicit perceptual target. The Depth Quality Index computes interocular discrepancy \(D=\lvert I_{dl}-I_{dr}\rvert\), converts it to CIE LAB, extracts equatorial viewports, decomposes them with a one-level discrete Haar wavelet transform, computes standard deviation and entropy statistics, and regresses depth quality with SVR [2408.10134]. On the SOLID stereoscopic omnidirectional database, adaptive-view DQI reported SROCC/KROCC/PLCC \(=0.930/0.781/0.948\), and fusing DQI with MS-SSIM improved overall QoE prediction to SROCC \(=0.930\), KROCC \(=0.789\), and PLCC \(=0.936\) [2408.10134]. The result indicates that “quality” in omnidirectional media is not exhausted by monocular distortion visibility.

## 7. Open challenges and research directions

Several open problems recur across the literature. One is the dependence of many methods on external saliency or scanpath priors. ST360IQ relies on ATSal and identifies joint end-to-end learning of saliency and quality as a natural extension [2303.06907]. GSR notes that current systems still rely on a fixed viewing condition \(\Omega\), and that automatically inferring or adapting \(\Omega\) remains open, as does personalization through user-specific scanpath priors [2309.03472]. Assessor360 likewise notes that its Recursive Probability Sampling uses fixed step sizes and directions [2305.10983].

Another challenge is dataset realism and scale. ST360IQ explicitly notes the lack of very large \(360^\circ\) IQA datasets and points to self-supervised pretraining on unlabeled omnidirectional imagery [2303.06907]. MTAOIQA observes that synthetic distortions in JUFE-10K and OIQ-10K may differ from real-world equirectangular capture artifacts [2501.11512]. The emergence of JUFE-10K, OIQ-10K, AIGCOIQA2024, and OHF2024 addresses the scale issue partially, but each targets a specific regime—non-uniform lens-level distortions, mixed spatial distributions, or AI-generated content—rather than a unified benchmark.

Computational efficiency remains a practical constraint. ST360IQ identifies its 14-layer transformer plus ResNet-50 as potentially heavy for real-time monitoring on head-mounted displays [2303.06907]. VGCN likewise describes its local-global architecture as costly for real-time VR applications [2002.09140]. The viewport-unaware paradigm [2503.06129] and the compact video representation of GSR [2309.03472] can be read as direct responses to this constraint.

Finally, temporal extension is still incomplete. ST360IQ explicitly states that it addresses only still-image IQA and leaves omnidirectional video with temporal modeling open [2303.06907]. The moving-camera framework and GSR demonstrate that even static OIQA already contains temporal browsing structure [2005.10547; 2309.03472]. A plausible implication is that future omnidirectional video quality assessment will likely need to combine spherical geometry, scanpath uncertainty, and distortion non-uniformity rather than adding temporal pooling to an otherwise image-centric model.

In aggregate, OIQA has evolved from spherical adaptations of planar metrics into a broader perceptual modeling discipline centered on geometry-aware sampling, viewport or patch interaction, behavior-conditioned aggregation, and increasingly multidimensional quality targets. The field’s present trajectory moves simultaneously toward larger and more heterogeneous databases, stronger no-reference models, and richer definitions of quality that include comfort, semantic correspondence, local distortion distribution, and, in stereoscopic settings, depth quality.

Source: https://www.emergentmind.com/topics/omnidirectional-image-quality-assessment-oiqa