JPEG AIC-3 Test Methodology
- JPEG AIC-3 is a subjective image quality assessment framework that uses fine-grained triplet comparisons and Thurstone Case V reconstruction to derive precise JND scales.
- It integrates plain and boosted comparisons to amplify subtle distortions, enabling reliable ratings in the 0–3 JND range across diverse codecs and content types.
- The methodology applies rigorous observer control and bootstrap-based confidence analysis, making it versatile for HDR, SDR, and light-field compression evaluations.
Searching arXiv for the cited JPEG AIC-3 papers to ground the article in the current literature. Search query: JPEG AIC-3 triplet comparison high-fidelity compressed images arXiv JPEG AIC-3 Test Methodology is a subjective image quality assessment framework for the very high-fidelity regime, designed to obtain extremely fine-grained, just-noticeable-difference (JND)-based quality scales of compressed images when conventional five-point, continuous, DSIS, or DSCQS procedures lose resolution and observers struggle to distinguish “visually lossless” from “nearly lossless.” Within ISO/IEC JTC1 SC29 WG1, it is described as the core of Part 3 of the Advanced Image Coding and Evaluation series, ISO/IEC 29170-3, and it builds on the earlier AIC-1 and AIC-2 activities by combining triplet comparisons, visual stimulus boosting, Thurstone Case V reconstruction, and bootstrap confidence analysis into a repeatable methodology for aligning multiple codecs and content types on a common JND scale (Testolina et al., 2024, Jenadeleh et al., 14 Jun 2025).
1. Standardization Context and Intended Operating Regime
The central objective of JPEG AIC-3 is to provide a subjective test that is sensitive to quality differences in the $0$–$3$ JND range and to supply a data-analysis pipeline that reconstructs linear quality scales in JND units, aligned across multiple codecs and content types. The methodology targets compressed images whose artifacts are extremely subtle, including conditions below or near the visually-lossless threshold. In the standardization narrative reported by Testolina et al., AIC-3 “sits on top of” AIC-1, which provides guidelines for image coding evaluation, and AIC-2, which provides a near-lossless flicker test; the distinguishing feature of AIC-3 is that it extends subjective testing into the very high-fidelity regime through fine-grained triplet judgments and explicit JND reconstruction (Testolina et al., 2024).
A persistent misconception is that AIC-3 is only a flicker-based visibility test. The literature instead describes a broader framework in which boosted judgments are coupled to plain judgments and jointly mapped onto a perceptual scale. In the HDR implementation, the methodology is explicitly described as a combination of Plain Triplet Comparisons (PTC) and Boosted Triplet Comparisons (BTC), with a unified functional model fitted over all responses; in the earlier SDR crowdsourced study, boosted and plain experiments are similarly paired so that boosted measurements can be rescaled back to the unboosted perceptual scale (Jenadeleh et al., 14 Jun 2025, Testolina et al., 2024).
The operating regime of the method is not limited to a single image modality. The documented applications include standard-dynamic-range high-fidelity compression, HDR and wide color gamut compression, learning-based compression with JPEG AI, and low-distortion coded light fields with view synthesis. This breadth does not imply identical stimulus presentation in every study; rather, it indicates that the AIC-3 framework has been adapted to several domains while preserving its reliance on comparative judgments and Thurstonian scale reconstruction (Jenadeleh et al., 14 Jun 2025, Saraiva et al., 18 Sep 2025).
2. Triplet Structure, Stimulus Construction, and Boosting
The canonical AIC-3 trial is a triplet , where is the uncompressed reference and and are two distorted versions that must be compared against that same reference. In PTC, observers judge which of the two distorted images is more degraded relative to the reference; in BTC, the distorted images are modified so that minute differences become more visible. Across studies, the response set is “Left,” “Right,” or “Not Sure,” and “Not Sure” is retained during data collection to reduce fatigue while being split equally between the two alternatives during analysis (Testolina et al., 2024, Jenadeleh et al., 14 Jun 2025).
The boosting concept is specified in detail in the SDR formulation. Each test image undergoes three transformations. First, a zoom operation center-crops the stimulus to half the original width and height and then up-scales it back to full size via a Lanczos filter. For an original stimulus of size , the boosted image is constructed as
- ,
- . Second, artifact amplification is applied relative to the reference $3$0:
$3$1
with $3$2, independently in $3$3, $3$4, and $3$5 channels. Third, the boosted test image flickers against the reference at $3$6 Hz, with $3$7 ms reference and $3$8 ms test, repeatedly (Testolina et al., 2024).
The HDR implementation uses the same distinction between plain and boosted comparisons but instantiates the stimuli differently. In PTC, two compressed versions of the same source are shown side by side on the central $3$9K region of a 0K HDR display, and the observer toggles each test image in place against the reference before deciding which is more distorted. In BTC, each compressed image is 1 zoomed and alternated, or flickered, at 2 Hz with the reference, and observers select which flickers more strongly. The stated rationale is that BTC’s zoom and flicker “boost visibility of minute differences that plain toggling might miss,” thereby extending reliable discrimination below the visually-lossless threshold (Jenadeleh et al., 14 Jun 2025).
Stimulus design is codec- and content-aware. In the HDR dataset AIC-HDR2025, five diverse HDR source images—landscape, night scene, flower, portrait, and architecture—each encoded as 3-bit Rec. 2100 PQ at 4 pixels, were compressed with JPEG XT, JPEG XL, AVIF, and JPEG AI at five roughly evenly spaced distortion levels per codec, covering 5–6 JND and beyond up to approximately 7 JND, for a total of 8 compressed stimuli plus 9 uncompressed references (Jenadeleh et al., 14 Jun 2025). In the earlier SDR reference set, five 0 crops were compressed with JPEG, JPEG 2000, VVC Intra, JPEG XL, and AVIF at ten rate points chosen so that their perceptual distances from the source were approximately 1 JND, yielding 2 stimuli (Testolina et al., 2024). A plausible implication is that AIC-3 is designed around dense local sampling of the quality axis rather than sparse endpoint comparisons.
3. Experimental Protocol, Batch Design, and Observer Control
AIC-3 organizes triplets into structured batches that serve both estimation and quality-control functions. In the HDR implementation, each source and codec defines six quality points, consisting of five distortion levels plus the reference. Same-codec triplets are all unordered pairs of distinct distortion levels within one codec, amounting to 3 per codec per method, while cross-codec triplets are randomly selected pairs across codecs, 4 per source, used to align scales. This produces 5 triplets per source per method and 6 PTC plus 7 BTC triplets across five sources. Triplets are grouped into six batches of 8 questions, and each triplet 9 is paired with its mirror 0 in the same batch to permit consistency checks (Jenadeleh et al., 14 Jun 2025).
Observer management is an explicit part of the methodology. In the HDR laboratory study, each triplet was judged by 1 independent observers. Observers completed at most two batches, with a forced three-minute break between batches; reported durations were 2 min per PTC batch and 3 min per BTC batch. The practical guideline derived from that study is to collect 4 independent judgments per triplet to achieve 5 JND precision at the 6 level, to limit observers to two 7-item batches per method, and to monitor each batch’s accuracy and consistency using mirror pairs, discarding batches with both below 8 (Jenadeleh et al., 14 Jun 2025).
The SDR crowdsourced studies implement analogous but not identical controls. In Testolina et al., triplet sets were split into 9 randomized batches of 0 BTC or 1 PTC trials, and batch filtering required “trap” triplets—level 2 versus level 3 same-codec—to be answered correctly at least 4 of the time. The JPEG AI study used 5 trap questions per batch, fully randomized question order, and up to two batches per MTurk worker; it further computed accuracy and consistency per batch, then applied Otsu thresholding on the average of the two measures, with thresholds 6 for PTC and 7 for BTC, before applying ISO JPEGAIC-3 outlier detection (Testolina et al., 2024, Jenadeleh et al., 7 Apr 2025).
The controlled laboratory conditions reported for HDR are specific and stringent: four fully controlled labs, three using MacBook Pro Liquid Retina XDR and one using Sony BVM-HX310, calibrated to 8 cd/m9 white (D65) and 0 cd/m1 peak with ambient light of 2 lux. Viewing distance was 3 stimulus height for PTC and 4 for BTC, citing ITU-R BT.2246-8. By contrast, the SDR and JPEG AI studies were run on Amazon Mechanical Turk under uncontrolled setups, albeit with qualification quizzes, trap trials, and display recommendations. This suggests that AIC-3 can be deployed under both crowdsourced and laboratory conditions, but the precision claims documented for HDR were obtained under fully controlled viewing (Jenadeleh et al., 14 Jun 2025, Mohammadi et al., 16 Sep 2025).
4. Psychometric Reconstruction and JND Scaling
The statistical core of AIC-3 is Thurstone Case V. In the HDR formulation, each stimulus 5 is assigned a latent JND impairment 6, and for a triplet 7 the probability that the observer chooses 8 as more distorted is modeled as
9
where 0 is the standard normal cumulative distribution function and 1 is the common judgment noise, set to 2 for JND units. “Not sure” responses are split equally between 3 and 4 (Jenadeleh et al., 14 Jun 2025).
To combine PTC and BTC, AIC-3 uses a unified functional model. For each codec and source, the plain, unboosted impairment is modeled as an exponential rate-distortion curve,
5
and the relation between the plain and boosted scales is modeled by a quadratic transform,
6
A compressed image at bitrate 7 therefore has PTC impairment 8 and BTC impairment 9. In the HDR study, the 0 parameters 1 were jointly optimized by maximum-likelihood estimation over all PTC and BTC responses. The JPEG AI study describes the same unified exponential-plus-quadratic model and likewise estimates 2 by MLE over the combined data (Jenadeleh et al., 14 Jun 2025, Jenadeleh et al., 7 Apr 2025).
An earlier SDR reconstruction describes the same psychometric basis but makes the JND conversion step explicit. After maximum-likelihood estimation under the Thurstonian Case V model, the raw scale values 3 satisfy
4
Because one JND corresponds to 5 units on that internal scale, conversion to JND units is
6
That study then fits, separately for each source-codec pair, a second-degree polynomial to recover the unboosted scale from the boosted JND values: 7 The later unified model can be read as an integrated parametric reformulation of the same plain-versus-boosted relation (Testolina et al., 2024).
A related adaptation appears in light-field quality assessment. There, pairwise counts are collected in a matrix 8, empirical probabilities are computed as 9, transformed to 0-scores via 1, and the scale values 2 are then obtained by maximum-likelihood estimation under Case V assumptions. The final values are linearly rescaled to 3 and inverted so that higher 4 indicates better quality (Saraiva et al., 18 Sep 2025). The commonality across these variants is the treatment of subjective comparisons as noisy pairwise evidence about an underlying latent scale.
5. Confidence Intervals, Precision, and Reported Reliability
Confidence estimation in AIC-3 is bootstrap-based. In the HDR study, triplets were resampled with replacement 5 times, the scale reconstruction was repeated, and at 6 grid points in bitrate 7 the distribution of 8 was recorded; the 9 and $3$00 percentiles of that distribution yielded the $3$01 confidence interval. The practical recommendation given in that work is to estimate $3$02 confidence intervals by nonparametric bootstrap with at least $3$03 resamples and to report rate-distortion curves with shaded confidence bands, including the average confidence-interval width at $3$04 JND (Jenadeleh et al., 14 Jun 2025).
The principal precision result reported for AIC-HDR2025 is that the average width of the $3$05 confidence interval around $3$06 JND is $3$07 JND, with the paper further stating that the confidence intervals are tightly bounded throughout a range from imperceptible distortion to clear distortion at $3$08–$3$09 JND. The authors interpret these narrow intervals as evidence that AIC-3 can resolve fractional JND steps even in HDR content where artifacts are extremely subtle (Jenadeleh et al., 14 Jun 2025).
Earlier SDR work reports the same bootstrap logic at larger scale. Testolina et al. process $3$10 bootstrap replicates end-to-end, resampling responses per triplet with replacement and obtaining $3$11 confidence bands on the final JND scale for each image and codec. The JPEG AI study uses $3$12 bootstrap resamples of the filtered PTC and BTC data, refits the model for each resample, and defines the $3$13 confidence interval at each bitrate from the $3$14th and $3$15th percentiles. Across the literature, the methodology therefore treats uncertainty as intrinsic to the reconstructed perceptual scale rather than as an afterthought attached only to summary statistics (Testolina et al., 2024, Jenadeleh et al., 7 Apr 2025).
A common oversimplification is to read a single JND score as an exact perceptual quantity. The published analyses do not support that interpretation. They report a mean scale value together with uncertainty, and in the objective-metric evaluation framework this uncertainty is elevated to a first-class variable through per-image $3$16 and $3$17, a high-fidelity subset defined by $3$18, and an outlier criterion based on $3$19 with $3$20 (Mohammadi et al., 16 Sep 2025). This suggests that the methodological unit of interest is not merely the point estimate, but the estimate together with its confidence structure.
6. Applications, Adaptations, and Relation to Objective Metric Benchmarking
AIC-3 has been used to build datasets and evaluation pipelines for several compression scenarios. The HDR study introduces AIC-HDR2025 as “the first such HDR dataset,” comprising $3$21 test images generated from five HDR sources, each compressed using four codecs at five compression levels, and reports $3$22 ratings from $3$23 participants across four fully controlled labs (Jenadeleh et al., 14 Jun 2025). The JPEG AI study constructs a dataset of $3$24 compressed images from five diverse sources, collects $3$25 triplet responses from $3$26 participants, and uses the resulting JND scales to compare JPEG AI with other codecs in the high-fidelity regime (Jenadeleh et al., 7 Apr 2025). Testolina et al. provide a $3$27-stimulus SDR dataset and code to reproduce the results (Testolina et al., 2024).
The methodology has also been adapted beyond conventional single-image coding. In the light-field study, the proposed JPEG AIC-3 test methodology is applied to low-distortion coded light fields with view synthesis. Two stimuli are displayed side by side, each alternating between an original and a coded view at $3$28 Hz, so that both sides flicker simultaneously; observers select the side whose flicker appears stronger. The test covers intra-method, cross-codec, encoding-method, bias-control, and attention-check triplets, with $3$29 unique triplets, $3$30 screened participants, $3$31 evaluations per stimulus pair, a $3$32-trial practice block, side inversion in $3$33 of presentations, and a mandatory mid-session rest (Saraiva et al., 18 Sep 2025). This does not redefine the core AIC-3 psychometric logic, but it demonstrates that the triplet-comparison framework can be specialized to synthesized-view artifacts.
AIC-3 is also used as a ground-truth substrate for objective image quality metric evaluation. In the metric-benchmarking study, subjective scale reconstruction on the JPEG AIC-3 dataset yields for each compressed image a mean $3$34 in JND and a standard deviation $3$35, enabling evaluation on the full range “All” and on high-fidelity and medium-fidelity subsets. Objective metric scores $3$36 are mapped to the JND domain through a four-parameter logistic,
$3$37
and then evaluated by RMSE, PLCC, SROCC, Kendall’s tau, Outlier Ratio, Perceptually-Weighted Rank Correlation, and the proposed Z-RMSE,
$3$38
The same study introduces the Meng–Rosenthal–Rubin test for comparing correlated SROCCs between metrics on the common subjective ground truth (Mohammadi et al., 16 Sep 2025). The JPEG AI study likewise uses a four-parameter logistic fit to align full-reference metrics to the JND scale and reports that CVVDP achieves the highest overall performance, while also noting that most metrics, including CVVDP, were overly optimistic for JPEG AI-compressed images (Jenadeleh et al., 7 Apr 2025).
Practical recommendations reported across these studies converge on several design principles: choose at least five distortion levels spanning $3$39–$3$40 JND per codec; include $3$41–$3$42 cross-codec triplets to align scales; permit “Not Sure”; filter unreliable observers using trap trials, mirror-pair consistency, or both; adopt the unified functional model $3$43 and $3$44 with Thurstone Case V MLE; bootstrap the entire pipeline; and report rate-distortion curves with confidence bands and per-codec or per-source breakdowns (Jenadeleh et al., 14 Jun 2025, Testolina et al., 2024). In that sense, JPEG AIC-3 is best understood not as a single display script, but as a standardized psychometric methodology for reconstructing fine-grained JND quality scales from carefully designed triplet judgments in the high-fidelity regime.