Papers
Topics
Authors
Recent
Search
2000 character limit reached

JPEG AIC-3 Test Methodology

Updated 12 July 2026
  • JPEG AIC-3 is a subjective image quality assessment framework that uses fine-grained triplet comparisons and Thurstone Case V reconstruction to derive precise JND scales.
  • It integrates plain and boosted comparisons to amplify subtle distortions, enabling reliable ratings in the 0–3 JND range across diverse codecs and content types.
  • The methodology applies rigorous observer control and bootstrap-based confidence analysis, making it versatile for HDR, SDR, and light-field compression evaluations.

Searching arXiv for the cited JPEG AIC-3 papers to ground the article in the current literature. Search query: JPEG AIC-3 triplet comparison high-fidelity compressed images arXiv JPEG AIC-3 Test Methodology is a subjective image quality assessment framework for the very high-fidelity regime, designed to obtain extremely fine-grained, just-noticeable-difference (JND)-based quality scales of compressed images when conventional five-point, continuous, DSIS, or DSCQS procedures lose resolution and observers struggle to distinguish “visually lossless” from “nearly lossless.” Within ISO/IEC JTC1 SC29 WG1, it is described as the core of Part 3 of the Advanced Image Coding and Evaluation series, ISO/IEC 29170-3, and it builds on the earlier AIC-1 and AIC-2 activities by combining triplet comparisons, visual stimulus boosting, Thurstone Case V reconstruction, and bootstrap confidence analysis into a repeatable methodology for aligning multiple codecs and content types on a common JND scale (Testolina et al., 2024, Jenadeleh et al., 14 Jun 2025).

1. Standardization Context and Intended Operating Regime

The central objective of JPEG AIC-3 is to provide a subjective test that is sensitive to quality differences in the $0$–$3$ JND range and to supply a data-analysis pipeline that reconstructs linear quality scales in JND units, aligned across multiple codecs and content types. The methodology targets compressed images whose artifacts are extremely subtle, including conditions below or near the visually-lossless threshold. In the standardization narrative reported by Testolina et al., AIC-3 “sits on top of” AIC-1, which provides guidelines for image coding evaluation, and AIC-2, which provides a near-lossless flicker test; the distinguishing feature of AIC-3 is that it extends subjective testing into the very high-fidelity regime through fine-grained triplet judgments and explicit JND reconstruction (Testolina et al., 2024).

A persistent misconception is that AIC-3 is only a flicker-based visibility test. The literature instead describes a broader framework in which boosted judgments are coupled to plain judgments and jointly mapped onto a perceptual scale. In the HDR implementation, the methodology is explicitly described as a combination of Plain Triplet Comparisons (PTC) and Boosted Triplet Comparisons (BTC), with a unified functional model fitted over all responses; in the earlier SDR crowdsourced study, boosted and plain experiments are similarly paired so that boosted measurements can be rescaled back to the unboosted perceptual scale (Jenadeleh et al., 14 Jun 2025, Testolina et al., 2024).

The operating regime of the method is not limited to a single image modality. The documented applications include standard-dynamic-range high-fidelity compression, HDR and wide color gamut compression, learning-based compression with JPEG AI, and low-distortion coded light fields with view synthesis. This breadth does not imply identical stimulus presentation in every study; rather, it indicates that the AIC-3 framework has been adapted to several domains while preserving its reliance on comparative judgments and Thurstonian scale reconstruction (Jenadeleh et al., 14 Jun 2025, Saraiva et al., 18 Sep 2025).

2. Triplet Structure, Stimulus Construction, and Boosting

The canonical AIC-3 trial is a triplet (Ii,I0,Ik)(I_i, I_0, I_k), where I0I_0 is the uncompressed reference and IiI_i and IkI_k are two distorted versions that must be compared against that same reference. In PTC, observers judge which of the two distorted images is more degraded relative to the reference; in BTC, the distorted images are modified so that minute differences become more visible. Across studies, the response set is “Left,” “Right,” or “Not Sure,” and “Not Sure” is retained during data collection to reduce fatigue while being split equally between the two alternatives during analysis (Testolina et al., 2024, Jenadeleh et al., 14 Jun 2025).

The boosting concept is specified in detail in the SDR formulation. Each test image undergoes three transformations. First, a zoom operation center-crops the stimulus to half the original width and height and then up-scales it back to full size via a Lanczos filter. For an original stimulus II of size W×HW \times H, the boosted image is constructed as

  1. crop(I,W4:3W4,H4:3H4)\mathrm{crop}(I,\tfrac{W}{4}:\tfrac{3W}{4},\tfrac{H}{4}:\tfrac{3H}{4}),
  2. Lanczos2(crop;W,H)\mathrm{Lanczos}_2(\text{crop}; W,H). Second, artifact amplification is applied relative to the reference $3$0:

$3$1

with $3$2, independently in $3$3, $3$4, and $3$5 channels. Third, the boosted test image flickers against the reference at $3$6 Hz, with $3$7 ms reference and $3$8 ms test, repeatedly (Testolina et al., 2024).

The HDR implementation uses the same distinction between plain and boosted comparisons but instantiates the stimuli differently. In PTC, two compressed versions of the same source are shown side by side on the central $3$9K region of a (Ii,I0,Ik)(I_i, I_0, I_k)0K HDR display, and the observer toggles each test image in place against the reference before deciding which is more distorted. In BTC, each compressed image is (Ii,I0,Ik)(I_i, I_0, I_k)1 zoomed and alternated, or flickered, at (Ii,I0,Ik)(I_i, I_0, I_k)2 Hz with the reference, and observers select which flickers more strongly. The stated rationale is that BTC’s zoom and flicker “boost visibility of minute differences that plain toggling might miss,” thereby extending reliable discrimination below the visually-lossless threshold (Jenadeleh et al., 14 Jun 2025).

Stimulus design is codec- and content-aware. In the HDR dataset AIC-HDR2025, five diverse HDR source images—landscape, night scene, flower, portrait, and architecture—each encoded as (Ii,I0,Ik)(I_i, I_0, I_k)3-bit Rec. 2100 PQ at (Ii,I0,Ik)(I_i, I_0, I_k)4 pixels, were compressed with JPEG XT, JPEG XL, AVIF, and JPEG AI at five roughly evenly spaced distortion levels per codec, covering (Ii,I0,Ik)(I_i, I_0, I_k)5–(Ii,I0,Ik)(I_i, I_0, I_k)6 JND and beyond up to approximately (Ii,I0,Ik)(I_i, I_0, I_k)7 JND, for a total of (Ii,I0,Ik)(I_i, I_0, I_k)8 compressed stimuli plus (Ii,I0,Ik)(I_i, I_0, I_k)9 uncompressed references (Jenadeleh et al., 14 Jun 2025). In the earlier SDR reference set, five I0I_00 crops were compressed with JPEG, JPEG 2000, VVC Intra, JPEG XL, and AVIF at ten rate points chosen so that their perceptual distances from the source were approximately I0I_01 JND, yielding I0I_02 stimuli (Testolina et al., 2024). A plausible implication is that AIC-3 is designed around dense local sampling of the quality axis rather than sparse endpoint comparisons.

3. Experimental Protocol, Batch Design, and Observer Control

AIC-3 organizes triplets into structured batches that serve both estimation and quality-control functions. In the HDR implementation, each source and codec defines six quality points, consisting of five distortion levels plus the reference. Same-codec triplets are all unordered pairs of distinct distortion levels within one codec, amounting to I0I_03 per codec per method, while cross-codec triplets are randomly selected pairs across codecs, I0I_04 per source, used to align scales. This produces I0I_05 triplets per source per method and I0I_06 PTC plus I0I_07 BTC triplets across five sources. Triplets are grouped into six batches of I0I_08 questions, and each triplet I0I_09 is paired with its mirror IiI_i0 in the same batch to permit consistency checks (Jenadeleh et al., 14 Jun 2025).

Observer management is an explicit part of the methodology. In the HDR laboratory study, each triplet was judged by IiI_i1 independent observers. Observers completed at most two batches, with a forced three-minute break between batches; reported durations were IiI_i2 min per PTC batch and IiI_i3 min per BTC batch. The practical guideline derived from that study is to collect IiI_i4 independent judgments per triplet to achieve IiI_i5 JND precision at the IiI_i6 level, to limit observers to two IiI_i7-item batches per method, and to monitor each batch’s accuracy and consistency using mirror pairs, discarding batches with both below IiI_i8 (Jenadeleh et al., 14 Jun 2025).

The SDR crowdsourced studies implement analogous but not identical controls. In Testolina et al., triplet sets were split into IiI_i9 randomized batches of IkI_k0 BTC or IkI_k1 PTC trials, and batch filtering required “trap” triplets—level IkI_k2 versus level IkI_k3 same-codec—to be answered correctly at least IkI_k4 of the time. The JPEG AI study used IkI_k5 trap questions per batch, fully randomized question order, and up to two batches per MTurk worker; it further computed accuracy and consistency per batch, then applied Otsu thresholding on the average of the two measures, with thresholds IkI_k6 for PTC and IkI_k7 for BTC, before applying ISO JPEGAIC-3 outlier detection (Testolina et al., 2024, Jenadeleh et al., 7 Apr 2025).

The controlled laboratory conditions reported for HDR are specific and stringent: four fully controlled labs, three using MacBook Pro Liquid Retina XDR and one using Sony BVM-HX310, calibrated to IkI_k8 cd/mIkI_k9 white (D65) and II0 cd/mII1 peak with ambient light of II2 lux. Viewing distance was II3 stimulus height for PTC and II4 for BTC, citing ITU-R BT.2246-8. By contrast, the SDR and JPEG AI studies were run on Amazon Mechanical Turk under uncontrolled setups, albeit with qualification quizzes, trap trials, and display recommendations. This suggests that AIC-3 can be deployed under both crowdsourced and laboratory conditions, but the precision claims documented for HDR were obtained under fully controlled viewing (Jenadeleh et al., 14 Jun 2025, Mohammadi et al., 16 Sep 2025).

4. Psychometric Reconstruction and JND Scaling

The statistical core of AIC-3 is Thurstone Case V. In the HDR formulation, each stimulus II5 is assigned a latent JND impairment II6, and for a triplet II7 the probability that the observer chooses II8 as more distorted is modeled as

II9

where W×HW \times H0 is the standard normal cumulative distribution function and W×HW \times H1 is the common judgment noise, set to W×HW \times H2 for JND units. “Not sure” responses are split equally between W×HW \times H3 and W×HW \times H4 (Jenadeleh et al., 14 Jun 2025).

To combine PTC and BTC, AIC-3 uses a unified functional model. For each codec and source, the plain, unboosted impairment is modeled as an exponential rate-distortion curve,

W×HW \times H5

and the relation between the plain and boosted scales is modeled by a quadratic transform,

W×HW \times H6

A compressed image at bitrate W×HW \times H7 therefore has PTC impairment W×HW \times H8 and BTC impairment W×HW \times H9. In the HDR study, the crop(I,W4:3W4,H4:3H4)\mathrm{crop}(I,\tfrac{W}{4}:\tfrac{3W}{4},\tfrac{H}{4}:\tfrac{3H}{4})0 parameters crop(I,W4:3W4,H4:3H4)\mathrm{crop}(I,\tfrac{W}{4}:\tfrac{3W}{4},\tfrac{H}{4}:\tfrac{3H}{4})1 were jointly optimized by maximum-likelihood estimation over all PTC and BTC responses. The JPEG AI study describes the same unified exponential-plus-quadratic model and likewise estimates crop(I,W4:3W4,H4:3H4)\mathrm{crop}(I,\tfrac{W}{4}:\tfrac{3W}{4},\tfrac{H}{4}:\tfrac{3H}{4})2 by MLE over the combined data (Jenadeleh et al., 14 Jun 2025, Jenadeleh et al., 7 Apr 2025).

An earlier SDR reconstruction describes the same psychometric basis but makes the JND conversion step explicit. After maximum-likelihood estimation under the Thurstonian Case V model, the raw scale values crop(I,W4:3W4,H4:3H4)\mathrm{crop}(I,\tfrac{W}{4}:\tfrac{3W}{4},\tfrac{H}{4}:\tfrac{3H}{4})3 satisfy

crop(I,W4:3W4,H4:3H4)\mathrm{crop}(I,\tfrac{W}{4}:\tfrac{3W}{4},\tfrac{H}{4}:\tfrac{3H}{4})4

Because one JND corresponds to crop(I,W4:3W4,H4:3H4)\mathrm{crop}(I,\tfrac{W}{4}:\tfrac{3W}{4},\tfrac{H}{4}:\tfrac{3H}{4})5 units on that internal scale, conversion to JND units is

crop(I,W4:3W4,H4:3H4)\mathrm{crop}(I,\tfrac{W}{4}:\tfrac{3W}{4},\tfrac{H}{4}:\tfrac{3H}{4})6

That study then fits, separately for each source-codec pair, a second-degree polynomial to recover the unboosted scale from the boosted JND values: crop(I,W4:3W4,H4:3H4)\mathrm{crop}(I,\tfrac{W}{4}:\tfrac{3W}{4},\tfrac{H}{4}:\tfrac{3H}{4})7 The later unified model can be read as an integrated parametric reformulation of the same plain-versus-boosted relation (Testolina et al., 2024).

A related adaptation appears in light-field quality assessment. There, pairwise counts are collected in a matrix crop(I,W4:3W4,H4:3H4)\mathrm{crop}(I,\tfrac{W}{4}:\tfrac{3W}{4},\tfrac{H}{4}:\tfrac{3H}{4})8, empirical probabilities are computed as crop(I,W4:3W4,H4:3H4)\mathrm{crop}(I,\tfrac{W}{4}:\tfrac{3W}{4},\tfrac{H}{4}:\tfrac{3H}{4})9, transformed to Lanczos2(crop;W,H)\mathrm{Lanczos}_2(\text{crop}; W,H)0-scores via Lanczos2(crop;W,H)\mathrm{Lanczos}_2(\text{crop}; W,H)1, and the scale values Lanczos2(crop;W,H)\mathrm{Lanczos}_2(\text{crop}; W,H)2 are then obtained by maximum-likelihood estimation under Case V assumptions. The final values are linearly rescaled to Lanczos2(crop;W,H)\mathrm{Lanczos}_2(\text{crop}; W,H)3 and inverted so that higher Lanczos2(crop;W,H)\mathrm{Lanczos}_2(\text{crop}; W,H)4 indicates better quality (Saraiva et al., 18 Sep 2025). The commonality across these variants is the treatment of subjective comparisons as noisy pairwise evidence about an underlying latent scale.

5. Confidence Intervals, Precision, and Reported Reliability

Confidence estimation in AIC-3 is bootstrap-based. In the HDR study, triplets were resampled with replacement Lanczos2(crop;W,H)\mathrm{Lanczos}_2(\text{crop}; W,H)5 times, the scale reconstruction was repeated, and at Lanczos2(crop;W,H)\mathrm{Lanczos}_2(\text{crop}; W,H)6 grid points in bitrate Lanczos2(crop;W,H)\mathrm{Lanczos}_2(\text{crop}; W,H)7 the distribution of Lanczos2(crop;W,H)\mathrm{Lanczos}_2(\text{crop}; W,H)8 was recorded; the Lanczos2(crop;W,H)\mathrm{Lanczos}_2(\text{crop}; W,H)9 and $3$00 percentiles of that distribution yielded the $3$01 confidence interval. The practical recommendation given in that work is to estimate $3$02 confidence intervals by nonparametric bootstrap with at least $3$03 resamples and to report rate-distortion curves with shaded confidence bands, including the average confidence-interval width at $3$04 JND (Jenadeleh et al., 14 Jun 2025).

The principal precision result reported for AIC-HDR2025 is that the average width of the $3$05 confidence interval around $3$06 JND is $3$07 JND, with the paper further stating that the confidence intervals are tightly bounded throughout a range from imperceptible distortion to clear distortion at $3$08–$3$09 JND. The authors interpret these narrow intervals as evidence that AIC-3 can resolve fractional JND steps even in HDR content where artifacts are extremely subtle (Jenadeleh et al., 14 Jun 2025).

Earlier SDR work reports the same bootstrap logic at larger scale. Testolina et al. process $3$10 bootstrap replicates end-to-end, resampling responses per triplet with replacement and obtaining $3$11 confidence bands on the final JND scale for each image and codec. The JPEG AI study uses $3$12 bootstrap resamples of the filtered PTC and BTC data, refits the model for each resample, and defines the $3$13 confidence interval at each bitrate from the $3$14th and $3$15th percentiles. Across the literature, the methodology therefore treats uncertainty as intrinsic to the reconstructed perceptual scale rather than as an afterthought attached only to summary statistics (Testolina et al., 2024, Jenadeleh et al., 7 Apr 2025).

A common oversimplification is to read a single JND score as an exact perceptual quantity. The published analyses do not support that interpretation. They report a mean scale value together with uncertainty, and in the objective-metric evaluation framework this uncertainty is elevated to a first-class variable through per-image $3$16 and $3$17, a high-fidelity subset defined by $3$18, and an outlier criterion based on $3$19 with $3$20 (Mohammadi et al., 16 Sep 2025). This suggests that the methodological unit of interest is not merely the point estimate, but the estimate together with its confidence structure.

6. Applications, Adaptations, and Relation to Objective Metric Benchmarking

AIC-3 has been used to build datasets and evaluation pipelines for several compression scenarios. The HDR study introduces AIC-HDR2025 as “the first such HDR dataset,” comprising $3$21 test images generated from five HDR sources, each compressed using four codecs at five compression levels, and reports $3$22 ratings from $3$23 participants across four fully controlled labs (Jenadeleh et al., 14 Jun 2025). The JPEG AI study constructs a dataset of $3$24 compressed images from five diverse sources, collects $3$25 triplet responses from $3$26 participants, and uses the resulting JND scales to compare JPEG AI with other codecs in the high-fidelity regime (Jenadeleh et al., 7 Apr 2025). Testolina et al. provide a $3$27-stimulus SDR dataset and code to reproduce the results (Testolina et al., 2024).

The methodology has also been adapted beyond conventional single-image coding. In the light-field study, the proposed JPEG AIC-3 test methodology is applied to low-distortion coded light fields with view synthesis. Two stimuli are displayed side by side, each alternating between an original and a coded view at $3$28 Hz, so that both sides flicker simultaneously; observers select the side whose flicker appears stronger. The test covers intra-method, cross-codec, encoding-method, bias-control, and attention-check triplets, with $3$29 unique triplets, $3$30 screened participants, $3$31 evaluations per stimulus pair, a $3$32-trial practice block, side inversion in $3$33 of presentations, and a mandatory mid-session rest (Saraiva et al., 18 Sep 2025). This does not redefine the core AIC-3 psychometric logic, but it demonstrates that the triplet-comparison framework can be specialized to synthesized-view artifacts.

AIC-3 is also used as a ground-truth substrate for objective image quality metric evaluation. In the metric-benchmarking study, subjective scale reconstruction on the JPEG AIC-3 dataset yields for each compressed image a mean $3$34 in JND and a standard deviation $3$35, enabling evaluation on the full range “All” and on high-fidelity and medium-fidelity subsets. Objective metric scores $3$36 are mapped to the JND domain through a four-parameter logistic,

$3$37

and then evaluated by RMSE, PLCC, SROCC, Kendall’s tau, Outlier Ratio, Perceptually-Weighted Rank Correlation, and the proposed Z-RMSE,

$3$38

The same study introduces the Meng–Rosenthal–Rubin test for comparing correlated SROCCs between metrics on the common subjective ground truth (Mohammadi et al., 16 Sep 2025). The JPEG AI study likewise uses a four-parameter logistic fit to align full-reference metrics to the JND scale and reports that CVVDP achieves the highest overall performance, while also noting that most metrics, including CVVDP, were overly optimistic for JPEG AI-compressed images (Jenadeleh et al., 7 Apr 2025).

Practical recommendations reported across these studies converge on several design principles: choose at least five distortion levels spanning $3$39–$3$40 JND per codec; include $3$41–$3$42 cross-codec triplets to align scales; permit “Not Sure”; filter unreliable observers using trap trials, mirror-pair consistency, or both; adopt the unified functional model $3$43 and $3$44 with Thurstone Case V MLE; bootstrap the entire pipeline; and report rate-distortion curves with confidence bands and per-codec or per-source breakdowns (Jenadeleh et al., 14 Jun 2025, Testolina et al., 2024). In that sense, JPEG AIC-3 is best understood not as a single display script, but as a standardized psychometric methodology for reconstructing fine-grained JND quality scales from carefully designed triplet judgments in the high-fidelity regime.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to JPEG AIC-3 Test Methodology.