JPEG AIC-3: High-Fidelity Image Quality Dataset
- JPEG AIC-3 Dataset is a high-fidelity image quality benchmark that uses triplet-comparison experiments to quantify subtle compression artifacts in JND units.
- It includes multiple releases like BTC/PTC SDR and JPEG AI SDR25, offering diverse codec images and extensive crowdsourced subjective ratings.
- The dataset enables robust benchmarking of objective image quality metrics, supporting standards development and precision codec optimization.
JPEG AIC-3 Dataset denotes the high-fidelity image-quality datasets and derivative releases developed around the JPEG AIC-3 activity for subjective assessment of compressed images in the visually near-lossless regime. Their defining features are triplet-comparison experiments, reconstruction of perceptual quality in Just Noticeable Difference (JND) units, and public benchmark material for codec evaluation and objective image-quality assessment; in current literature, the term covers the core SDR BTC/PTC dataset, the JPEG AI-specific SDR25 release, and later extensions such as AIC-HDR2025 (Testolina et al., 2024, Jenadeleh et al., 7 Apr 2025).
1. Standards context and scope
JPEG AIC-3 belongs to the ISO/IEC 29170 “Advanced Image Coding” series. In the literature summarized here, AIC-1 is described as guidelines for image coding system evaluation, AIC-2 as a flicker-based procedure for nearly lossless coding and JND-threshold estimation, and AIC-3 as the protocol and dataset activity for subjective quality assessment of high-fidelity images. Its explicit target is the quality range in which conventional subjective protocols based on absolute category ratings become insensitive to subtle impairments, while binary visually-lossless tests do not provide a continuous scale (Testolina et al., 2024, Mohammadi et al., 16 Sep 2025).
AIC-3 therefore replaces coarse category scores with fine-grained pairwise or triplet judgments and expresses the resulting perceptual scale in JND units. In the subsequent benchmarking literature, this JND-based formulation is used both for codec comparison and for evaluating whether objective image-quality metrics remain discriminative in the range below and around $1$ JND, which is the operational boundary between imperceptible or barely perceptible distortion and clearly visible impairment (Mohammadi et al., 16 Sep 2025).
The term “JPEG AIC-3 dataset” is not used in only one way. Some papers use it for the broader high-fidelity source/test material defined by the AIC-3 activity, while others use it for specific public releases derived from that material, such as the BTC/PTC crowdsourced dataset or the JPEG AI SDR25 subset. This usage reflects the fact that AIC-3 is simultaneously a test methodology, a source-image pool, and a set of released annotations and compressed stimuli (Testolina et al., 2024, Jenadeleh et al., 7 Apr 2025).
2. Core SDR composition and released variants
The core SDR AIC-3 study uses 5 source images selected from the JPEG AIC-3 dataset to represent a diverse range of image types and content. For subjective testing, each source is cropped to 620 × 800 pixels. In the original BTC/PTC release, each source is compressed with JPEG, JPEG 2000, VVC Intra, JPEG XL, and AVIF at 10 bitrates each, producing 250 decoded images (stimuli). These distortion levels correspond approximately to JND values equally spaced from 0.25 to 2.5, with level 0 denoting the source and level 10 the strongest artifacts (Testolina et al., 2024).
The later JPEG AI-focused release reuses five source images from the JPEG AIC-3 dataset—00002, 00006, 00007, 00009, and 00010—described as Human face, Scene with water, Night scene, Landscape, and Buildings. For crowdsourcing, one crop per source is again set to 620 × 800 pixels. Each source is compressed with JPEG AI at 10 distinct bitrates from 0.3 to 1.65 bpp with steps of 0.15 bpp, producing 50 compressed JPEG-AI images (Jenadeleh et al., 7 Apr 2025).
AIC-3 methodology has also been extended to HDR/WCG data. The HDR instance, AIC-HDR2025, contains 100 test images generated from five HDR sources, each compressed using four codecs at five compression levels. Although distinct from the SDR AIC-3 release, it preserves the same psychophysical design principles and is explicitly described as an HDR implementation of the JPEG AIC-3 methodology (Jenadeleh et al., 14 Jun 2025).
| Release | Content | Distinctive property |
|---|---|---|
| Core BTC/PTC SDR dataset (Testolina et al., 2024) | 5 source crops; 250 decoded images; 5 codecs × 10 levels | About 440,000 triplet responses |
| JPEG AI SDR25 (Jenadeleh et al., 7 Apr 2025) | 5 source crops; 50 JPEG AI images; 10 bitrates | 96,200 triplet responses from 459 participants |
| AIC-HDR2025 (Jenadeleh et al., 14 Jun 2025) | 5 HDR sources; 100 test images; 4 codecs × 5 levels | 34,560 ratings in four controlled labs |
The core SDR BTC/PTC release is announced at https://github.com/jpeg-aic/dataset-BTC-PTC-24, while the JPEG AI subset is released as https://github.com/jpeg-aic/dataset-JPEG-AI-SDR25 (Testolina et al., 2024, Jenadeleh et al., 7 Apr 2025).
3. JPEG AI-specific instantiation
The JPEG AI-specific AIC-3 dataset is tailored to high-fidelity evaluation of the JPEG AI verification model. It is explicitly framed as a comprehensive subjective visual quality assessment of JPEG AI-compressed images using the JPEG AIC-3 methodology, with perceptual differences quantified in JND units. The abstract reports 50 compressed images with fine-grained distortion levels from five diverse sources, 96,200 triplet responses, and 459 participants (Jenadeleh et al., 7 Apr 2025).
For this release, JPEG AI is operated in high operating point mode. The technical configuration is stated as YUV444, BT.709, advanced analysis/synthesis transforms with attention (IDs 0/2), and all content-adaptive tools, including post-processing filters and adaptive scaling of residuals, enabled. The software reference is Verification model VM7.0. The bitrate span from 0.3 to 1.65 bpp was chosen so that JPEG AI lies in the same perceived distortion range as high-quality JPEG, JPEG 2000, AVIF, VVC intra, and JPEG XL in the preceding AIC-3 study, with distortion levels approximately spanning 0.25–2.5 JND (Jenadeleh et al., 7 Apr 2025).
The crowdsourcing design separates Boosted Triplet Comparisons (BTC) from Plain Triplet Comparisons (PTC). In the JPEG AI study, BTC contains 660 triplets split into five batches of 132 questions, and PTC contains 180 triplets split into two batches of 90 questions. The response density is intentionally asymmetric: 120 responses per triplet in BTC and 49 responses per triplet in PTC. The worker counts are 386 for BTC and 73 for PTC, with subjects allowed to complete up to two batches and question order randomized (Jenadeleh et al., 7 Apr 2025).
The resulting dataset is a JPEG AI-specific JND benchmark rather than a generic codec corpus. Its primary use is to anchor JPEG AI images on the same perceptual scale as the traditional codecs from the original AIC-3 SDR study, and to expose codec-dependent biases of objective metrics in the high-fidelity range (Jenadeleh et al., 7 Apr 2025).
4. Subjective protocol and JND scale reconstruction
The AIC-3 protocol is built around triplets , where is the source image and are compressed variants. In Boosted Triplet Comparison, the two test images are shown side by side and each alternates with the source at 10 Hz, using 100 ms source and 100 ms distorted frames. Boosting combines three operations: zooming, in which images are cropped to half the size in each dimension and upscaled with Lanczos resampling; artifact amplification, in which the pixel-wise difference between original and distorted image is scaled by a factor of 2; and flicker, which converts small residual differences into a strong or weak temporal shimmer. In Plain Triplet Comparison, no boosting is applied; observers toggle between source and distorted image, must toggle at least once, and are limited to 2 Hz toggling within a 30 s response window (Testolina et al., 2024).
The original SDR study includes multiple question types: same-codec comparisons, cross-codec comparisons, bias-checking comparisons, and trap questions. BTC contains 2750 same-codec, 550 cross-codec, 100 bias-checking, and 200 trap questions, for 3600 questions in total; PTC contains 750 same-codec, 150 cross-codec, 50 bias-checking, and 100 trap questions, for 1050 questions in total. About 440,000 triplet question responses were collected on Amazon Mechanical Turk, after which 615 BTC batches from 423 subjects and 260 PTC batches from 208 subjects were retained by reliability filtering (Testolina et al., 2024).
The reconstruction of perceptual quality uses a Thurstonian Case V model. “Not sure” responses are split half-half between left and right, which turns triplet judgments into a two-alternative forced-choice representation. In the original formulation, JND conversion is anchored by the condition that a distortion difference of 1 JND corresponds to 75% correct discrimination under the scaling model. In the later unified JPEG AI formulation, the non-boosted distortion-rate relation is modeled as
and the mapping from plain to boosted distortion as
The coefficients are estimated by maximum likelihood estimation (MLE) from joint BTC/PTC data, and confidence intervals are obtained by 1000 bootstrap samples (Jenadeleh et al., 7 Apr 2025).
A notable property of the JPEG AI release is its precision. The reported confidence interval width for an image at JND is smaller than $0.1 + 0.05x$ for every codec including JPEG AI and for every source. In the HDR extension, the same methodology yields 95% confidence intervals averaging a width of 0.27 at 1 JND, which further supports the claim that AIC-3 is optimized for the high-fidelity regime rather than for coarse impairment grading (Jenadeleh et al., 7 Apr 2025, Jenadeleh et al., 14 Jun 2025).
5. Benchmarking objective quality metrics
A central use of the JPEG AIC-3 dataset is benchmarking full-reference IQA metrics against JND-scale ground truth. The JPEG AI SDR25 study evaluates 15 metrics, including PSNR-Y, SSIM, MS-SSIM, PSNR-HVS, IW-SSIM, GMSD, NLPD, Butteraugli-pnorm, SSIMULACRA1, SSIMULACRA2, VMAF, VMAF-neg, HDR-VDP-2 Q, HDR-VDP-3 Q, and CVVDP. The later large-scale evaluation broadens this to conventional, fusion, and learning-based metrics, including PSNRY, PSNR-HVS, SSIM, MS-SSIM, IW-SSIM, FSIM, FSIMc, GMSD, UQI, NLPD, VIF, HaarPSI, VSI, FLIP, CVVDP, HDR-VDP-2 Q, HDR-VDP-3 Q, BUTTERAUGLI, SSIMULACRA1/2, VMAF, VMAF-NEG, LPIPS, DISTS, A-DISTS, PieAPP, and TOPIQ (Jenadeleh et al., 7 Apr 2025, Mohammadi et al., 16 Sep 2025).
The later evaluation treats the AIC-3 benchmark as 300 compressed images and partitions it into the full range All with JND scores spanning approximately 0 to 3.8 JND, a High-Fidelity (HF) subset JND with 115 scores, and a Medium-Fidelity (MF) subset with 185 scores. This split is important because metric saturation is most severe in HF, where output values often become nearly constant despite perceptually meaningful changes (Mohammadi et al., 16 Sep 2025).
The JPEG AI-specific study reports that CVVDP achieves the overall highest performance across all codecs, with overall PLCC 0.960 and overall SRCC 0.961, while VMAF-neg attains the highest correlation on JPEG AI only, with PLCC 0.958 and SRCC 0.963. At the same time, most metrics, including CVVDP, are described as overly optimistic for JPEG AI-compressed images: for a given subjective JND level, they assign better quality scores to JPEG AI than to traditional codecs. A stated explanation is that existing metrics were not designed for the artifact structures produced by learning-based compression (Jenadeleh et al., 7 Apr 2025).
The later study introduces uncertainty-aware evaluation via Z-RMSE,
0
where 1 is the subjective JND mean, 2 the transformed objective score, and 3 the subjective standard deviation. It also applies the Meng–Rosenthal–Rubin test and the Wilcoxon Signed-Rank test for pairwise metric comparison. In these analyses, CVVDP is reported as statistically better than every other metric, and IW-SSIM as the next-best performer. The same study also finds that about 50% of metrics show no statistically significant difference between evaluation on full-resolution images and evaluation on the subjective-test crops, although using the exact crops is slightly better when differences do exist (Mohammadi et al., 16 Sep 2025).
6. Extensions, ambiguities, and limitations
The AIC-3 methodology has already been generalized beyond SDR. AIC-HDR2025 is described as the first such HDR dataset for fine-grained HDR image quality assessment from noticeably distorted to very high fidelity. It comprises 100 test images from five HDR sources, compressed with JPEG XT, JPEG XL, JPEG AI, and AVIF at five levels, and collects 34,560 ratings from 151 participants in four fully controlled laboratories. This extension confirms that AIC-3 is not limited to SDR photographic content, but a transferable psychophysical framework for codec evaluation in the high-fidelity regime (Jenadeleh et al., 14 Jun 2025).
A persistent ambiguity in the literature concerns the relation between the high-fidelity AIC-3 datasets and the separate JPEG AI Image Compression Visual Artifacts Dataset. The latter processes about 350,000 unique images from Open Images and produces 46,440 crowd-validated artifact examples for the JPEG AI verification model, with labels for texture and boundary degradation, color change, and text corruption. Its purpose is detection, localization, and quantification of localized neural-compression failures relative to HM-18.0, not JND-scale subjective quality reconstruction. The paper itself states that it does not literally use the term “AIC-3”, even though standards-adjacent discussion may treat it as closely aligned with later JPEG AI/AIC testing (Tsereh et al., 2024).
The core AIC-3 SDR benchmark is deliberately small. It uses 5 source images and 620 × 800 crops, and the later metric-evaluation study explicitly characterizes it as “a deliberately small but very carefully constructed benchmark.” This small-4, high-precision design is a strength for psychometric scale reconstruction but also a limitation for broad content coverage. A plausible implication is that AIC-3 should be read as a precision benchmark for near-lossless perception rather than as a general-purpose image corpus. The same papers also show that crop-based subjective testing introduces a full-image versus crop mismatch for objective metrics, although the measured impact is usually small (Mohammadi et al., 16 Sep 2025).
In practice, the JPEG AIC-3 dataset is most significant where compression artifacts are subtle, subjective uncertainty matters, and objective metrics must be evaluated below or around the JND threshold. That includes standards development, high-fidelity codec tuning, metric design, and any application in which perceptually minor errors remain operationally important (Testolina et al., 2024, Mohammadi et al., 16 Sep 2025).