---
title: 'JPEG AIC-3: High-Fidelity Image Quality Dataset'
url: https://www.emergentmind.com/topics/jpeg-aic-3-dataset
type: topic
---

# JPEG AIC-3: High-Fidelity Image Quality Dataset

JPEG AIC-3 Dataset denotes the high-fidelity image-quality datasets and derivative releases developed around the JPEG AIC-3 activity for subjective assessment of compressed images in the visually near-lossless regime. Their defining features are triplet-comparison experiments, reconstruction of perceptual quality in Just Noticeable Difference (JND) units, and public benchmark material for codec evaluation and objective image-quality assessment; in current literature, the term covers the core SDR BTC/PTC dataset, the JPEG AI-specific SDR25 release, and later extensions such as AIC-HDR2025 [2410.09501][2504.06301].

## 1. Standards context and scope

JPEG AIC-3 belongs to the ISO/IEC 29170 “Advanced Image Coding” series. In the literature summarized here, AIC-1 is described as guidelines for image coding system evaluation, AIC-2 as a flicker-based procedure for nearly lossless coding and JND-threshold estimation, and AIC-3 as the protocol and dataset activity for subjective quality assessment of high-fidelity images. Its explicit target is the quality range in which conventional subjective protocols based on absolute category ratings become insensitive to subtle impairments, while binary visually-lossless tests do not provide a continuous scale [2410.09501][2509.13150].

AIC-3 therefore replaces coarse category scores with fine-grained pairwise or triplet judgments and expresses the resulting perceptual scale in JND units. In the subsequent benchmarking literature, this JND-based formulation is used both for codec comparison and for evaluating whether objective image-quality metrics remain discriminative in the range below and around \(1\) JND, which is the operational boundary between imperceptible or barely perceptible distortion and clearly visible impairment [2509.13150].

The term “JPEG AIC-3 dataset” is not used in only one way. Some papers use it for the broader high-fidelity source/test material defined by the AIC-3 activity, while others use it for specific public releases derived from that material, such as the BTC/PTC crowdsourced dataset or the JPEG AI SDR25 subset. This usage reflects the fact that AIC-3 is simultaneously a test methodology, a source-image pool, and a set of released annotations and compressed stimuli [2410.09501][2504.06301].

## 2. Core SDR composition and released variants

The core SDR AIC-3 study uses **5** source images selected from the JPEG AIC-3 dataset to represent a diverse range of image types and content. For subjective testing, each source is cropped to **620 × 800 pixels**. In the original BTC/PTC release, each source is compressed with **JPEG**, **JPEG 2000**, **VVC Intra**, **JPEG XL**, and **AVIF** at **10** bitrates each, producing **250 decoded images (stimuli)**. These distortion levels correspond approximately to JND values equally spaced from **0.25 to 2.5**, with level **0** denoting the source and level **10** the strongest artifacts [2410.09501].

The later JPEG AI-focused release reuses **five source images** from the JPEG AIC-3 dataset—**00002**, **00006**, **00007**, **00009**, and **00010**—described as **Human face**, **Scene with water**, **Night scene**, **Landscape**, and **Buildings**. For crowdsourcing, one crop per source is again set to **620 × 800 pixels**. Each source is compressed with JPEG AI at **10 distinct bitrates** from **0.3 to 1.65 bpp** with steps of **0.15 bpp**, producing **50 compressed JPEG-AI images** [2504.06301].

AIC-3 methodology has also been extended to HDR/WCG data. The HDR instance, AIC-HDR2025, contains **100 test images** generated from **five HDR sources**, each compressed using **four codecs** at **five compression levels**. Although distinct from the SDR AIC-3 release, it preserves the same psychophysical design principles and is explicitly described as an HDR implementation of the JPEG AIC-3 methodology [2506.12505].

| Release | Content | Distinctive property |
|---|---|---|
| Core BTC/PTC SDR dataset [2410.09501] | 5 source crops; 250 decoded images; 5 codecs × 10 levels | About 440,000 triplet responses |
| JPEG AI SDR25 [2504.06301] | 5 source crops; 50 JPEG AI images; 10 bitrates | 96,200 triplet responses from 459 participants |
| AIC-HDR2025 [2506.12505] | 5 HDR sources; 100 test images; 4 codecs × 5 levels | 34,560 ratings in four controlled labs |

The core SDR BTC/PTC release is announced at `https://github.com/jpeg-aic/dataset-BTC-PTC-24`, while the JPEG AI subset is released as `https://github.com/jpeg-aic/dataset-JPEG-AI-SDR25` [2410.09501][2504.06301].

## 3. JPEG AI-specific instantiation

The JPEG AI-specific AIC-3 dataset is tailored to high-fidelity evaluation of the JPEG AI verification model. It is explicitly framed as a comprehensive subjective visual quality assessment of JPEG AI-compressed images using the JPEG AIC-3 methodology, with perceptual differences quantified in JND units. The abstract reports **50 compressed images with fine-grained distortion levels from five diverse sources**, **96,200 triplet responses**, and **459 participants** [2504.06301].

For this release, JPEG AI is operated in **high operating point** mode. The technical configuration is stated as **YUV444**, **BT.709**, **advanced analysis/synthesis transforms with attention (IDs 0/2)**, and **all content-adaptive tools, including post-processing filters and adaptive scaling of residuals, enabled**. The software reference is **Verification model VM7.0**. The bitrate span from **0.3 to 1.65 bpp** was chosen so that JPEG AI lies in the same perceived distortion range as high-quality JPEG, JPEG 2000, AVIF, VVC intra, and JPEG XL in the preceding AIC-3 study, with distortion levels approximately spanning **0.25–2.5 JND** [2504.06301].

The crowdsourcing design separates **Boosted Triplet Comparisons (BTC)** from **Plain Triplet Comparisons (PTC)**. In the JPEG AI study, BTC contains **660 triplets** split into **five batches of 132 questions**, and PTC contains **180 triplets** split into **two batches of 90 questions**. The response density is intentionally asymmetric: **120 responses per triplet** in BTC and **49 responses per triplet** in PTC. The worker counts are **386** for BTC and **73** for PTC, with subjects allowed to complete up to two batches and question order randomized [2504.06301].

The resulting dataset is a JPEG AI-specific JND benchmark rather than a generic codec corpus. Its primary use is to anchor JPEG AI images on the same perceptual scale as the traditional codecs from the original AIC-3 SDR study, and to expose codec-dependent biases of objective metrics in the high-fidelity range [2504.06301].

## 4. Subjective protocol and JND scale reconstruction

The AIC-3 protocol is built around triplets \((I_i, I_0, I_k)\), where \(I_0\) is the source image and \(I_i, I_k\) are compressed variants. In **Boosted Triplet Comparison**, the two test images are shown side by side and each alternates with the source at **10 Hz**, using **100 ms** source and **100 ms** distorted frames. Boosting combines three operations: **zooming**, in which images are cropped to half the size in each dimension and upscaled with **Lanczos** resampling; **artifact amplification**, in which the pixel-wise difference between original and distorted image is scaled by a factor of **2**; and **flicker**, which converts small residual differences into a strong or weak temporal shimmer. In **Plain Triplet Comparison**, no boosting is applied; observers toggle between source and distorted image, must toggle at least once, and are limited to **2 Hz** toggling within a **30 s** response window [2410.09501].

The original SDR study includes multiple question types: **same-codec comparisons**, **cross-codec comparisons**, **bias-checking comparisons**, and **trap questions**. BTC contains **2750** same-codec, **550** cross-codec, **100** bias-checking, and **200** trap questions, for **3600** questions in total; PTC contains **750** same-codec, **150** cross-codec, **50** bias-checking, and **100** trap questions, for **1050** questions in total. About **440,000 triplet question responses** were collected on Amazon Mechanical Turk, after which **615 BTC batches from 423 subjects** and **260 PTC batches from 208 subjects** were retained by reliability filtering [2410.09501].

The reconstruction of perceptual quality uses a **Thurstonian Case V** model. “Not sure” responses are split half-half between left and right, which turns triplet judgments into a two-alternative forced-choice representation. In the original formulation, JND conversion is anchored by the condition that a distortion difference of **1 JND** corresponds to **75%** correct discrimination under the scaling model. In the later unified JPEG AI formulation, the non-boosted distortion-rate relation is modeled as
$$
d(r) = \alpha \exp(-\beta r),
$$
and the mapping from plain to boosted distortion as
$$
t(d) = \gamma_1 d + \gamma_2 d^2.
$$
The coefficients are estimated by **maximum likelihood estimation (MLE)** from joint BTC/PTC data, and confidence intervals are obtained by **1000 bootstrap samples** [2504.06301].

A notable property of the JPEG AI release is its precision. The reported confidence interval width for an image at \(x\) JND is smaller than \(0.1 + 0.05x\) for every codec including JPEG AI and for every source. In the HDR extension, the same methodology yields **95% confidence intervals averaging a width of 0.27 at 1 JND**, which further supports the claim that AIC-3 is optimized for the high-fidelity regime rather than for coarse impairment grading [2504.06301][2506.12505].

## 5. Benchmarking objective quality metrics

A central use of the JPEG AIC-3 dataset is benchmarking full-reference IQA metrics against JND-scale ground truth. The JPEG AI SDR25 study evaluates **15** metrics, including **PSNR-Y**, **SSIM**, **MS-SSIM**, **PSNR-HVS**, **IW-SSIM**, **GMSD**, **NLPD**, **Butteraugli-pnorm**, **SSIMULACRA1**, **SSIMULACRA2**, **VMAF**, **VMAF-neg**, **HDR-VDP-2 Q**, **HDR-VDP-3 Q**, and **CVVDP**. The later large-scale evaluation broadens this to conventional, fusion, and learning-based metrics, including **PSNRY**, **PSNR-HVS**, **SSIM**, **MS-SSIM**, **IW-SSIM**, **FSIM**, **FSIMc**, **GMSD**, **UQI**, **NLPD**, **VIF**, **HaarPSI**, **VSI**, **FLIP**, **CVVDP**, **HDR-VDP-2 Q**, **HDR-VDP-3 Q**, **BUTTERAUGLI**, **SSIMULACRA1/2**, **VMAF**, **VMAF-NEG**, **LPIPS**, **DISTS**, **A-DISTS**, **PieAPP**, and **TOPIQ** [2504.06301][2509.13150].

The later evaluation treats the AIC-3 benchmark as **300 compressed images** and partitions it into the full range **All** with JND scores spanning approximately **0 to 3.8 JND**, a **High-Fidelity (HF)** subset \([0,1]\) JND with **115** scores, and a **Medium-Fidelity (MF)** subset \([1,+\infty)\) with **185** scores. This split is important because metric saturation is most severe in HF, where output values often become nearly constant despite perceptually meaningful changes [2509.13150].

The JPEG AI-specific study reports that **CVVDP** achieves the overall highest performance across all codecs, with **overall PLCC 0.960** and **overall SRCC 0.961**, while **VMAF-neg** attains the highest correlation on **JPEG AI only**, with **PLCC 0.958** and **SRCC 0.963**. At the same time, most metrics, including CVVDP, are described as **overly optimistic** for JPEG AI-compressed images: for a given subjective JND level, they assign better quality scores to JPEG AI than to traditional codecs. A stated explanation is that existing metrics were not designed for the artifact structures produced by learning-based compression [2504.06301].

The later study introduces uncertainty-aware evaluation via **Z-RMSE**,
$$
\text{Z-RMSE} = \sqrt{\frac{1}{n}\sum_{i=1}^{n}\left(\frac{S_{\text{trans},i}-S_{\text{subj},i}}{\sigma_i}\right)^2},
$$
where \(S_{\text{subj},i}\) is the subjective JND mean, \(S_{\text{trans},i}\) the transformed objective score, and \(\sigma_i\) the subjective standard deviation. It also applies the **Meng–Rosenthal–Rubin** test and the **Wilcoxon Signed-Rank** test for pairwise metric comparison. In these analyses, **CVVDP** is reported as statistically better than every other metric, and **IW-SSIM** as the next-best performer. The same study also finds that about **50%** of metrics show no statistically significant difference between evaluation on full-resolution images and evaluation on the subjective-test crops, although using the exact crops is slightly better when differences do exist [2509.13150].

## 6. Extensions, ambiguities, and limitations

The AIC-3 methodology has already been generalized beyond SDR. AIC-HDR2025 is described as the first such HDR dataset for fine-grained HDR image quality assessment from noticeably distorted to very high fidelity. It comprises **100** test images from **five HDR sources**, compressed with **JPEG XT**, **JPEG XL**, **JPEG AI**, and **AVIF** at **five** levels, and collects **34,560 ratings** from **151 participants** in **four fully controlled laboratories**. This extension confirms that AIC-3 is not limited to SDR photographic content, but a transferable psychophysical framework for codec evaluation in the high-fidelity regime [2506.12505].

A persistent ambiguity in the literature concerns the relation between the high-fidelity AIC-3 datasets and the separate **JPEG AI Image Compression Visual Artifacts Dataset**. The latter processes about **350,000 unique images** from **Open Images** and produces **46,440** crowd-validated artifact examples for the JPEG AI verification model, with labels for **texture and boundary degradation**, **color change**, and **text corruption**. Its purpose is detection, localization, and quantification of localized neural-compression failures relative to **HM-18.0**, not JND-scale subjective quality reconstruction. The paper itself states that it does not literally use the term “AIC-3”, even though standards-adjacent discussion may treat it as closely aligned with later JPEG AI/AIC testing [2411.06810].

The core AIC-3 SDR benchmark is deliberately small. It uses **5** source images and **620 × 800** crops, and the later metric-evaluation study explicitly characterizes it as “a deliberately small but very carefully constructed benchmark.” This small-\(N\), high-precision design is a strength for psychometric scale reconstruction but also a limitation for broad content coverage. A plausible implication is that AIC-3 should be read as a precision benchmark for near-lossless perception rather than as a general-purpose image corpus. The same papers also show that crop-based subjective testing introduces a full-image versus crop mismatch for objective metrics, although the measured impact is usually small [2509.13150].

In practice, the JPEG AIC-3 dataset is most significant where compression artifacts are subtle, subjective uncertainty matters, and objective metrics must be evaluated below or around the JND threshold. That includes standards development, high-fidelity codec tuning, metric design, and any application in which perceptually minor errors remain operationally important [2410.09501][2509.13150].

Source: https://www.emergentmind.com/topics/jpeg-aic-3-dataset