Papers
Topics
Authors
Recent
Search
2000 character limit reached

RawGen: Diffusion-based Raw Image Generation

Updated 4 July 2026
  • RawGen is a diffusion-based framework that transforms text prompts and sRGB inputs into physically meaningful, scene-referred raw outputs.
  • It combines conditional denoising with a specialized decoder and deterministic camera calibration to produce precise, camera-specific raw representations.
  • The framework enables sRGB inversion and many-to-one supervision to robustly suppress photo-finishing variability for improved downstream low-level vision tasks.

Searching arXiv for “RawGen” and closely related papers to ground the article with current citations. {"query":"ti:RawGen OR all:RawGen", "max_results": 10} Searching for inverse-ISP and raw image generation context to support positioning and related-work framing. {"query":"all:\"inverse ISP\" OR all:\"raw image generation\" OR all:\"camera raw\"", "max_results": 10} RawGen is a diffusion-based framework for learning camera raw image generation from either text prompts or sRGB inputs, producing scene-referred linear outputs in CIE XYZ and camera-specific raw representations for arbitrary target cameras (Kim et al., 31 Mar 2026). It is presented as, to the authors’ knowledge, the first diffusion-based framework enabling text-to-raw generation for arbitrary target cameras, alongside sRGB-to-raw inversion. Its central premise is that large-scale sRGB diffusion priors can be repurposed into physically meaningful linear image generation by combining conditional denoising, a specialized decoder, and deterministic camera calibration derived from DNG metadata (Kim et al., 31 Mar 2026).

1. Problem setting and nomenclature

RawGen is motivated by the distinction between scene-referred raw images and display-referred sRGB images. Scene-referred raw data retains linear radiometric information about the scene and is the appropriate domain for denoising, demosaicing, HDR reconstruction, illuminant estimation, and learning ISPs, because white balance, exposure scaling, and tone mapping have physically meaningful interpretations only on linear signals. By contrast, sRGB outputs are 8-bit, nonlinear, and already entangled with gamma, tone curves, contrast, color styling, and other photo-finishing operations. Those properties make sRGB a poor substrate for physics-aware image analysis and restoration (Kim et al., 31 Mar 2026).

The practical obstacle is dataset scarcity. Large-scale raw datasets are limited, camera-specific, and coupled to proprietary in-camera ISP behavior. As a result, each new sensor or ISP variant typically requires new capture and relabeling, while existing diffusion models remain optimized for photo-finished sRGB imagery rather than scene-referred linear representations. Classical inverse-ISP methods assume a fixed pipeline and therefore struggle when the input has passed through diverse, unknown, and nonlinear photo-finishing stages (Kim et al., 31 Mar 2026).

The name “RawGen” also appears in unrelated literature. In one large-systems paper, “RawGen” is interpreted as Retrieval-Augmented Generation rather than raw-image generation, so the term is not unique across arXiv usage (Naikov et al., 23 Jan 2025).

2. Core objectives and conceptual contributions

RawGen repurposes a large sRGB text-to-image prior to produce physically meaningful scene-referred outputs. Its first contribution is text-to-raw generation for arbitrary cameras: the framework generates linear CIE XYZ first, then deterministically maps that output to a target camera raw-RGB space using calibration metadata such as ForwardMatrix and ColorMatrix from DNG files. This design makes the linear synthesis camera-agnostic while still allowing camera-specific raw outputs without retraining (Kim et al., 31 Mar 2026).

Its second contribution is sRGB-to-raw inversion. Given an sRGB image containing unknown photo-finishing, RawGen uses image conditioning to invert the input to a canonical linear XYZ anchor, after which the result can be mapped to a device raw space if desired. The framework therefore addresses both generative and inverse formulations within the same latent-space pipeline (Kim et al., 31 Mar 2026).

A third contribution is the many-to-one inverse-ISP dataset. For each underlying scene, multiple sRGB renditions generated with diverse ISP and photo-finishing parameters are anchored to a single scene-referred target. This supervision explicitly teaches the model to suppress stylistic variability and recover the common linear representation. The authors describe the resulting outputs as camera-centric linear reconstructions that outperform fixed-ISP inverse methods on heterogeneous inputs (Kim et al., 31 Mar 2026).

A fourth contribution is methodological rather than purely architectural: RawGen is designed as a data source for downstream low-level vision. Text-driven synthetic raw generation is used to scale training data for illuminant estimation, raw denoising, and neural ISP training, with reported gains over prior synthetic sources (Kim et al., 31 Mar 2026).

3. Architecture and learning formulation

RawGen builds on a pretrained rectified-flow DiT with native image conditioning, specifically the FLUX.1-Kontext backbone. The framework operates in latent space. A frozen VAE encoder EVAEE_{\mathrm{VAE}} encodes sRGB inputs into latents, while a specialized VAE decoder DVAED_{\mathrm{VAE}} is fine-tuned so that it decodes latents into linear CIE XYZ images rather than sRGB (Kim et al., 31 Mar 2026).

For sRGB-to-raw inversion, the input image is encoded into zsRGBz_{\mathrm{sRGB}}, and those conditioning tokens are concatenated with noisy target tokens. The combined sequence is then jointly processed through the DiT with LoRA adapters on attention projections; only the target tokens are supervised. For text-to-raw generation, the framework uses the base model’s text-to-latent path to obtain zsRGBz_{\mathrm{sRGB}}, and then runs the same conditional generation procedure to produce an XYZ latent. In both modes, the latent prediction stage is followed by XYZ decoding and then by deterministic camera-specific mapping (Kim et al., 31 Mar 2026).

The denoiser is trained with a rectified-flow vv-prediction objective. Let ϵ∼N(0,I)\epsilon \sim \mathcal{N}(0, I) and define the straight-line noising path

zt=(1−t) zXYZ+t ϵ,t∈[0,1].z_t = (1 - t)\, z_{\mathrm{XYZ}} + t\, \epsilon, \quad t \in [0,1].

The ground-truth velocity is

vgt=ϵ−zXYZ,v_{\mathrm{gt}} = \epsilon - z_{\mathrm{XYZ}},

and the denoising loss is

Ldenoise=Ek, t, ϵ∥vgt−vθ(zt, t; zsRGB(k))∥22.\mathcal{L}_{\mathrm{denoise}} = \mathbb{E}_{k,\, t,\, \epsilon} \left\| v_{\mathrm{gt}} - v_{\theta}(z_t,\, t;\, z_{\mathrm{sRGB}^{(k)}}) \right\|_2^2.

The decoder is fine-tuned with an L1L_1 reconstruction objective,

DVAED_{\mathrm{VAE}}0

which retargets the decoder from the sRGB domain to the XYZ domain while preserving spatial representation capacity (Kim et al., 31 Mar 2026).

4. Many-to-one inverse-ISP dataset and forward imaging model

The many-to-one dataset is central to RawGen’s invariance claim. For each raw scene, a canonical scene-referred target DVAED_{\mathrm{VAE}}1 in CIE XYZ is derived by applying per-scene white balance and camera-to-XYZ conversion before any sRGB rendering. Multiple sRGB variants are then produced by varying white balance gains, tone mapping curves, and contrast around realistic ranges. The training set is therefore

DVAED_{\mathrm{VAE}}2

where each DVAED_{\mathrm{VAE}}3 is a different photo-finished rendition of the same underlying scene (Kim et al., 31 Mar 2026).

The representative forward ISP used to motivate the inverse problem begins from camera raw with black level DVAED_{\mathrm{VAE}}4 and white balance gains DVAED_{\mathrm{VAE}}5:

DVAED_{\mathrm{VAE}}6

where DVAED_{\mathrm{VAE}}7 is demosaicing, DVAED_{\mathrm{VAE}}8 is a tone map, and DVAED_{\mathrm{VAE}}9 is the sRGB gamma or OETF. Optional modules include denoising, sharpening, gamut mapping, and contrast or saturation adjustments (Kim et al., 31 Mar 2026).

In the reproducibility details, the randomized photo-finishing process is specified more concretely. White balance uses red and blue gains sampled as zsRGBz_{\mathrm{sRGB}}0, with green fixed at zsRGBz_{\mathrm{sRGB}}1. Tone mapping is applied per channel as

zsRGBz_{\mathrm{sRGB}}2

and global contrast uses

zsRGBz_{\mathrm{sRGB}}3

This construction explicitly exposes the model to ISP and photo-finishing diversity while keeping the target scene-referred representation fixed (Kim et al., 31 Mar 2026).

5. Camera-specific mapping and inference modes

RawGen separates canonical linear generation from device-specific rendering by decoding into CIE XYZ first and only then mapping into the target camera’s raw-RGB space. The mapping uses DNG calibration metadata and interpolates matrices between two reference illuminants:

zsRGBz_{\mathrm{sRGB}}4

zsRGBz_{\mathrm{sRGB}}5

The decoded XYZ image is mapped to white-balanced camera RGB by

zsRGBz_{\mathrm{sRGB}}6

the illuminant is converted as

zsRGBz_{\mathrm{sRGB}}7

and device raw-RGB is obtained through

zsRGBz_{\mathrm{sRGB}}8

Optional heteroscedastic noise and CFA remosaicing can then be applied for more realistic raw outputs (Kim et al., 31 Mar 2026).

The sRGB-to-raw inversion mode takes an input image zsRGBz_{\mathrm{sRGB}}9, target camera zsRGBz_{\mathrm{sRGB}}0, and optional text prompt zsRGBz_{\mathrm{sRGB}}1, and first encodes the sRGB image as

zsRGBz_{\mathrm{sRGB}}2

It then integrates the reverse rectified-flow ODE conditioned on zsRGBz_{\mathrm{sRGB}}3:

zsRGBz_{\mathrm{sRGB}}4

decodes the scene-referred XYZ image, and maps it to the target camera’s raw-RGB. The text-to-raw mode replaces the image encoder stage with the base model’s text-to-latent path,

zsRGBz_{\mathrm{sRGB}}5

after which the same conditional generation, decoding, and XYZ-to-raw mapping are used (Kim et al., 31 Mar 2026).

The implementation is intentionally lightweight in calibration requirements. The reported inputs are an sRGB image or a text prompt together with the target camera’s DNG metadata, and the paper states that using a single DNG file of the target camera suffices to obtain the required matrices (Kim et al., 31 Mar 2026).

6. Training setup and empirical evaluation

The training data combines the MIT-Adobe FiveK and RAISE raw DNG collections. XYZ anchors are computed using DNG AsShotNeutral and ForwardMatrix, while multiple sRGB variants are rendered with a physically grounded software ISP and randomized photo-finishing. The denoiser uses the FLUX.1-Kontext DiT backbone with LoRA of rank zsRGBz_{\mathrm{sRGB}}6 and zsRGBz_{\mathrm{sRGB}}7 on attention projections; the decoder is fine-tuned to XYZ using zsRGBz_{\mathrm{sRGB}}8 loss. Training uses zsRGBz_{\mathrm{sRGB}}9 crops, with XYZ anchors stored as 16-bit PNG and sRGB variants as 8-bit PNG (Kim et al., 31 Mar 2026).

In the many-to-one invertability evaluation based on FiveK expert-retouched variations from editors A–E, RawGen is reported to achieve the best CIE XYZ reconstruction across all five styles against CIE XYZ Net, InvISP, and Raw-Diffusion. The reported PSNR/SSIM examples are A: 23.20/0.8432, B: 24.35/0.8581, C: 23.37/0.8387, D: 23.51/0.8531, and E: 23.89/0.8500. Typical baselines are summarized as approximately 19–21 dB PSNR and approximately 0.78–0.84 SSIM, and an ablated one-to-one RawGen is described as substantially worse, which the authors use to highlight the role of many-to-one training (Kim et al., 31 Mar 2026).

The suppression of photo-finishing variability is also evaluated in latent space using PCA, t-SNE, and UMAP compactness over 100 graded variants per prompt. Mean distance to centroid is reported as PCA 160.0, t-SNE 10.69, and UMAP 1.067 for RawGen, compared with larger values for alternatives such as XYZNet and Raw-Diffusion. The paper interprets this as evidence that RawGen more strongly suppresses photo-finishing variability while preserving a common scene-referred anchor (Kim et al., 31 Mar 2026).

For device-specific synthesis, the decoded XYZ outputs are mapped to the Samsung Galaxy S24 main camera raw-RGB space with optional noise. Pre-trained neural ISPs trained only on real S24 data produce plausible sRGB from RawGen raw inputs without retraining, which the paper presents as an indicator of distribution alignment between the synthetic raw outputs and real-device raw data (Kim et al., 31 Mar 2026).

7. Downstream use, limitations, and significance

A major claim of RawGen is that synthetic raw data can improve downstream low-level vision systems. Using 3K generated samples and evaluating on real test splits, the paper reports gains over Graphics2RAW in three tasks. For illuminant estimation on NUS-8 with nine cameras, Graphics2RAW yields mean 4.21°, median 3.38°, and worst 25% 8.57°, whereas RawGen yields mean 3.14°, median 2.11°, and worst 25% 7.37°, approaching the real-data model at mean 3.02°, median 2.17°, and worst 25% 6.77°. For neural ISP training on a nighttime dataset, Graphics2RAW gives PSNR 38.10, SSIM 0.974, and vv0 2.301, while RawGen gives PSNR 38.42, SSIM 0.970, and vv1 2.183, comparable to training on real raw at PSNR 38.32, SSIM 0.974, and vv2 2.133. For raw denoising, RawGen reports 50.63/0.994 at ISO 1600 and 48.57/0.992 at ISO 3200, compared with 49.37/0.991 and 48.16/0.989 for Graphics2RAW (Kim et al., 31 Mar 2026).

The reported advantages follow directly from the framework’s representation choice. By generating canonical XYZ, RawGen decouples scene synthesis from rendering, so white balance, exposure, and tone mapping remain reliable and camera-agnostic in the linear domain. Deterministic mapping to arbitrary camera raw spaces provides device-specific data without retraining, and optional noise plus CFA remosaicing makes the outputs suitable for raw-domain models. The use of text prompts further scales scene diversity without capture campaigns or graphics asset preparation (Kim et al., 31 Mar 2026).

The limitations are also explicit. Device fidelity beyond color remains incomplete, because accurate raw synthesis depends not only on color calibration but also on sensor noise, lens shading, point spread functions, and optical blur. RawGen currently injects heteroscedastic noise but does not model complex spatially varying characteristics. Inversion remains many-to-one, so extreme photo-finishing or heavy local edits can challenge recovery. Mapping quality also depends on correct DNG calibration; inaccurate or incomplete ForwardMatrix or ColorMatrix metadata degrades XYZ-to-raw accuracy. The paper therefore positions future work around richer device priors and learned physics for noise and optics (Kim et al., 31 Mar 2026).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (2)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to RawGen.