DiTTO: Diffusion-Based Storage Trace Generator
- DiTTO is a generative framework that uses diffusion models to synthesize realistic and configurable multi-device storage traces from structured image representations.
- It leverages a conditioning mechanism called CHIP to align quantitative workload configurations with trace images for precise control over parameters like read/write ratios and intensity.
- The framework incorporates sparsity-aware strategies and temporal outpainting to extend trace generation to arbitrary lengths while preserving temporal coherence and inter-device dependencies.
Searching arXiv for the specified paper and closely related storage-trace generation context. arxiv_search.query({"search_query":"id:(Kim et al., 2 Sep 2025) OR ti:\"A Diffusion-Based Framework for Configurable and Realistic Multi-Storage Trace Generation\"","start":0,"max_results":5}) DiTTO is a generative framework for synthesizing realistic, configurable, and diverse multi-device storage traces by applying diffusion models to structured image-like encodings of storage activity (Kim et al., 2 Sep 2025). It addresses the difficulty of collecting and sharing real storage traces, and the limited expressiveness of template-based synthetic generators, by learning from real multi-device traces while allowing precise user control over properties such as read/write ratio, workload intensity, and device utilization patterns. The framework combines diffusion-based generation, a conditioning mechanism named CHIP (Contrastive Hyperconfiguration-Image Pretraining), a sparsity-aware representation for highly sparse I/O workloads, and an outpainting procedure for generating arbitrarily long traces (Kim et al., 2 Sep 2025).
1. Problem formulation and scope
Storage traces are logs of I/O operations and are used for designing and evaluating storage systems, testing schedulers, caching policies, and load balancers, evaluating device technologies such as SSDs, HDDs, and NVRAM/PCM, and conducting “what-if” performance studies under different workloads (Kim et al., 2 Sep 2025). Real traces are hard to collect because they incur instrumentation overhead, raise privacy concerns, and are often not shared. Template-based generators such as Filebench and SPECstorage do not capture complex, evolving, multi-device workload behavior, while prior ML-based generators still rely heavily on hand-crafted features and often lack precise control over key workload parameters (Kim et al., 2 Sep 2025).
DiTTO is designed to learn from real multi-device traces and then generate workloads that are realistic, precisely configurable, and diverse. It is, to the authors’ knowledge, the first application of diffusion models to storage trace generation (Kim et al., 2 Sep 2025). Its stated objectives are to capture temporal dynamics and inter-device dependencies, align generated traces with user-defined configurations, and extend generated sequences to arbitrary length via temporal outpainting.
A plausible implication is that DiTTO is positioned as a systems-oriented generative model rather than a generic sequence synthesizer: its target object is not merely an event stream, but a workload surrogate suitable for storage-system evaluation under controlled experimental conditions.
2. Diffusion formulation and trace representation
DiTTO follows the denoising diffusion probabilistic model paradigm, adapted to 2D multi-channel “images” that encode trace segments (Kim et al., 2 Sep 2025). If denotes the image representation of a trace window, the forward diffusion process adds Gaussian noise according to
with the closed-form relation
where (Kim et al., 2 Sep 2025). The reverse process is learned conditionally as , with a neural network trained under the simplified DDPM objective
The central representational step is the conversion of raw storage logs into a multi-channel image. Time is discretized into fixed-length bins along the -axis, devices occupy rows on the -axis, and separate channels encode read and write intensity (Kim et al., 2 Sep 2025). This yields an image representation in which temporal dynamics appear along the horizontal axis, multi-device structure appears along the vertical axis, and cross-device correlations appear as spatial patterns spanning multiple rows and channels.
Because storage workloads are extremely sparse, DiTTO does not use direct binary encoding alone. Instead, it adds local Gaussian noise around access pixels, so that an event at becomes a small region with Gaussian falloff rather than a single isolated pixel (Kim et al., 2 Sep 2025). This representation smooths gradients, mitigates extreme sparsity, and provides contextual structure for diffusion training. The paper states that this sparsity-aware strategy is important for training stability and for modeling idle periods.
The core denoiser is a U-Net–style decoder operating on these image-like trace representations and conditioned on the diffusion timestep and a configuration embedding (Kim et al., 2 Sep 2025). The framework is described as drawing inspiration from DALL·E 2’s hierarchical latent diffusion over CLIP-like embeddings, and from image diffusion models such as Palette and RePaint.
3. Conditioning, configurability, and CHIP
A defining feature of DiTTO is explicit configurability. The framework supports user control over global and per-trace read/write ratio, workload intensity, device utilization patterns, and related I/O characteristics such as the number of total requests in a window (Kim et al., 2 Sep 2025). This conditioning mechanism is implemented through CHIP, short for Contrastive Hyperconfiguration-Image Pretraining.
CHIP is described as analogous to CLIP, but with system configurations instead of natural language and trace images instead of photographs (Kim et al., 2 Sep 2025). For each trace window, DiTTO computes a numeric “hyperconfiguration” vector 0 containing descriptors such as overall read percentage, total I/O count, and fraction of accesses per device. CHIP then aligns configuration embeddings and image embeddings through contrastive pretraining, using a configuration encoder 1 and an image encoder 2. At generation time, a desired configuration 3 is mapped to an embedding 4, which conditions the denoiser 5 (Kim et al., 2 Sep 2025).
This design gives DiTTO a quantitatively specified control surface rather than a qualitative prompt interface. The paper states that traces with similar configurations are mapped close in embedding space, while mismatched pairs are pushed apart. This suggests that the conditioning space is intended to preserve semantically meaningful workload neighborhoods rather than merely act as an auxiliary label vector.
The same section also reports an emergent property: even characteristics not explicitly included in the configuration vector, such as spatial locality, are reproduced in a configuration-dependent way (Kim et al., 2 Sep 2025). In the evaluation, traces clustered by locality metrics yield regenerated traces with distinct locality distributions, implying that the model captures higher-order workload statistics beyond the explicitly conditioned fields.
4. Data pipeline, preprocessing, and outpainting
DiTTO is evaluated on the Alibaba Block IO trace dataset, identified as a real-world block-level trace set capturing large-scale I/O request streams, multi-device behaviors, and real production workloads (Kim et al., 2 Sep 2025). Although the paper cites other trace sources, including SNIA IOTTA RocksDB SSD traces and FIU filesystem syscall traces, the implemented experiments focus on Alibaba Block IO.
The preprocessing pipeline consists of event extraction, fixed-duration windowing, image construction, Gaussian smoothing, configuration computation, and normalization (Kim et al., 2 Sep 2025). Event extraction retains timestamp, operation type, and device ID for each I/O. Each time window becomes one training example. Discretization places time bins on the horizontal axis and device IDs on the vertical axis, with separate read and write channels. Gaussian smoothing is then applied around access pixels, and numeric workload metrics are computed for CHIP.
The training implementation uses PyTorch 2.4 (Kim et al., 2 Sep 2025). Exact hyperparameters such as the number of diffusion steps, noise schedule, and network width are not detailed in the short paper, but the training is described as following standard DDPM practices for image-size inputs and incorporating sparsity-aware strategies for the large proportion of empty time bins.
A further architectural element is temporal outpainting. DiTTO adapts image outpainting to trace generation so that trace segments can be extended beyond a fixed window size (Kim et al., 2 Sep 2025). In practical use, the tail region of a previously generated segment is provided as context for the next segment, and the process is repeated until the desired duration is reached. The paper describes this as enabling arbitrarily long traces while preserving temporal coherence across boundaries.
A plausible implication is that outpainting provides a compromise between fixed-window image modeling and open-ended temporal generation: the model learns on bounded windows but can be deployed for long-horizon synthetic trace construction without reverting to a wholly different sequential generator.
5. Evaluation: configurability, fidelity, and diversity
The reported evaluation targets three properties: configurable adherence, fidelity, and diversity (Kim et al., 2 Sep 2025). For configurable adherence, the authors generate 50 traces from user-defined configurations and report average read/write ratio error of less than 8%, with an example of a target of 82.99% reads yielding approximately 81.24% reads. Device utilization pattern error is reported at around 2% on average (Kim et al., 2 Sep 2025). The paper defines component-wise relative error as
6
Fidelity is evaluated qualitatively through visual comparisons of original and generated trace images and quantitatively through t-SNE embedding analysis (Kim et al., 2 Sep 2025). The synthetic images are reported to exhibit similar burstiness, idle periods, and multi-device access patterns. In the t-SNE plots, generated traces cluster near original traces but not exactly on top of them, which the paper interprets as simultaneous evidence of fidelity and diversity rather than mode collapse.
Temporal coherence under outpainting is shown by extended traces whose segment boundaries exhibit no obvious discontinuities (Kim et al., 2 Sep 2025). Spatial locality is analyzed by clustering real traces into two groups according to read/write locality metrics and then regenerating traces from configurations sampled from those clusters. The resulting locality distributions are reported to match cluster-wise, despite locality not being explicitly included in the conditioning vector.
The paper does not report formal statistical distances such as KL divergence or Wasserstein distance (Kim et al., 2 Sep 2025). It relies instead on configuration error, image inspection, t-SNE structure, and locality-distribution analysis. This suggests that the evaluation emphasizes systems-facing workload plausibility and controllability over benchmark conventions from generative modeling.
6. Applications, comparison context, and limitations
DiTTO is presented as useful for benchmarking and evaluation of file systems, distributed storage systems, RAID, and CephFS; for “what-if” and stress testing under controlled trace characteristics; for trace scaling through outpainting; and for generating workloads for different numbers of devices or utilization patterns by adjusting configuration vectors (Kim et al., 2 Sep 2025). A practical generation workflow consists of choosing a base dataset, specifying a desired configuration 7, computing the CHIP embedding, running conditional diffusion to generate a trace segment, outpainting additional segments if needed, and decoding the generated image back into timestamped events (Kim et al., 2 Sep 2025).
In contextual comparison, the paper contrasts DiTTO with template-based generators such as Filebench and SPECstorage, which rely on fixed templates and provide only coarse control, and with prior ML-based generators that often use Markov models, RNNs, or other sequence models and require substantial feature engineering (Kim et al., 2 Sep 2025). DiTTO is described as data-driven, explicitly multi-device, and capable of arbitrary-length generation with low configuration error.
The paper also identifies or implies several limitations. Realism depends heavily on the training traces, with experiments centered on Alibaba Block IO; generalization to completely different workloads or device types is not guaranteed without retraining (Kim et al., 2 Sep 2025). Rare behaviors may be underrepresented unless present in the data. Conditioning currently targets metrics such as read/write ratio and basic utilization patterns, not more complex workload descriptors such as queue depth distributions, per-file access patterns, or detailed QoS constraints. The work also does not address extremely large multi-node or multi-rack storage systems in detail.
Future directions proposed in the paper include richer conditioning, such as queue depth distributions or time-varying configuration trajectories; integration with simulators and online tools; extensions to broader multi-node or cross-datacenter storage domains; hierarchical or multi-resolution representations; and richer architectures or guidance methods, including transformers and classifier-free guidance variants (Kim et al., 2 Sep 2025). This suggests that DiTTO is best understood as a proof that diffusion-based generative modeling can be adapted to sparse, structured systems data, with configurable workload synthesis as the immediate application and broader systems-data generation as a plausible extension.