Papers
Topics
Authors
Recent
Search
2000 character limit reached

HA30K: Helmholtz Acoustic 30K Dataset

Updated 14 July 2026
  • HA30K is a dataset of 31,000 2D acoustic material configurations paired with pressure field images from Helmholtz equation solutions.
  • It enables rapid surrogate modeling by framing acoustic simulation as a conditional image-generation problem using Stable Diffusion and ControlNet.
  • The dataset supports applications in sound design, noise control, and inverse design while facilitating studies on speed-quality trade-offs via standard image metrics.

HA30K, short for Helmholtz Acoustic 30K, is a dataset of 31,000 acoustic materials introduced to support data-driven simulation of acoustic wave propagation by learning solutions of the Helmholtz equation from geometric and material configurations (Gramaccioni et al., 7 Oct 2025). It is motivated by the computational burden of traditional numerical solvers such as FEM and FDM, which are accurate but expensive for large-scale, high-resolution, or repeated simulations. The dataset pairs geometric/material configurations with corresponding pressure-field solutions, and the accompanying baseline frames Helmholtz solution prediction as a conditional image-generation problem using Stable Diffusion + ControlNet. The stated goal is not universal replacement of numerical solvers, but a practical foundation for fast acoustic simulation, learning mappings from geometry to pressure field, and future work in inverse design and physics-informed generative modeling.

1. Scope and motivation

HA30K was created as a benchmark for repeated Helmholtz-equation solves in settings where many acoustic configurations must be evaluated quickly (Gramaccioni et al., 7 Oct 2025). The motivating application areas identified for the dataset are sound design, noise control, acoustic material engineering, and early-stage exploration where many candidate designs must be evaluated quickly. In this framing, the central problem is not the formulation of the Helmholtz equation itself, but the expense of repeatedly solving it for changing geometries, source locations, frequencies, and material conditions.

The dataset is explicitly positioned as infrastructure for deep-learning-based surrogates. This suggests a workflow in which numerical simulation is shifted from repeated online solving to offline dataset generation followed by amortized inference. A plausible implication is that HA30K is intended to support comparative evaluation across surrogate architectures, not only the specific diffusion baseline reported in the paper.

2. Dataset structure and representation

HA30K contains 31,000 samples of 2D acoustic material configurations (Gramaccioni et al., 7 Oct 2025). Each sample includes a geometric/material configuration and a corresponding pressure-field solution. The geometric component is represented as an image of a square domain containing 1 to 6 square obstacles, a sound source location, and material types associated with the domain and obstacles. The solution component is obtained by solving the Helmholtz equation for the corresponding configuration.

The simulation outputs are stored in the form

[X,Y,Re(p),Im(p)],[X, Y, \mathrm{Re}(p), \mathrm{Im}(p)],

where XX and YY are spatial coordinates, pp is the complex acoustic pressure field, and Re(p)\mathrm{Re}(p) and Im(p)\mathrm{Im}(p) are its real and imaginary parts (Gramaccioni et al., 7 Oct 2025). For the dataset’s image target, the real part Re(p)\mathrm{Re}(p) is converted into an image.

The image encoding is deliberately simple. The input/source image is derived from obstacle placement, with black pixels = obstacles and white pixels = free space, while the target image is generated from Re(p)\mathrm{Re}(p) using the viridis colormap. Both are stored as 256 × 256 images. This representation recasts a PDE solution operator as an image-to-image mapping augmented by physical metadata.

3. Physical model and numerical formulation

The underlying PDE is the Helmholtz equation in the strong form

Δp+k2p=fin Ω,\Delta p + k^2 p = f \quad \text{in } \Omega,

where pp denotes the acoustic pressure field, XX0 the wave number, XX1 the source term, and XX2 the spatial domain (Gramaccioni et al., 7 Oct 2025). The formulation includes two boundary conditions.

On the outer boundary XX3, the paper uses the Sommerfeld radiation condition

XX4

which is used to simulate free radiation and suppress artificial reflections at the boundary. On obstacle boundaries XX5, it uses the Robin boundary condition

XX6

which models impedance effects at obstacle surfaces through the parameter XX7.

The paper also derives a weak formulation by multiplying by a test function XX8, integrating over XX9, and applying integration by parts. The resulting expression includes a volume term for the Helmholtz operator, a source term YY0, and boundary integrals associated with the Sommerfeld and Robin conditions (Gramaccioni et al., 7 Oct 2025). This weak form is what FreeFEM solves numerically.

The choice of boundary conditions is consequential for the interpretation of the dataset. The outer Sommerfeld condition corresponds to an open-radiation setting rather than a closed cavity, while the Robin condition makes obstacle behavior material-dependent through impedance. This suggests that HA30K is designed to capture both geometric scattering and surface interaction effects.

4. Data generation protocol

The dataset was generated by solving the Helmholtz equation using FreeFEM, described as an open-source finite element solver (Gramaccioni et al., 7 Oct 2025). The computational domain is a unit square, discretized with a 256×256 mesh. Obstacles are square, randomly placed without overlap, and their number ranges from 1 to 6. Obstacle sizes are sampled from [0.1, 0.4] with step 0.05, and on average the obstacles occupy about 32% of the domain area.

The physical parameters varied during generation include sound frequency, material properties, and sound source position. Frequency ranges from 100 to 4000 Hz in increments of 100 Hz, and the sound source is placed randomly in non-occupied regions. The paper provides the following speeds of sound for several materials:

Material Speed of sound
Rubber 60 m/s
Air at 20°C 343 m/s
Lead 1210 m/s
Gold 3240 m/s
Glass 4540 m/s
Aluminum 6320 m/s

For obstacle materials, the impedance parameter YY1 is used to model different surfaces:

Obstacle material Impedance YY2
foam 150
rubber 600
wood 1500
metal YY3

These design choices make the dataset structurally constrained but systematically varied. The setup is strictly 2D, limited to square domains and square obstacles, which narrows geometric diversity but also makes the benchmark well defined and reproducible.

5. Baseline generative modeling with Stable Diffusion and ControlNet

To demonstrate that HA30K can support fast surrogate modeling, the paper presents a baseline based on Stable Diffusion + ControlNet (Gramaccioni et al., 7 Oct 2025). The model takes the material/domain image as conditioning input and a text prompt encoding global physical parameters. The prompt has the form:

“The material is made of {material_domain}, has {num_obstacles} obstacles made of {material_obstacles}. A sound source is located in (x={s_x}, y={s_y}) having frequency {source_freq} Hz”

This text is used as cross-attention conditioning. The formulation therefore combines spatial conditioning through the image and global parameter conditioning through language.

The training configuration is specified as follows: Stable Diffusion weights are frozen, only ControlNet parameters are trained, and the checkpoint used is Stable Diffusion 2.1. The data split is 80% train, 15% validation, and 5% test. Training is performed on a single 48 GB NVIDIA RTX A6000 with batch size 4, learning rate 1e-5, and 100k diffusion steps, with each step defined using 1000 timesteps.

The baseline’s conceptual move is to represent the pressure solution as an image and thereby treat Helmholtz prediction as conditional image generation rather than explicit PDE solving. The stated benefits are no need to run a full PDE solver at inference time, GPU-friendly parallel inference, compatibility with batch processing, and an adjustable quality/speed trade-off via diffusion steps. A plausible implication is that the method is especially suited to throughput-oriented exploratory pipelines rather than solver-grade verification.

6. Evaluation, speed–quality trade-offs, and stated use

The paper evaluates generated outputs using standard image-quality metrics: MSE, FID, and SSIM (Gramaccioni et al., 7 Oct 2025). These are intended to measure both numerical pixel-level error and perceptual or structural fidelity. No separate custom training objective beyond the diffusion-model training setup implied by Stable Diffusion and ControlNet is described.

A reported comparison across DDIM sampling steps is given below:

DDIM steps FID SSIM MSE
20 53.31 0.657 0.1156
50 40.93 0.662 0.1290
75 43.62 0.669 0.1163

These results support the paper’s claim that inference quality can be traded against speed by changing the number of sampling steps. The same section emphasizes a speedup comparison between parallelized deep-learning inference and sequential FreeFEM simulations, using batch sizes of 1, 3, 5, 10, and 25 together with DDIM sampling at different step counts. The main takeaway stated in the paper is that the baseline can achieve substantial acceleration through GPU parallelization, particularly when many simulations are required.

The intended use cases of HA30K are listed as surrogate modeling of acoustic wave propagation, benchmarking generative models for PDE solutions, fast simulation of acoustic materials, inverse design of acoustic structures, and physics-informed deep learning research. The dataset is also positioned as a resource for studying how well image-generation models can learn physically meaningful fields from geometry.

7. Limitations and interpretation

The paper explicitly describes the baseline as an approximate surrogate, not a replacement for exact numerical solvers in all settings (Gramaccioni et al., 7 Oct 2025). It also notes that the target output is the real part of the pressure field represented as an image, which simplifies the full complex-valued solution. Additional constraints are that the setup is 2D and restricted to square obstacles in a square domain.

Performance is also tied to diffusion sampling steps, so faster inference may reduce fidelity. The paper therefore positions the approach as especially useful for early-stage exploration, where rapid screening of many candidate geometries matters more than perfect physical fidelity, rather than for high-precision engineering validation.

A common misconception would be to read HA30K as a claim that generative models supersede numerical simulation. The paper does not make that claim. Instead, it presents a dataset and baseline showing that image-based surrogates can approximate Helmholtz pressure fields quickly and in parallel, while classical solvers remain the reference mechanism for exact numerical computation. In that sense, HA30K functions both as a benchmark for surrogate accuracy and as a testbed for research on physics-aware generative modeling.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to HA30K.