Papers
Topics
Authors
Recent
Search
2000 character limit reached

SynthCloner: Factorized Synth Preset Conversion

Updated 14 July 2026
  • SynthCloner is a factorized neural audio codec that decomposes audio into distinct ADSR envelope, timbre, and content components.
  • The model employs three parallel encoders and a unified decoder to independently manipulate these factors, ensuring effective preset conversion.
  • Paired with the extensive SynthCAT dataset, it demonstrates superior performance in preserving targeted audio attributes compared to baseline systems.

SynthCloner is a factorized neural audio codec designed specifically for synthesizer preset conversion: given a source performance and a reference preset, it reconstructs the source’s musical content using the target’s timbre and dynamic envelope. Its central technical claim is that synthesizer conversion should not be treated as timbre transfer alone. Instead, audio is decomposed into three explicitly modeled attributes—ADSR envelope, timbre, and content—so that these factors can be edited and recombined independently. The model is paired with SynthCAT, a Serum-based dataset covering 250 timbres, 120 ADSR envelopes, and 100 MIDI sequences, and is evaluated on both objective and subjective criteria for timbre, envelope, and content preservation (Liu et al., 29 Sep 2025).

1. Problem formulation and scope

In SynthCloner, a synthesizer preset is treated as a compound object comprising parameters controlling oscillators, filters, and the master amplitude envelope. Preset conversion seeks to render a source performance using a reference preset’s timbre and envelope while preserving the source content, namely notes, f0f_0 trajectory, and timing. This formulation is motivated by the observation that many musical timbre-transfer systems emphasize spectral similarity while leaving the temporal amplitude envelope implicit, an assumption that is often misaligned with electronic synthesis, where the master amplitude envelope is itself a primary design variable (Liu et al., 29 Sep 2025).

A common conflation is between preset conversion and parameter inversion. Differentiable DSP or parameter estimation methods aim at interpretable synthesis controls, whereas SynthCloner learns directly from audio to perform many-to-many preset conversion with explicit envelope control. Earlier synthesizer-control work framed the task as learning an organized latent audio space with an invertible mapping to synthesizer parameters, addressing automatic parameter inference, macro-control learning, and audio-based preset exploration in one model (Esling et al., 2019). Other systems pursued practical cloning through multimodal retrieval, user-centered genetic mixing, and guided parameter editing rather than direct factorized conversion (Brade et al., 2023). SynthCloner occupies a different position: it is an audio-domain conversion model whose core design variable is explicit separation of content, timbre, and ADSR.

2. Factorized codec architecture

The architecture uses three parallel encoders and a unified decoder. The ADSR path takes xex_e, computes log-RMS energy, and maps it through a temporal multi-scale Conv-BiLSTM to an envelope latent ze∈RD×Tz_e \in \mathbb{R}^{D \times T}. The content path takes xcx_c through a shared encoder and residual vector quantization, producing zc∈RD×Tz_c \in \mathbb{R}^{D \times T}. The timbre path takes xtx_t through a shared encoder and a Conformer-based timbre extractor with global mean pooling, producing a global timbre embedding zt∈RDz_t \in \mathbb{R}^{D}. Fusion is additive at the frame level, with

zmix=zc+ze,z_{\text{mix}} = z_c + z_e,

followed by timbre conditioning through conditional layer normalization before waveform decoding (Liu et al., 29 Sep 2025).

The timbre modulation is written as

yt=γ(zt)⊙LN(ht)+β(zt),y_t = \gamma(z_t) \odot \mathrm{LN}(h_t) + \beta(z_t),

where γ(⋅)\gamma(\cdot) and xex_e0 are learned affine projections of the timbre embedding, xex_e1 is layer normalization, and xex_e2 denotes elementwise multiplication. This arrangement gives SynthCloner three distinct control surfaces: frame-synchronous content, frame-synchronous envelope, and global timbre. The content representation uses a codebook size of 1024 and 8 residual quantization layers; the timbre encoder uses a Conformer extractor; and the decoder maps the modulated features back to waveform (Liu et al., 29 Sep 2025).

This factorization is operational rather than merely descriptive. Each attribute can be swapped independently at inference. Envelope-only swap retains source timbre and content while replacing xex_e3 from a reference. Timbre-only swap replaces xex_e4 while retaining source envelope and content. Standard preset conversion uses source content together with reference timbre and reference ADSR. The reported ablations show that omitting either timbre transfer or ADSR transfer selectively degrades the corresponding metric, which suggests that the separation is functionally meaningful rather than a bookkeeping convenience (Liu et al., 29 Sep 2025).

3. ADSR formalization and the SynthCAT dataset

SynthCloner’s envelope modeling is grounded in a standard ADSR+Hold parameterization with Attack xex_e5, Decay xex_e6, Hold xex_e7, Sustain xex_e8, and Release xex_e9. In SynthCAT, these are sampled uniformly from the ranges

ze∈RD×Tz_e \in \mathbb{R}^{D \times T}0

The rendering pipeline uses a piecewise-linear envelope aligned to each MIDI note:

  • for ze∈RD×Tz_e \in \mathbb{R}^{D \times T}1, ze∈RD×Tz_e \in \mathbb{R}^{D \times T}2;
  • for ze∈RD×Tz_e \in \mathbb{R}^{D \times T}3, ze∈RD×Tz_e \in \mathbb{R}^{D \times T}4;
  • for ze∈RD×Tz_e \in \mathbb{R}^{D \times T}5, ze∈RD×Tz_e \in \mathbb{R}^{D \times T}6;
  • for ze∈RD×Tz_e \in \mathbb{R}^{D \times T}7, ze∈RD×Tz_e \in \mathbb{R}^{D \times T}8;
  • for ze∈RD×Tz_e \in \mathbb{R}^{D \times T}9, xcx_c0 (Liu et al., 29 Sep 2025).

SynthCAT is rendered from Serum VST. The dataset construction begins with commercial preset packs, from which long sustained tones are rendered and a 1-second one-shot is extracted from the flattest region. The flatness score is defined as

xcx_c1

Segments with flatness xcx_c2 are retained as timbre exemplars. These one-shots are then pitch-shifted and duration-aligned to 100 monophonic MIDI phrases from the mono-midi-transposition dataset, with per-note ADSR shaping applied during rendering (Liu et al., 29 Sep 2025).

The resulting corpus is the Cartesian product of 250 timbres, 120 envelopes, and 100 MIDI files, yielding approximately 3,000,000 mono audio samples at 44.1 kHz and about 2,500 hours in total. For evaluation, 50,000 test samples are formed using seen timbres together with 20 unseen envelopes and 10 unseen MIDI sequences, and each test source is paired with 10 randomly selected references. The dataset therefore supports both factor-specific perturbation during training and explicit ground-truth conversion targets during testing (Liu et al., 29 Sep 2025).

4. Training objectives and disentanglement strategy

The central training mechanism is attribute-specific perturbation. Each path sees inputs that preserve its target factor while perturbing the others. For ADSR, xcx_c3 shares envelope xcx_c4 but uses different content and timbre; analogous constructions are used for the content and timbre paths. This enforces invariance structurally, using the combinatorial coverage of SynthCAT rather than relying solely on latent regularization. Auxiliary supervision is added in the form of frame-level pitch labels from MIDI for content, and categorical classification losses over 120 envelope identities and 250 timbre identities for ADSR and timbre (Liu et al., 29 Sep 2025).

The reconstruction objective follows FACodec-style training. The multi-scale mel loss is

xcx_c5

with FFT windows xcx_c6, mel bins xcx_c7, and hop length equal to one quarter of the window length. Feature matching and adversarial losses are included as xcx_c8 and xcx_c9. Quantization uses standard VQ-VAE-style terms,

zc∈RD×Tz_c \in \mathbb{R}^{D \times T}0

and the classification heads contribute

zc∈RD×Tz_c \in \mathbb{R}^{D \times T}1

The total objective is

zc∈RD×Tz_c \in \mathbb{R}^{D \times T}2

with zc∈RD×Tz_c \in \mathbb{R}^{D \times T}3, zc∈RD×Tz_c \in \mathbb{R}^{D \times T}4, zc∈RD×Tz_c \in \mathbb{R}^{D \times T}5, zc∈RD×Tz_c \in \mathbb{R}^{D \times T}6, zc∈RD×Tz_c \in \mathbb{R}^{D \times T}7, and zc∈RD×Tz_c \in \mathbb{R}^{D \times T}8 (Liu et al., 29 Sep 2025).

Evaluation combines objective and subjective measures. The reported objective metrics are MSTFT, LRMSD, and F0RMSE. LRMSD is defined on log-RMS contours as

zc∈RD×Tz_c \in \mathbb{R}^{D \times T}9

and F0RMSE uses TorchCrepe pitch tracks:

xtx_t0

On these measures, SynthCloner reports MSTFT xtx_t1 versus xtx_t2 for CTD and xtx_t3 for SS-VAE; LRMSD xtx_t4 versus xtx_t5 and xtx_t6; and F0RMSE xtx_t7 Hz versus xtx_t8 and xtx_t9. Subjective tests with 20 listeners report TMOS zt∈RDz_t \in \mathbb{R}^{D}0, ADSRMOS zt∈RDz_t \in \mathbb{R}^{D}1, and CMOS zt∈RDz_t \in \mathbb{R}^{D}2, compared with ground-truth values of zt∈RDz_t \in \mathbb{R}^{D}3, zt∈RDz_t \in \mathbb{R}^{D}4, and zt∈RDz_t \in \mathbb{R}^{D}5 respectively. Removing the ADSR path degrades both objective and subjective performance, with MSTFT zt∈RDz_t \in \mathbb{R}^{D}6, LRMSD zt∈RDz_t \in \mathbb{R}^{D}7, F0RMSE zt∈RDz_t \in \mathbb{R}^{D}8, TMOS zt∈RDz_t \in \mathbb{R}^{D}9, ADSRMOS zmix=zc+ze,z_{\text{mix}} = z_c + z_e,0, and CMOS zmix=zc+ze,z_{\text{mix}} = z_c + z_e,1 (Liu et al., 29 Sep 2025).

5. Conversion modes and practical operation

SynthCloner supports several inference modes. In the standard many-to-many preset conversion setting, the model takes source audio for content and reference audio for timbre plus envelope:

  1. zmix=zc+ze,z_{\text{mix}} = z_c + z_e,2
  2. zmix=zc+ze,z_{\text{mix}} = z_c + z_e,3
  3. zmix=zc+ze,z_{\text{mix}} = z_c + z_e,4
  4. zmix=zc+ze,z_{\text{mix}} = z_c + z_e,5
  5. zmix=zc+ze,z_{\text{mix}} = z_c + z_e,6 The output preserves source content while adopting reference timbre and ADSR. Envelope-only swap uses zmix=zc+ze,z_{\text{mix}} = z_c + z_e,7 from a reference and zmix=zc+ze,z_{\text{mix}} = z_c + z_e,8 from the source; timbre-only swap does the reverse. Multiple references can also be combined by taking the envelope from one reference and the timbre from another (Liu et al., 29 Sep 2025).

These operating modes clarify a frequent misconception in synthesizer-cloning discussions: matching spectral color alone is not sufficient for preset fidelity. In SynthCloner, omitting ADSR conversion while keeping source envelope raises LRMSD from zmix=zc+ze,z_{\text{mix}} = z_c + z_e,9 to yt=γ(zt)⊙LN(ht)+β(zt),y_t = \gamma(z_t) \odot \mathrm{LN}(h_t) + \beta(z_t),0; omitting timbre conversion while keeping source timbre raises MSTFT from yt=γ(zt)⊙LN(ht)+β(zt),y_t = \gamma(z_t) \odot \mathrm{LN}(h_t) + \beta(z_t),1 to yt=γ(zt)⊙LN(ht)+β(zt),y_t = \gamma(z_t) \odot \mathrm{LN}(h_t) + \beta(z_t),2. At the same time, F0RMSE remains comparatively stable under these partial swaps, which indicates that content preservation is largely decoupled from the two style attributes (Liu et al., 29 Sep 2025).

The released implementation provides code, pretrained checkpoint, and audio demos. Reproduction settings specified in the paper use PyTorch and torchaudio, 1-second mono inputs at 44.1 kHz, AdamW with learning rate yt=γ(zt)⊙LN(ht)+β(zt),y_t = \gamma(z_t) \odot \mathrm{LN}(h_t) + \beta(z_t),3 and exponential decay yt=γ(zt)⊙LN(ht)+β(zt),y_t = \gamma(z_t) \odot \mathrm{LN}(h_t) + \beta(z_t),4, batch size yt=γ(zt)⊙LN(ht)+β(zt),y_t = \gamma(z_t) \odot \mathrm{LN}(h_t) + \beta(z_t),5, RVQ codebook size yt=γ(zt)⊙LN(ht)+β(zt),y_t = \gamma(z_t) \odot \mathrm{LN}(h_t) + \beta(z_t),6, 8 residual quantization layers, and 400k training steps (Liu et al., 29 Sep 2025).

6. Position within the broader cloning literature

SynthCloner belongs to a broader family of systems concerned with audio cloning, inverse synthesis, and high-level control, but its emphasis is distinct. Earlier work on synthesizer cloning has treated the problem as bidirectional mapping between audio and parameter space, using VAEs and Normalizing Flows to support automatic parameter inference, macro-controls, and preset exploration (Esling et al., 2019). Sounderfeit approached the cloning of a physical modeling bowed-string synthesizer through a conditional adversarial autoencoder that jointly learned parameter estimation and resynthesis from recorded data (Sinclair, 2018). Torchsynth and synth1B1 supplied a GPU-enabled modular synthesis environment with one billion synthesized sounds paired with exact synthesis parameters, together with inverse-synthesis and rank-based evaluation workflows (Turian et al., 2021). These approaches are parameter-centric or system-identification-centric; SynthCloner is instead an audio-domain preset-conversion model with explicit envelope factorization.

A second line of work has focused on practical, interactive cloning. SynthScribe combines multimodal retrieval, user-centered genetic mixing, and guided parameter editing to make cloning practical from text or audio references, but it does not perform automatic inverse synthesis against arbitrary external audio (Brade et al., 2023). White-box Serum systems such as "White-box Audio VST Effect Programming" and SerumRNN provide step-by-step effect programming instructions that iteratively transform a current sound toward a target sound, emphasizing interpretability and effect-order discovery rather than factorized latent control (Mitcheltree et al., 2021, Mitcheltree et al., 2021).

A third line concerns zero-shot instrument or voice cloning. TokenSynth uses a decoder-only transformer to generate neural audio codec tokens from MIDI tokens and CLAP embeddings, supporting instrument cloning, text-to-instrument synthesis, and text-guided timbre manipulation without fine-tuning (Kim et al., 13 Feb 2025). Anysynth removes the fixed timbre embedding bottleneck entirely, conditioning a Diffusion Transformer directly on uncompressed reference audio and target MIDI through in-context flow matching and Asymmetric Hierarchical CFG (Jing et al., 13 Jul 2026). In speech, expressive neural voice cloning conditions a Mellotron-style Tacotron 2 system on speaker encoding, pitch contour, and latent style tokens to transfer style and control expressiveness for unseen speakers (Neekhara et al., 2021). SynthCloner differs from all three by targeting synthesizer preset conversion specifically and by making the ADSR envelope a first-class factor rather than an implicit by-product.

The model’s stated limitations are also specific. It is monophonic, excludes presets with LFOs, coupled envelope routings, and time-varying effects, and leaves room for improvement on truly unseen timbres and more complex preset behaviors. Latency, model size, and exact inference runtime are not reported. A plausible implication is that its strongest current domain is controlled synthesizer audio with explicit one-shot timbres and linear ADSR shaping, rather than the full procedural complexity of contemporary synthesizer programming (Liu et al., 29 Sep 2025).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to SynthCloner.