---
title: 'SynthCloner: Factorized Synth Preset Conversion'
url: https://www.emergentmind.com/topics/synthcloner
type: topic
---

# SynthCloner: Factorized Synth Preset Conversion

SynthCloner is a factorized neural audio codec designed specifically for synthesizer preset conversion: given a source performance and a reference preset, it reconstructs the source’s musical content using the target’s timbre and dynamic envelope. Its central technical claim is that synthesizer conversion should not be treated as timbre transfer alone. Instead, audio is decomposed into three explicitly modeled attributes—ADSR envelope, timbre, and content—so that these factors can be edited and recombined independently. The model is paired with SynthCAT, a Serum-based dataset covering 250 timbres, 120 ADSR envelopes, and 100 MIDI sequences, and is evaluated on both objective and subjective criteria for timbre, envelope, and content preservation [2509.24286].

## 1. Problem formulation and scope

In SynthCloner, a synthesizer preset is treated as a compound object comprising parameters controlling oscillators, filters, and the master amplitude envelope. Preset conversion seeks to render a source performance using a reference preset’s timbre and envelope while preserving the source content, namely notes, $f_0$ trajectory, and timing. This formulation is motivated by the observation that many musical timbre-transfer systems emphasize spectral similarity while leaving the temporal amplitude envelope implicit, an assumption that is often misaligned with electronic synthesis, where the master amplitude envelope is itself a primary design variable [2509.24286].

A common conflation is between preset conversion and parameter inversion. Differentiable DSP or parameter estimation methods aim at interpretable synthesis controls, whereas SynthCloner learns directly from audio to perform many-to-many preset conversion with explicit envelope control. Earlier synthesizer-control work framed the task as learning an organized latent audio space with an invertible mapping to synthesizer parameters, addressing automatic parameter inference, macro-control learning, and audio-based preset exploration in one model [1907.00971]. Other systems pursued practical cloning through multimodal retrieval, user-centered genetic mixing, and guided parameter editing rather than direct factorized conversion [2312.04690]. SynthCloner occupies a different position: it is an audio-domain conversion model whose core design variable is explicit separation of content, timbre, and ADSR.

## 2. Factorized codec architecture

The architecture uses three parallel encoders and a unified decoder. The ADSR path takes $x_e$, computes log-RMS energy, and maps it through a temporal multi-scale Conv-BiLSTM to an envelope latent $z_e \in \mathbb{R}^{D \times T}$. The content path takes $x_c$ through a shared encoder and residual vector quantization, producing $z_c \in \mathbb{R}^{D \times T}$. The timbre path takes $x_t$ through a shared encoder and a Conformer-based timbre extractor with global mean pooling, producing a global timbre embedding $z_t \in \mathbb{R}^{D}$. Fusion is additive at the frame level, with
$$
z_{\text{mix}} = z_c + z_e,
$$
followed by timbre conditioning through conditional layer normalization before waveform decoding [2509.24286].

The timbre modulation is written as
$$
y_t = \gamma(z_t) \odot \mathrm{LN}(h_t) + \beta(z_t),
$$
where $\gamma(\cdot)$ and $\beta(\cdot)$ are learned affine projections of the timbre embedding, $\mathrm{LN}$ is layer normalization, and $\odot$ denotes elementwise multiplication. This arrangement gives SynthCloner three distinct control surfaces: frame-synchronous content, frame-synchronous envelope, and global timbre. The content representation uses a codebook size of 1024 and 8 residual quantization layers; the timbre encoder uses a Conformer extractor; and the decoder maps the modulated features back to waveform [2509.24286].

This factorization is operational rather than merely descriptive. Each attribute can be swapped independently at inference. Envelope-only swap retains source timbre and content while replacing $z_e$ from a reference. Timbre-only swap replaces $z_t$ while retaining source envelope and content. Standard preset conversion uses source content together with reference timbre and reference ADSR. The reported ablations show that omitting either timbre transfer or ADSR transfer selectively degrades the corresponding metric, which suggests that the separation is functionally meaningful rather than a bookkeeping convenience [2509.24286].

## 3. ADSR formalization and the SynthCAT dataset

SynthCloner’s envelope modeling is grounded in a standard ADSR+Hold parameterization with Attack $T_A$, Decay $T_D$, Hold $T_H$, Sustain $S$, and Release $T_R$. In SynthCAT, these are sampled uniformly from the ranges
$$
T_A \in [10, 100]\ \mathrm{ms},\quad
T_D \in [50, 300]\ \mathrm{ms},\quad
T_H \in [0, 200]\ \mathrm{ms},\quad
S \in [0.0, 0.80],\quad
T_R \in [30, 300]\ \mathrm{ms}.
$$
The rendering pipeline uses a piecewise-linear envelope aligned to each MIDI note:
- for $t \in [t_0, t_0 + T_A)$, $E(t) = (t - t_0)/T_A$;
- for $t \in [t_0 + T_A, t_0 + T_A + T_H)$, $E(t) = 1$;
- for $t \in [t_0 + T_A + T_H, t_0 + T_A + T_H + T_D)$, $E(t) = 1 - (1 - S)\cdot (t - (t_0 + T_A + T_H))/T_D$;
- for $t \in [t_0 + T_A + T_H + T_D, t_{\mathrm{off}})$, $E(t) = S$;
- for $t \ge t_{\mathrm{off}}$, $E(t) = S \cdot \max(0, 1 - (t - t_{\mathrm{off}})/T_R)$ [2509.24286].

SynthCAT is rendered from Serum VST. The dataset construction begins with commercial preset packs, from which long sustained tones are rendered and a 1-second one-shot is extracted from the flattest region. The flatness score is defined as
$$
\mathrm{Flatness}(x) = \frac{1}{1 + \mathrm{Var}(\mu(x))}, \qquad
\mu(x) = \frac{1}{N}\sum_{n=1}^{N} x[n].
$$
Segments with flatness $> 0.95$ are retained as timbre exemplars. These one-shots are then pitch-shifted and duration-aligned to 100 monophonic MIDI phrases from the mono-midi-transposition dataset, with per-note ADSR shaping applied during rendering [2509.24286].

The resulting corpus is the Cartesian product of 250 timbres, 120 envelopes, and 100 MIDI files, yielding approximately 3,000,000 mono audio samples at 44.1 kHz and about 2,500 hours in total. For evaluation, 50,000 test samples are formed using seen timbres together with 20 unseen envelopes and 10 unseen MIDI sequences, and each test source is paired with 10 randomly selected references. The dataset therefore supports both factor-specific perturbation during training and explicit ground-truth conversion targets during testing [2509.24286].

## 4. Training objectives and disentanglement strategy

The central training mechanism is attribute-specific perturbation. Each path sees inputs that preserve its target factor while perturbing the others. For ADSR, $x_e$ shares envelope $e_0$ but uses different content and timbre; analogous constructions are used for the content and timbre paths. This enforces invariance structurally, using the combinatorial coverage of SynthCAT rather than relying solely on latent regularization. Auxiliary supervision is added in the form of frame-level pitch labels from MIDI for content, and categorical classification losses over 120 envelope identities and 250 timbre identities for ADSR and timbre [2509.24286].

The reconstruction objective follows FACodec-style training. The multi-scale mel loss is
$$
L_{\mathrm{mel}} = \sum_{s \in S} \|M_s(x) - M_s(\hat{y})\|_1,
$$
with FFT windows $[32, 64, 128, 256, 512, 1024, 2048]$, mel bins $[5, 10, 20, 40, 80, 160, 320]$, and hop length equal to one quarter of the window length. Feature matching and adversarial losses are included as $L_{\mathrm{feat}}$ and $L_{\mathrm{adv}}$. Quantization uses standard VQ-VAE-style terms,
$$
L_{\mathrm{commit}} = \|\mathrm{sg}[e] - z\|_2^2,\qquad
L_{\mathrm{codebook}} = \|e - \mathrm{sg}[z]\|_2^2,
$$
and the classification heads contribute
$$
L_{\mathrm{timbre}} = \mathrm{CE}(p_{\mathrm{timbre}}(\cdot|z_t), y_{\mathrm{timbre}}),\quad
L_{\mathrm{content}} = \mathrm{CE}(p_{\mathrm{pitch}}(\cdot|z_c), y_{\mathrm{pitch}}),\quad
L_{\mathrm{adsr}} = \mathrm{CE}(p_{\mathrm{env}}(\cdot|z_e), y_{\mathrm{env}}).
$$
The total objective is
$$
L_{\mathrm{total}} =
\lambda_{\mathrm{mel}} L_{\mathrm{mel}} +
\lambda_{\mathrm{feat}} L_{\mathrm{feat}} +
\lambda_{\mathrm{adv}} L_{\mathrm{adv}} +
\lambda_{\mathrm{commit}} L_{\mathrm{commit}} +
\lambda_{\mathrm{codebook}} L_{\mathrm{codebook}} +
\lambda_{\mathrm{timbre}} L_{\mathrm{timbre}} +
\lambda_{\mathrm{content}} L_{\mathrm{content}} +
\lambda_{\mathrm{adsr}} L_{\mathrm{adsr}},
$$
with $\lambda_{\mathrm{mel}}=15.0$, $\lambda_{\mathrm{feat}}=2.0$, $\lambda_{\mathrm{adv}}=1.0$, $\lambda_{\mathrm{commit}}=0.25$, $\lambda_{\mathrm{codebook}}=1.0$, and $\lambda_{\mathrm{timbre}}=\lambda_{\mathrm{content}}=\lambda_{\mathrm{adsr}}=5.0$ [2509.24286].

Evaluation combines objective and subjective measures. The reported objective metrics are MSTFT, LRMSD, and F0RMSE. LRMSD is defined on log-RMS contours as
$$
\mathrm{LRMSD} = \frac{1}{T}\sum_{t=1}^{T} |\hat{R}(t) - R(t)|,
$$
and F0RMSE uses TorchCrepe pitch tracks:
$$
\mathrm{F0RMSE} = \sqrt{\frac{1}{N}\sum_{n=1}^{N}(\hat{f}_0[n] - f_0[n])^2}.
$$
On these measures, SynthCloner reports MSTFT $3.00$ versus $5.69$ for CTD and $7.22$ for SS-VAE; LRMSD $0.17$ versus $0.89$ and $0.92$; and F0RMSE $20.64$ Hz versus $583.01$ and $641.62$. Subjective tests with 20 listeners report TMOS $3.91$, ADSRMOS $3.94$, and CMOS $4.11$, compared with ground-truth values of $4.08$, $3.96$, and $4.25$ respectively. Removing the ADSR path degrades both objective and subjective performance, with MSTFT $3.84$, LRMSD $0.42$, F0RMSE $29.04$, TMOS $3.09$, ADSRMOS $2.40$, and CMOS $3.76$ [2509.24286].

## 5. Conversion modes and practical operation

SynthCloner supports several inference modes. In the standard many-to-many preset conversion setting, the model takes source audio for content and reference audio for timbre plus envelope:
1. $z_c = \mathrm{ContentEncoderRVQ}(x_{\mathrm{src}})$  
2. $z_e = \mathrm{EnvelopeEncoder}(\log\mathrm{RMS}(x_{\mathrm{ref}}))$  
3. $z_t = \mathrm{TimbreEncoder}(x_{\mathrm{ref}})$  
4. $z_{\mathrm{mix}} = z_c + z_e$  
5. $\hat{y} = \mathrm{Decoder}(\mathrm{AdaLN}(z_{\mathrm{mix}}; z_t))$  
The output preserves source content while adopting reference timbre and ADSR. Envelope-only swap uses $z_e$ from a reference and $z_t$ from the source; timbre-only swap does the reverse. Multiple references can also be combined by taking the envelope from one reference and the timbre from another [2509.24286].

These operating modes clarify a frequent misconception in synthesizer-cloning discussions: matching spectral color alone is not sufficient for preset fidelity. In SynthCloner, omitting ADSR conversion while keeping source envelope raises LRMSD from $0.17$ to $0.39$; omitting timbre conversion while keeping source timbre raises MSTFT from $3.00$ to $5.97$. At the same time, F0RMSE remains comparatively stable under these partial swaps, which indicates that content preservation is largely decoupled from the two style attributes [2509.24286].

The released implementation provides code, pretrained checkpoint, and audio demos. Reproduction settings specified in the paper use PyTorch and torchaudio, 1-second mono inputs at 44.1 kHz, AdamW with learning rate $10^{-4}$ and exponential decay $0.999996$, batch size $8$, RVQ codebook size $1024$, 8 residual quantization layers, and 400k training steps [2509.24286].

## 6. Position within the broader cloning literature

SynthCloner belongs to a broader family of systems concerned with audio cloning, inverse synthesis, and high-level control, but its emphasis is distinct. Earlier work on synthesizer cloning has treated the problem as bidirectional mapping between audio and parameter space, using VAEs and Normalizing Flows to support automatic parameter inference, macro-controls, and preset exploration [1907.00971]. Sounderfeit approached the cloning of a physical modeling bowed-string synthesizer through a conditional adversarial autoencoder that jointly learned parameter estimation and resynthesis from recorded data [1806.09617]. Torchsynth and synth1B1 supplied a GPU-enabled modular synthesis environment with one billion synthesized sounds paired with exact synthesis parameters, together with inverse-synthesis and rank-based evaluation workflows [2104.12922]. These approaches are parameter-centric or system-identification-centric; SynthCloner is instead an audio-domain preset-conversion model with explicit envelope factorization.

A second line of work has focused on practical, interactive cloning. SynthScribe combines multimodal retrieval, user-centered genetic mixing, and guided parameter editing to make cloning practical from text or audio references, but it does not perform automatic inverse synthesis against arbitrary external audio [2312.04690]. White-box Serum systems such as "White-box Audio VST Effect Programming" and SerumRNN provide step-by-step effect programming instructions that iteratively transform a current sound toward a target sound, emphasizing interpretability and effect-order discovery rather than factorized latent control [2102.03170] [2104.03876].

A third line concerns zero-shot instrument or voice cloning. TokenSynth uses a decoder-only transformer to generate neural audio codec tokens from MIDI tokens and CLAP embeddings, supporting instrument cloning, text-to-instrument synthesis, and text-guided timbre manipulation without fine-tuning [2502.08939]. Anysynth removes the fixed timbre embedding bottleneck entirely, conditioning a Diffusion Transformer directly on uncompressed reference audio and target MIDI through in-context flow matching and Asymmetric Hierarchical CFG [2607.11143]. In speech, expressive neural voice cloning conditions a Mellotron-style Tacotron 2 system on speaker encoding, pitch contour, and latent style tokens to transfer style and control expressiveness for unseen speakers [2102.00151]. SynthCloner differs from all three by targeting synthesizer preset conversion specifically and by making the ADSR envelope a first-class factor rather than an implicit by-product.

The model’s stated limitations are also specific. It is monophonic, excludes presets with LFOs, coupled envelope routings, and time-varying effects, and leaves room for improvement on truly unseen timbres and more complex preset behaviors. Latency, model size, and exact inference runtime are not reported. A plausible implication is that its strongest current domain is controlled synthesizer audio with explicit one-shot timbres and linear ADSR shaping, rather than the full procedural complexity of contemporary synthesizer programming [2509.24286].

Source: https://www.emergentmind.com/topics/synthcloner