Papers
Topics
Authors
Recent
Search
2000 character limit reached

Generating Financial Time Series by Matching Random Convolutional Features

Published 3 Jun 2026 in cs.LG and q-fin.ST | (2606.05138v1)

Abstract: Generating realistic financial time series is challenging as training data is often limited to a single historical path. With such scarce data, overfitting is hard to avoid, especially under adversarial training where a trained discriminator can memorize the training samples. To mitigate this, recent approaches train generators to minimize the discrepancy between untrained feature representations of real and generated time series. In these works, the feature maps are based on path signatures, which can fail to capture relevant time series properties at tractable truncation depths. In this work, we instead train generators by matching random convolutional features of real and generated time series. Existing random convolutional feature maps, such as Rocket and Hydra, have been shown to provide informative representations of real-world time series, but cannot supervise generative models because they are non-differentiable. We introduce SOCK (SOft Competing Kernels), a fully differentiable random convolutional feature map, suited to train generative time series models. We show that generators trained by matching random SOCK features consistently outperform signature and diffusion baselines across a wide range of small-sample financial datasets. We further demonstrate SOCK's expressiveness on two-sample hypothesis testing and time series classification tasks, where SOCK matches or outperforms existing unsupervised feature maps.

Summary

  • The paper introduces SOCK, the first fully differentiable random convolutional feature map, achieving high-fidelity matching of financial time series using soft competing pooling.
  • It employs a resampling strategy during training to mitigate overfitting in scarce-data regimes, outperforming signature-based and diffusion-based methods on synthetic and real datasets.
  • Empirical evaluations demonstrate SOCK’s competitive discriminative power in hypothesis testing and classification, with scalable integration for financial risk management and simulation.

Generative Modeling of Financial Time Series via Differentiable Random Convolutional Features

Motivation and Problem Setting

Financial time series generation is fundamentally constrained by limited sample availability; typically, only a single historical sequence is accessible for training. Traditional adversarial generative paradigms, such as GANs, are prone to overfitting, especially in the small-sample regime, due to discriminator memorization. Recent work has shifted towards feature-matching approaches that leverage fixed, untrained feature representations (e.g., truncated or randomized path signatures), but signature-based embeddings often omit relevant statistical properties at any tractable truncation depth.

This paper establishes a new protocol for conditional financial time series generation: learning a generator that, given qq recent observations, produces plausible futures by matching convolutional statistics of real and synthetic joint past/future path segments. The protocol targets evaluation in practical scenarios—training and testing on single realized paths with context-conditioned rollouts.

Random Convolutional Feature Matching and the SOCK Feature Map

Random convolutional feature maps (e.g., Rocket, Hydra) have empirically dominated small-sample time series classification benchmarks by extracting high-dimensional, unsupervised features from convolutional responses and pooling over time. However, non-differentiable pooling operations (e.g., argmax, indicator count) inhibit their use in training generative models via gradient descent.

The paper introduces SOCK (SOft Competing Kernels), the first fully differentiable random convolutional feature map that retains the competitive dynamics of Hydra-style random convolutions. SOCK operates as follows:

  • Input preprocessing: Augments, normalizes, and projects input paths via randomized Gaussian matrices.
  • Grouped random convolutions: Channels are partitioned, convolutions performed per group with ℓ1\ell_1-normalized, zero-sum Gaussian kernels (group width W=2W=2 is optimal).
  • Soft competing pooling: Kernel competitions at each time-step are encoded via softmax win probabilities (temperature τ\tau), aggregated by temporal statistics (e.g., soft-deviation—the std of win probabilities across time). Multi-scale statistics are obtained via multiple dilations.

SOCK leverages resampling during training—randomly refreshing feature map parameters—to expose the generator to a distribution of supervisory signals, thus mitigating overfitting to any single representation.

Empirical Results: Synthesis and Real Datasets

Experimental evaluation spans synthetic benchmarks (V1, V10, TGH, SV, FGN) and real-world datasets (sector baskets, FX, index, crypto), all trained and evaluated under single-path, conditional generation constraints. Discrepancy metrics include discriminative classifiers (SIG, MLP, RNN, SRNN) and distributional assessments (autocorrelation discrepancy, cross-correlation, marginal fit, tail fit via PnL expected shortfall).

SOCK-trained generators consistently yield the smallest distributional discrepancies:

Figure 1

Figure 1: Evaluation summary across synthetic and real datasets; larger, lighter dots denote lower discrepancy between real and generated distributions.

Marginal distributions from SOCK closely track out-of-sample empirical densities, even for log-variance processes with strong non-Gaussian structure (Tukey gg-hh, SV); autocorrelation in complex oscillatory regimes is faithfully reproduced. Generators trained with SOCK outperform signature-matching and diffusion baselines quantitatively. Tail fit under trading strategy PnL further corroborates SOCK's superior fidelity.

Figure 2

Figure 2: Visual comparison of real vs. generated distributions for various metrics (marginals, autocorrelation, cumulative variance, PnL tail); SOCK maintains closest alignment to real distributions.

Ablation studies indicate SOCK's robustness to architectural and training choices. Division of convolutional groups (kernel width), choice of data augmentation (integrated path, posneg), and pooling function (soft-deviation, soft-count/value) all show minimal sensitivity—highlighting the efficacy of randomized feature matching.

Discriminative Feature Expressiveness

SOCK's utility as a discriminative representation is validated in two-sample hypothesis tests for stochastic processes (e.g., detecting transitions in Hurst exponent, tail heaviness, correlation), where it consistently achieves high test power. Time series classification tasks on the UCR 112-dataset archive further demonstrate SOCK's competitiveness, equaling or exceeding MultiRocket and Hydra with an order-of-magnitude fewer random features and a single pooling operation.

Figure 3

Figure 3: Permutation test power for feature- and kernel-based MMDs; SOCK achieves maximal sensitivity to parameter shifts in stochastic processes.

Figure 4

Figure 4: UCR time series classification benchmarks; SOCK matches or exceeds state-of-the-art ranking with efficient feature allocation.

Parameter and feature budget analyses evidence the scalability and efficiency of SOCK in resource-constrained settings. Pooling ablations confirm that soft-deviation pooling is universally competitive, and the model is insensitive to random bias injections.

Figure 5

Figure 5

Figure 5: Sensitivity of SOCK's mean rank on UCR datasets to softmax temperature τ\tau and competing kernels kk.

Figure 6

Figure 6

Figure 6: Feature budget analysis; SOCK maintains high mean rank for varying kernel/group count and feature dimensionality.

Implications and Future Directions

The findings establish random convolutional feature matching (with resampling) as a robust paradigm for conditional time series generation in finite-data financial regimes, superseding signature-based and diffusion approaches. SOCK delivers strong generative fidelity and discriminative expressiveness, applicable to risk measurement, trading simulation, and hedging policy optimization.

From a theoretical perspective, the approach leverages ensemble randomness and high-dimensional convolutional statistics to approximate conditional laws—suggesting directions for universal embedding theories in stochastic process modeling. Practically, SOCK is amenable to integration in time series pipelines for synthetic data augmentation, automated financial stress-testing, and robust candidate policy evaluation.

Future work may extend SOCK to higher-dimensional regimes, investigate hybrid learned/randomized feature maps, and exploit differentiability for end-to-end fine-tuning within deep generative architectures for scenario simulation and risk management.

Conclusion

This manuscript substantiates that differentiable random convolutional feature matching, instantiated via SOCK, enables high-fidelity conditional generation of financial time series from scarce data. The methodology achieves strong quantitative and qualitative performance across synthetic and real-world metrics and exhibits competitive expressive power in discriminative settings. Its robust architectural simplicity and scalability position SOCK as a viable candidate for practical financial modeling and synthetic data applications (2606.05138).

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Tweets

Sign up for free to view the 1 tweet with 9 likes about this paper.