---
title: 'CauKer: Synthetic Data for TSFMs Pretraining'
url: https://www.emergentmind.com/topics/cauker
type: topic
---

# CauKer: Synthetic Data for TSFMs Pretraining

Search arXiv for CauKer and related time-series foundation model papers.
CauKer is a synthetic data generation and pretraining framework for **time series classification foundation models (TSFMs)**. Its defining claim is unusually specific: state-of-the-art classification TSFMs can be pretrained **only on CauKer synthetic data**, can nearly match or reach state-of-the-art performance on real benchmarks, and exhibit **clean scaling laws** in both dataset size and model capacity that are not visible when the same models are pretrained on standard real-world classification corpora [2508.02879]. The framework is explicitly classification-oriented rather than forecasting-oriented: it combines Gaussian Process kernel composition with Structural Causal Models (SCMs) to generate diverse, causally coherent synthetic time series with realistic trends, seasonality, noise, anomalies, and nonlinear cross-variable interactions.

## 1. Research setting and defining objective

CauKer targets **zero-shot time series classification**. The intended pipeline is to pretrain an encoder \(F : \mathbb{R}^t \to \mathbb{R}^q\) on large unlabeled synthetic corpora, freeze the encoder, and then fit a light downstream classifier on top of the learned embeddings for benchmark datasets such as UCR and UEA [2508.02879]. In this formulation, the downstream classifier is not part of pretraining; supervised labels are used only after pretraining when fitting the probe on real data.

The framework is motivated by several constraints of existing TSFM practice. The strongest existing classification models depend on computationally costly pretraining on large-scale, carefully curated real-world collections, sometimes including the same datasets used later for evaluation. The CauKer formulation treats this as problematic for at least three reasons stated in the source material: data scarcity and curation cost, privacy and regulation in domains such as healthcare or industrial sensing, and the difficulty of studying scaling laws when real corpora have limited diversity or domain mismatch [2508.02879].

A central conceptual distinction is that CauKer is **not** a TSFM architecture and **not** a self-supervised learning objective. It is a synthetic-data generator and pretraining corpus design. Later work using LeNEPA makes this distinction explicit: in that setting, CauKer is the **pretraining distribution** for a LeNEPA encoder rather than a model component or loss function [2607.00958]. This helps avoid a common misunderstanding in secondary discussions, where the synthetic corpus, the encoder, and the training objective are sometimes conflated.

## 2. Synthetic generation mechanism

CauKer combines two ingredients: **Gaussian Process (GP) kernel composition** for realistic root time series and an **SCM over a random directed acyclic graph (DAG)** for causal coupling among variables [2508.02879]. The GP stage produces univariate temporal structure; the SCM stage converts isolated roots into causally coupled multivariate systems.

For each root time series, CauKer samples from a GP prior
\[
f(t) \sim \mathcal{GP}\bigl(\mu(t), k(t,t')\bigr),
\]
using a composite kernel and a non-zero mean. The construction uses a kernel bank \(\mathcal{K}\), a mean bank \(\mathcal{M}\), and an activation bank \(\mathcal{A}\). The kernel bank includes RBF, RationalQuadratic, ExpSineSquared, DotProduct, WhiteKernel, and ConstantKernel. Composite kernels are formed by random additive and multiplicative composition:
\[
k^\ast(t,t') = \kappa_1(t,t') \star_1 \kappa_2(t,t') \star_2 \dots \star_{K-1} \kappa_K(t,t'),
\]
where \(K \sim \mathcal{U}(1, n_\mathcal{K})\), the \(\kappa_i\) are sampled from \(\mathcal{K}\), and each \(\star_i\) is either \(+\) or \(\times\). Additive composition supports combinations such as trend plus seasonality, while multiplicative composition supports locally periodic behavior.

A notable design choice is the explicit use of **non-zero mean functions**, unlike forecasting synthetic generators that often assume zero mean. The means used are linear, exponential, and an “anomaly” mean that is piecewise with occasional spikes where values at random indices follow \(\mathcal{U}(-5, 5)\). The root GP for variable \(i\) is
\[
f_i(t) \sim \mathcal{GP}\bigl(\mu_i(t), k_i^\ast(t,t')\bigr),
\]
producing a time series \(t_i \in \mathbb{R}^L\), with \(L=512\) in the reported experiments.

The SCM stage overlays a random DAG \((\mathcal{V}, \mathcal{E})\) with \(V = |\mathcal{V}|\) nodes, \(E = |\mathcal{E}|\) edges, and \(M < V\) root nodes. Root nodes receive independent GP time series. Each edge is assigned an activation from \(\mathcal{A}\), whose contents are affine linear \(a x + b\), ReLU, LeakyReLU, Sigmoid, Sine, and elementwise modulo \(x \bmod c\), with the parameter ranges given in the source material. For a non-root node \(v_j\), the structural assignment is
\[
t_{v_j} = W_j \cdot \bigl[ \, \phi(e_{i_1 j})(t_{u_{i_1}}), \dots, \phi(e_{i_r j})(t_{u_{i_r}}) \bigr] + b_j,
\]
with \(W_j, b_j \sim \mathcal{N}(0,1)\). This instantiates SCM equations of the form
\[
X_j(t) = f_j\bigl(\{X_i(t)\}_{i \in \mathrm{Pa}(j)}\bigr).
\]

This design lets CauKer represent trend, seasonality, non-stationarity, noise and roughness, causal cross-variable interactions, and abrupt changes or anomalies [2508.02879]. The source material also reports a qualitative diagnostic: hierarchical clustering on pairwise DTW distances of 200 CauKer series reveals clear clusters and anomalous outliers, which is consistent with a classification-oriented corpus rather than a purely smooth forecasting corpus.

## 3. Pretraining regime and supported TSFMs

CauKer is used to pretrain two families of TSFMs that instantiate the two main SSL paradigms studied in the paper: **Mantis**, an encoder-only contrastive model, and **MOMENT**, a T5-like masked encoder-decoder model [2508.02879]. In both cases pretraining is **task-agnostic**: no synthetic class labels are used.

Mantis is a Vision Transformer-like encoder specialized for 1D series, with a parameter range from approximately \(0.75\)M to \(114\)M and a default 8M configuration. It is pretrained with contrastive SSL in the style of SimCLR or TS2Vec. The setup uses augmentations \(\phi,\psi\), an encoder \(F\), a projection head \(g\), cosine similarity
\[
s_{\cos}(\mathbf{a},\mathbf{b}) =
\frac{\mathbf{a}^\top \mathbf{b}}{\|\mathbf{a}\| \|\mathbf{b}\|},
\]
and a temperature \(T=0.1\). After pretraining, the encoder is frozen and a **Random Forest** is trained on the embeddings for each UCR dataset.

MOMENT uses flan-T5-small, base, and large encoders with 77M, 248M, and 783M parameters, respectively. It is pretrained by masked reconstruction. A series \(\mathcal{T} \in \mathbb{R}^{1 \times T}\) is chunked into \(N\) patches of length \(P\), with \(T = N \cdot P\); some patches are replaced by a learnable \([MASK]\) embedding; and the decoder reconstructs the original series. The loss is MSE over masked patches:
\[
\mathcal{L}_{\text{masked}} =
\frac{1}{|\Omega|}
\sum_{n \in \Omega}
\left\| \mathcal{T}_n - h_{\text{rec}}(F(\text{[MASK]}))_n \right\|^2.
\]
After pretraining, the encoder is frozen and an **SVM** is trained on embeddings for each UCR dataset.

The paper emphasizes that CauKer is compatible with multiple architectures and pretraining approaches rather than being tied to one model family. This suggests that the main object of study is the **pretraining substrate**—the synthetic corpus and its structure—rather than an architecture-specific inductive bias.

## 4. Scaling behavior and transfer performance

The empirical centerpiece of CauKer is its scaling behavior. When the same classification TSFMs are pretrained on CauKer synthetic data, the paper reports **monotonic, regular scaling laws** in both dataset size and model size; when the same models are pretrained on real UEA data, scaling is described as **non-monotonic**, **flat**, or **irregular** [2508.02879].

For data scaling, the synthetic regime spans 10K, 50K, 100K, 500K, 1M, 5M, and 10M CauKer series. Mantis improves from **76.91%** at 10K to **79.09%** at 10M. MOMENT (77M) improves from **74.24%** at 100K to **77.49%** at 10M. By contrast, Mantis pretrained on UEA subsets moves from **75.67%** at 0.1% UEA to **76.33%** at 29% UEA and then drops to **71.93%** at 100% UEA. For MOMENT, accuracy fluctuates around approximately **70–72%** for 1–100% UEA data. The authors hypothesize two causes for this real-data behavior: **domain mismatch** from UEA to UCR and **limited diversity** in UEA, where additional samples may duplicate patterns rather than improve generalization.

For model scaling, CauKer again yields orderly behavior. With CauKer 1M, MOMENT accuracy rises from **75.21%** at 77M parameters to **76.16%** at 248M and **77.20%** at 783M. For Mantis, increasing model size from **0.75M** through **114.14M** shows overall improving accuracy, with saturation around **28M–114M** and **10M** training examples. Under UEA pretraining, increasing model capacity is reported as unstable or even detrimental; one example given is MOMENT **783M** on UEA 100%, which reaches only **66.07%**.

A further scaling result concerns training time. On CauKer 1M synthetic data, both Mantis and MOMENT show steady improvement with more epochs; on UEA 10%, further training gives flat or noisy curves. The paper interprets this as evidence that CauKer provides a cleaner and more learnable pretraining distribution.

The transfer claim is also quantitative. For **Mantis (8M)**, pretraining on the original real-data corpus of approximately **1.89M** real series yields **78.66%** average UCR accuracy, while pretraining on **100K** CauKer synthetic series yields **78.55%**, with no UCR overlap. For **MOMENT (77M)**, the original model pretrained on **13M** real series reaches **78.85%**, while the CauKer-pretrained reimplementation reaches **77.49%** at **10M** synthetic series. The paper summarizes this as synthetic-only pretraining being within approximately one point of state-of-the-art real-data TSFMs [2508.02879].

## 5. Comparative analyses, ablations, and representation diagnostics

Ablative comparison at **100K** pretraining samples clarifies which ingredients of CauKer matter. The source material compares pure SCM generation, forecasting-oriented synthetic generators, kernel-only generators, mean-augmented kernel generators, full CauKer, and real-data subsets [2508.02879].

| Pretraining source at 100K | Mantis | MOMENT |
|---|---:|---:|
| SCM-only | 73.49% | 59.23% |
| FPFN | 77.52% | 70.85% |
| KernelSynth | 77.70% | 69.31% |
| Mean+KernelSynth | 78.20% | 72.56% |
| CauKer | 78.31% | 74.24% |
| UEA real data | 76.73% | 73.55% |
| Forecasting real data | 75.81% | 73.93% |

Two conclusions are explicit in these numbers. First, **SCM-only** generation underperforms strongly, especially for MOMENT, indicating that causal coupling without realistic temporal motifs is insufficient. Second, **Mean+KernelSynth** improves on kernel-only generation, showing that non-zero means materially help classification-oriented pretraining. Full **CauKer (GP + non-zero mean + SCM)** gives the best synthetic corpus among the reported synthetic alternatives for both architectures.

The paper also includes internal representation diagnostics. Using the **AffineOT** non-linearity measure, Mantis pretrained on CauKer shows increasing non-linearity as dataset size grows, whereas the same measure is almost constant when Mantis is pretrained on UEA. Using **CKA** similarity between hidden layers, CauKer pretraining shows greater representational reorganization as dataset size increases, while UEA pretraining changes little from 600K to 12M [2508.02879]. These analyses are not presented as proof of causality, but they do support the claim that additional CauKer data remains usable by the model in a way that additional UEA data often does not.

## 6. Subsequent use, misconceptions, limitations, and outlook

Subsequent work uses CauKer as an **external pretraining corpus** rather than reimplementing its internals. In LeNEPA, for example, the model instance **LeNEPA-CauKer** is trained on the CauKer corpus, then frozen and evaluated on the full **UCR-128** benchmark with a Random Forest probe [2607.00958]. In that protocol, a single-seed, best-checkpoint run reaches **77.65%** mean UCR-128 Random-Forest accuracy, which is reported as within **1.16** points of **Mantis** and within **0.24** points of **MOMENT (77.89%)**. The authors of LeNEPA explicitly treat this result as an **existence proof** and not as a definitive leaderboard claim, because it uses one pretraining seed and a best checkpoint over the trajectory.

This later use also clarifies a recurring misconception: CauKer is not a latent-prediction objective, not an augmentation policy, and not a transformer backbone. It is the synthetic corpus and generation framework on which such models may be pretrained [2607.00958].

The limitations listed for CauKer are concrete. The original study investigates only **Mantis** and **MOMENT** in depth; it focuses on **classification**, not forecasting; the SCM is relatively simple, using linear aggregation after elementwise activations with Gaussian weights; GP kernels are best suited to relatively short and stationary-ish segments; and the synthetic distribution remains generic rather than domain-specific for ECG, EEG, finance, or industrial process data [2508.02879]. The authors suggest extending CauKer to forecasting TSFMs, integrating more domain-specific simulators, fitting explicit theoretical scaling laws, combining CauKer with real data in hybrid corpora, and using CauKer for controlled distribution-shift and intervention studies.

Taken together, the reported results position CauKer as a classification-oriented synthetic pretraining substrate whose importance lies less in any single architectural novelty than in a specific synthesis: **non-zero-mean GP temporal structure plus SCM-based causal coupling**. The empirical implication suggested by the available evidence is that properly structured synthetic corpora can make time-series foundation model pretraining more sample-efficient, more controllable, and more interpretable in scaling terms than standard real-world classification corpora [2508.02879].

Source: https://www.emergentmind.com/topics/cauker