CauKer: Synthetic Data for TSFMs Pretraining
- CauKer is a synthetic data generation system for time series classification foundation models that exclusively uses non-zero-mean Gaussian Processes and Structural Causal Models to create realistic series.
- It demonstrates clean, monotonic scaling laws in dataset size and model capacity, outperforming pretraining with standard real-world data.
- The framework enables zero-shot pretraining, reducing data curation, privacy concerns, and domain mismatch issues in applications like healthcare and industrial sensing.
Search arXiv for CauKer and related time-series foundation model papers. CauKer is a synthetic data generation and pretraining framework for time series classification foundation models (TSFMs). Its defining claim is unusually specific: state-of-the-art classification TSFMs can be pretrained only on CauKer synthetic data, can nearly match or reach state-of-the-art performance on real benchmarks, and exhibit clean scaling laws in both dataset size and model capacity that are not visible when the same models are pretrained on standard real-world classification corpora (Xie et al., 4 Aug 2025). The framework is explicitly classification-oriented rather than forecasting-oriented: it combines Gaussian Process kernel composition with Structural Causal Models (SCMs) to generate diverse, causally coherent synthetic time series with realistic trends, seasonality, noise, anomalies, and nonlinear cross-variable interactions.
1. Research setting and defining objective
CauKer targets zero-shot time series classification. The intended pipeline is to pretrain an encoder on large unlabeled synthetic corpora, freeze the encoder, and then fit a light downstream classifier on top of the learned embeddings for benchmark datasets such as UCR and UEA (Xie et al., 4 Aug 2025). In this formulation, the downstream classifier is not part of pretraining; supervised labels are used only after pretraining when fitting the probe on real data.
The framework is motivated by several constraints of existing TSFM practice. The strongest existing classification models depend on computationally costly pretraining on large-scale, carefully curated real-world collections, sometimes including the same datasets used later for evaluation. The CauKer formulation treats this as problematic for at least three reasons stated in the source material: data scarcity and curation cost, privacy and regulation in domains such as healthcare or industrial sensing, and the difficulty of studying scaling laws when real corpora have limited diversity or domain mismatch (Xie et al., 4 Aug 2025).
A central conceptual distinction is that CauKer is not a TSFM architecture and not a self-supervised learning objective. It is a synthetic-data generator and pretraining corpus design. Later work using LeNEPA makes this distinction explicit: in that setting, CauKer is the pretraining distribution for a LeNEPA encoder rather than a model component or loss function (Chemeris et al., 1 Jul 2026). This helps avoid a common misunderstanding in secondary discussions, where the synthetic corpus, the encoder, and the training objective are sometimes conflated.
2. Synthetic generation mechanism
CauKer combines two ingredients: Gaussian Process (GP) kernel composition for realistic root time series and an SCM over a random directed acyclic graph (DAG) for causal coupling among variables (Xie et al., 4 Aug 2025). The GP stage produces univariate temporal structure; the SCM stage converts isolated roots into causally coupled multivariate systems.
For each root time series, CauKer samples from a GP prior
using a composite kernel and a non-zero mean. The construction uses a kernel bank , a mean bank , and an activation bank . The kernel bank includes RBF, RationalQuadratic, ExpSineSquared, DotProduct, WhiteKernel, and ConstantKernel. Composite kernels are formed by random additive and multiplicative composition: where , the are sampled from , and each is either 0 or 1. Additive composition supports combinations such as trend plus seasonality, while multiplicative composition supports locally periodic behavior.
A notable design choice is the explicit use of non-zero mean functions, unlike forecasting synthetic generators that often assume zero mean. The means used are linear, exponential, and an “anomaly” mean that is piecewise with occasional spikes where values at random indices follow 2. The root GP for variable 3 is
4
producing a time series 5, with 6 in the reported experiments.
The SCM stage overlays a random DAG 7 with 8 nodes, 9 edges, and 0 root nodes. Root nodes receive independent GP time series. Each edge is assigned an activation from 1, whose contents are affine linear 2, ReLU, LeakyReLU, Sigmoid, Sine, and elementwise modulo 3, with the parameter ranges given in the source material. For a non-root node 4, the structural assignment is
5
with 6. This instantiates SCM equations of the form
7
This design lets CauKer represent trend, seasonality, non-stationarity, noise and roughness, causal cross-variable interactions, and abrupt changes or anomalies (Xie et al., 4 Aug 2025). The source material also reports a qualitative diagnostic: hierarchical clustering on pairwise DTW distances of 200 CauKer series reveals clear clusters and anomalous outliers, which is consistent with a classification-oriented corpus rather than a purely smooth forecasting corpus.
3. Pretraining regime and supported TSFMs
CauKer is used to pretrain two families of TSFMs that instantiate the two main SSL paradigms studied in the paper: Mantis, an encoder-only contrastive model, and MOMENT, a T5-like masked encoder-decoder model (Xie et al., 4 Aug 2025). In both cases pretraining is task-agnostic: no synthetic class labels are used.
Mantis is a Vision Transformer-like encoder specialized for 1D series, with a parameter range from approximately 8M to 9M and a default 8M configuration. It is pretrained with contrastive SSL in the style of SimCLR or TS2Vec. The setup uses augmentations 0, an encoder 1, a projection head 2, cosine similarity
3
and a temperature 4. After pretraining, the encoder is frozen and a Random Forest is trained on the embeddings for each UCR dataset.
MOMENT uses flan-T5-small, base, and large encoders with 77M, 248M, and 783M parameters, respectively. It is pretrained by masked reconstruction. A series 5 is chunked into 6 patches of length 7, with 8; some patches are replaced by a learnable 9 embedding; and the decoder reconstructs the original series. The loss is MSE over masked patches: 0 After pretraining, the encoder is frozen and an SVM is trained on embeddings for each UCR dataset.
The paper emphasizes that CauKer is compatible with multiple architectures and pretraining approaches rather than being tied to one model family. This suggests that the main object of study is the pretraining substrate—the synthetic corpus and its structure—rather than an architecture-specific inductive bias.
4. Scaling behavior and transfer performance
The empirical centerpiece of CauKer is its scaling behavior. When the same classification TSFMs are pretrained on CauKer synthetic data, the paper reports monotonic, regular scaling laws in both dataset size and model size; when the same models are pretrained on real UEA data, scaling is described as non-monotonic, flat, or irregular (Xie et al., 4 Aug 2025).
For data scaling, the synthetic regime spans 10K, 50K, 100K, 500K, 1M, 5M, and 10M CauKer series. Mantis improves from 76.91% at 10K to 79.09% at 10M. MOMENT (77M) improves from 74.24% at 100K to 77.49% at 10M. By contrast, Mantis pretrained on UEA subsets moves from 75.67% at 0.1% UEA to 76.33% at 29% UEA and then drops to 71.93% at 100% UEA. For MOMENT, accuracy fluctuates around approximately 70–72% for 1–100% UEA data. The authors hypothesize two causes for this real-data behavior: domain mismatch from UEA to UCR and limited diversity in UEA, where additional samples may duplicate patterns rather than improve generalization.
For model scaling, CauKer again yields orderly behavior. With CauKer 1M, MOMENT accuracy rises from 75.21% at 77M parameters to 76.16% at 248M and 77.20% at 783M. For Mantis, increasing model size from 0.75M through 114.14M shows overall improving accuracy, with saturation around 28M–114M and 10M training examples. Under UEA pretraining, increasing model capacity is reported as unstable or even detrimental; one example given is MOMENT 783M on UEA 100%, which reaches only 66.07%.
A further scaling result concerns training time. On CauKer 1M synthetic data, both Mantis and MOMENT show steady improvement with more epochs; on UEA 10%, further training gives flat or noisy curves. The paper interprets this as evidence that CauKer provides a cleaner and more learnable pretraining distribution.
The transfer claim is also quantitative. For Mantis (8M), pretraining on the original real-data corpus of approximately 1.89M real series yields 78.66% average UCR accuracy, while pretraining on 100K CauKer synthetic series yields 78.55%, with no UCR overlap. For MOMENT (77M), the original model pretrained on 13M real series reaches 78.85%, while the CauKer-pretrained reimplementation reaches 77.49% at 10M synthetic series. The paper summarizes this as synthetic-only pretraining being within approximately one point of state-of-the-art real-data TSFMs (Xie et al., 4 Aug 2025).
5. Comparative analyses, ablations, and representation diagnostics
Ablative comparison at 100K pretraining samples clarifies which ingredients of CauKer matter. The source material compares pure SCM generation, forecasting-oriented synthetic generators, kernel-only generators, mean-augmented kernel generators, full CauKer, and real-data subsets (Xie et al., 4 Aug 2025).
| Pretraining source at 100K | Mantis | MOMENT |
|---|---|---|
| SCM-only | 73.49% | 59.23% |
| FPFN | 77.52% | 70.85% |
| KernelSynth | 77.70% | 69.31% |
| Mean+KernelSynth | 78.20% | 72.56% |
| CauKer | 78.31% | 74.24% |
| UEA real data | 76.73% | 73.55% |
| Forecasting real data | 75.81% | 73.93% |
Two conclusions are explicit in these numbers. First, SCM-only generation underperforms strongly, especially for MOMENT, indicating that causal coupling without realistic temporal motifs is insufficient. Second, Mean+KernelSynth improves on kernel-only generation, showing that non-zero means materially help classification-oriented pretraining. Full CauKer (GP + non-zero mean + SCM) gives the best synthetic corpus among the reported synthetic alternatives for both architectures.
The paper also includes internal representation diagnostics. Using the AffineOT non-linearity measure, Mantis pretrained on CauKer shows increasing non-linearity as dataset size grows, whereas the same measure is almost constant when Mantis is pretrained on UEA. Using CKA similarity between hidden layers, CauKer pretraining shows greater representational reorganization as dataset size increases, while UEA pretraining changes little from 600K to 12M (Xie et al., 4 Aug 2025). These analyses are not presented as proof of causality, but they do support the claim that additional CauKer data remains usable by the model in a way that additional UEA data often does not.
6. Subsequent use, misconceptions, limitations, and outlook
Subsequent work uses CauKer as an external pretraining corpus rather than reimplementing its internals. In LeNEPA, for example, the model instance LeNEPA-CauKer is trained on the CauKer corpus, then frozen and evaluated on the full UCR-128 benchmark with a Random Forest probe (Chemeris et al., 1 Jul 2026). In that protocol, a single-seed, best-checkpoint run reaches 77.65% mean UCR-128 Random-Forest accuracy, which is reported as within 1.16 points of Mantis and within 0.24 points of MOMENT (77.89%). The authors of LeNEPA explicitly treat this result as an existence proof and not as a definitive leaderboard claim, because it uses one pretraining seed and a best checkpoint over the trajectory.
This later use also clarifies a recurring misconception: CauKer is not a latent-prediction objective, not an augmentation policy, and not a transformer backbone. It is the synthetic corpus and generation framework on which such models may be pretrained (Chemeris et al., 1 Jul 2026).
The limitations listed for CauKer are concrete. The original study investigates only Mantis and MOMENT in depth; it focuses on classification, not forecasting; the SCM is relatively simple, using linear aggregation after elementwise activations with Gaussian weights; GP kernels are best suited to relatively short and stationary-ish segments; and the synthetic distribution remains generic rather than domain-specific for ECG, EEG, finance, or industrial process data (Xie et al., 4 Aug 2025). The authors suggest extending CauKer to forecasting TSFMs, integrating more domain-specific simulators, fitting explicit theoretical scaling laws, combining CauKer with real data in hybrid corpora, and using CauKer for controlled distribution-shift and intervention studies.
Taken together, the reported results position CauKer as a classification-oriented synthetic pretraining substrate whose importance lies less in any single architectural novelty than in a specific synthesis: non-zero-mean GP temporal structure plus SCM-based causal coupling. The empirical implication suggested by the available evidence is that properly structured synthetic corpora can make time-series foundation model pretraining more sample-efficient, more controllable, and more interpretable in scaling terms than standard real-world classification corpora (Xie et al., 4 Aug 2025).