Papers
Topics
Authors
Recent
Search
2000 character limit reached

DARSD: Rep Space Decomposition for UDA

Updated 7 July 2026
  • DARSD is a framework for unsupervised domain adaptation (UDA) that decomposes time-series features into domain-invariant and domain-specific components to improve transfer learning.
  • It leverages adversarial learning, prototypical pseudo-labeling, and a hybrid contrastive optimization strategy to align and refine invariant representations.
  • Empirical results on HAR and industrial fault diagnosis benchmarks demonstrate that DARSD consistently outperforms several state-of-the-art UDA techniques.

DARSD, short for Domain Adaptation via Representation Space Decomposition, is a framework for unsupervised domain adaptation (UDA) on time series. It addresses the setting in which a classifier is trained on a labeled source domain and transferred to an unlabeled target domain with a similar label space but a different data distribution. The central premise is that time-series representations should not be treated as indivisible features to be globally aligned; instead, they should be decomposed into domain-invariant and domain-specific components, with transfer operating primarily on the invariant part. DARSD operationalizes this view through an Adversarial Learnable Common Invariant Basis, a prototypical pseudo-labeling mechanism, and a hybrid contrastive optimization strategy (Cai et al., 28 Jul 2025).

1. Problem setting and motivation

DARSD is formulated for UDA problems with a labeled source dataset

S={(xis,yis)}i=1ns∼Ds\mathcal{S} = \{(x_i^s, y_i^s)\}_{i=1}^{n_s} \sim \mathcal{D}_s

and an unlabeled target dataset

T={xit}i=1nt∼Dt,\mathcal{T} = \{x_i^t\}_{i=1}^{n_t} \sim \mathcal{D}_t,

where each sample satisfies

xis,xit∈RT×D.x_i^s, x_i^t \in \mathbb{R}^{T \times D}.

The source and target share the same label set C\mathcal{C}, but Ds≠Dt\mathcal{D}_s \neq \mathcal{D}_t. The goal is to learn a feature extractor FE(⋅)FE(\cdot) and classifier CLF(⋅)CLF(\cdot) that generalize well on the target domain despite the absence of target labels (Cai et al., 28 Jul 2025).

The paper situates this problem in human activity recognition (HAR) and industrial fault diagnosis. In these settings, domain shift can arise from changes in device type, sensor placement, subject population, operating condition, or environment. The paper argues that time series are especially sensitive to such shifts because low-level signal statistics such as amplitude, noise, and sampling artifacts vary across domains even when higher-level temporal semantics remain comparable. A walking pattern, for example, may retain periodic structure while differing substantially in channel statistics due to a change from one wearable device to another (Cai et al., 28 Jul 2025).

A core critique of prior UDA methods is that they typically align entire feature distributions as though the representation were homogeneous. DARSD is motivated by the claim that such whole-vector alignment can either over-align, suppressing semantic content, or under-align, allowing domain-specific contamination to persist. This suggests a more granular view of the feature space, in which adaptation proceeds by explicitly isolating transferable information from domain-specific residue.

2. Representation space decomposition

DARSD models a feature vector f=FE(x)∈Rdf = FE(x) \in \mathbb{R}^d as the sum of an invariant component and a domain-specific component:

f=finv+fspe.f = f^{inv} + f^{spe}.

The ambient space is assumed to admit an orthogonal decomposition

Rd=Sinv⊕Sspe,\mathbb{R}^d = \mathcal{S}^{inv} \oplus \mathcal{S}^{spe},

where T={xit}i=1nt∼Dt,\mathcal{T} = \{x_i^t\}_{i=1}^{n_t} \sim \mathcal{D}_t,0 contains semantics that should transfer across domains and T={xit}i=1nt∼Dt,\mathcal{T} = \{x_i^t\}_{i=1}^{n_t} \sim \mathcal{D}_t,1 contains domain-dependent artifacts (Cai et al., 28 Jul 2025).

Let T={xit}i=1nt∼Dt,\mathcal{T} = \{x_i^t\}_{i=1}^{n_t} \sim \mathcal{D}_t,2 and T={xit}i=1nt∼Dt,\mathcal{T} = \{x_i^t\}_{i=1}^{n_t} \sim \mathcal{D}_t,3 denote orthonormal bases for the invariant and specific subspaces, respectively. Then

T={xit}i=1nt∼Dt,\mathcal{T} = \{x_i^t\}_{i=1}^{n_t} \sim \mathcal{D}_t,4

Because the true invariant basis is unknown, DARSD introduces a learnable approximation

T={xit}i=1nt∼Dt,\mathcal{T} = \{x_i^t\}_{i=1}^{n_t} \sim \mathcal{D}_t,5

Under the paper’s subspace assumption, projection onto this basis recovers the invariant coordinates:

T={xit}i=1nt∼Dt,\mathcal{T} = \{x_i^t\}_{i=1}^{n_t} \sim \mathcal{D}_t,6

and the invariant reconstruction becomes

T={xit}i=1nt∼Dt,\mathcal{T} = \{x_i^t\}_{i=1}^{n_t} \sim \mathcal{D}_t,7

The theoretical argument in the paper states that if T={xit}i=1nt∼Dt,\mathcal{T} = \{x_i^t\}_{i=1}^{n_t} \sim \mathcal{D}_t,8 spans the same subspace as T={xit}i=1nt∼Dt,\mathcal{T} = \{x_i^t\}_{i=1}^{n_t} \sim \mathcal{D}_t,9, then the domain-specific contribution is eliminated in the projection and the invariant coordinates are recovered exactly (Cai et al., 28 Jul 2025).

In practice, the method adds a softmax filtering step to suppress noisy directions:

xis,xit∈RT×D.x_i^s, x_i^t \in \mathbb{R}^{T \times D}.0

xis,xit∈RT×D.x_i^s, x_i^t \in \mathbb{R}^{T \times D}.1

The reconstructed feature xis,xit∈RT×D.x_i^s, x_i^t \in \mathbb{R}^{T \times D}.2 is then treated as the operative domain-invariant representation for pseudo-labeling and contrastive learning. The paper’s interpretation is that the softmax emphasizes dominant invariant directions while attenuating leakage from domain-specific coordinates.

3. Constituent mechanisms

DARSD is organized around three interacting modules (Cai et al., 28 Jul 2025).

Component Function Output
Adv-LCIB Learns a shared invariant basis and reconstructs invariant features xis,xit∈RT×D.x_i^s, x_i^t \in \mathbb{R}^{T \times D}.3
PPGCE Assigns prototype-based target pseudo-labels with confidence partitioning xis,xit∈RT×D.x_i^s, x_i^t \in \mathbb{R}^{T \times D}.4, xis,xit∈RT×D.x_i^s, x_i^t \in \mathbb{R}^{T \times D}.5
Hybrid contrastive optimization Clusters, regularizes, and aligns source and target invariant features xis,xit∈RT×D.x_i^s, x_i^t \in \mathbb{R}^{T \times D}.6, xis,xit∈RT×D.x_i^s, x_i^t \in \mathbb{R}^{T \times D}.7, xis,xit∈RT×D.x_i^s, x_i^t \in \mathbb{R}^{T \times D}.8

The Adversarial Learnable Common Invariant Basis (Adv-LCIB) is the explicit subspace-learning stage. It projects source and target features into a low-dimensional invariant basis and reconstructs them there. To prevent arbitrary or degenerate reconstructions, DARSD introduces a discriminator xis,xit∈RT×D.x_i^s, x_i^t \in \mathbb{R}^{T \times D}.9 trained to distinguish original features C\mathcal{C}0 from reconstructed features C\mathcal{C}1. The adversarial loss is

C\mathcal{C}2

where C\mathcal{C}3 for original features and C\mathcal{C}4 for reconstructed features. The discriminator is trained to separate C\mathcal{C}5 from C\mathcal{C}6, while the basis-learning mechanism is trained adversarially to make C\mathcal{C}7 informationally close to C\mathcal{C}8. The intended effect is preservation of semantic content within the low-dimensional invariant subspace (Cai et al., 28 Jul 2025).

The Prototypical Pseudo-label Generation with Confidence Evaluation (PPGCE) module operates in that invariant space. For each class C\mathcal{C}9, DARSD maintains a prototype Ds≠Dt\mathcal{D}_s \neq \mathcal{D}_t0 updated from source invariant features by momentum:

Ds≠Dt\mathcal{D}_s \neq \mathcal{D}_t1

Each target invariant feature Ds≠Dt\mathcal{D}_s \neq \mathcal{D}_t2 is pseudo-labeled by maximum cosine similarity to the prototypes:

Ds≠Dt\mathcal{D}_s \neq \mathcal{D}_t3

with confidence

Ds≠Dt\mathcal{D}_s \neq \mathcal{D}_t4

Rather than trusting all pseudo-labels equally, DARSD partitions target samples into a confident subset and a distrusted subset using a time-varying confidence ratio. The paper states that the schedule begins conservatively and expands over training; in the reported implementation, the confidence ratio starts at Ds≠Dt\mathcal{D}_s \neq \mathcal{D}_t5 and increases by Ds≠Dt\mathcal{D}_s \neq \mathcal{D}_t6 every 15 batches (Cai et al., 28 Jul 2025).

The hybrid contrastive optimization strategy uses these partitions differently. Confident target samples are combined with labeled source samples for supervised contrastive aggregation. Distrusted target samples are not discarded; instead, they are optimized through self-supervised consistency and an anti-divergence regularizer that anchors them to source features. This design is meant to reduce pseudo-label error accumulation while still exploiting the full target set.

4. Objective function and training procedure

The supervised contrastive term Ds≠Dt\mathcal{D}_s \neq \mathcal{D}_t7 acts on the combined reliably labeled set

Ds≠Dt\mathcal{D}_s \neq \mathcal{D}_t8

For an anchor Ds≠Dt\mathcal{D}_s \neq \mathcal{D}_t9, positives are features with the same class label and negatives are those with different labels. The loss is defined as

FE(⋅)FE(\cdot)0

Its role is to pull same-class source and target invariant features together while separating different classes (Cai et al., 28 Jul 2025).

For the distrusted target subset FE(⋅)FE(\cdot)1, DARSD applies a self-supervised consistency loss. Each feature is paired with an augmented view produced through standard time-series augmentations such as jittering and scaling. The resulting objective is

FE(⋅)FE(\cdot)2

This term is intended to improve target structure without relying on noisy pseudo-labels (Cai et al., 28 Jul 2025).

The paper also identifies a possible divergence between the supervised branch and the self-supervised branch. To counter this, DARSD adds an anti-divergence term. For each distrusted target feature FE(⋅)FE(\cdot)3, it finds the nearest source invariant feature by cosine similarity,

FE(⋅)FE(\cdot)4

and then applies

FE(⋅)FE(\cdot)5

The stated purpose is to keep distrusted target features anchored to the source distribution in the invariant space and to avoid an “emerging divergence” between differently optimized target subsets (Cai et al., 28 Jul 2025).

The full objective is

FE(⋅)FE(\cdot)6

with FE(⋅)FE(\cdot)7 in the reported experiments. The implementation uses a 4-layer TCN as the shared feature extractor, with hidden size FE(⋅)FE(\cdot)8 and dropout FE(⋅)FE(\cdot)9. The feature dimension is CLF(⋅)CLF(\cdot)0, the invariant subspace dimension is CLF(⋅)CLF(\cdot)1, and the classifier is a 2-layer MLP with hidden size CLF(⋅)CLF(\cdot)2 and dropout CLF(⋅)CLF(\cdot)3 (Cai et al., 28 Jul 2025).

The training loop proceeds by sampling source and target mini-batches, extracting features, reconstructing invariant representations through the learnable basis, updating the discriminator, refreshing class prototypes from source invariant features, partitioning target samples by pseudo-label confidence, and finally minimizing the combined contrastive and adversarial objective. The paper describes final classification as operating on invariant features after this pre-training stage (Cai et al., 28 Jul 2025).

5. Empirical evaluation and ablations

DARSD is evaluated on four benchmark datasets: WISDM, HAR, HHAR, and MFD. The first three are HAR benchmarks with varying device, subject, and placement heterogeneity; MFD is a machine fault diagnosis dataset based on vibration signals from electromechanical drive systems. The paper uses Macro-F1 as the main metric because of class imbalance and evaluates a total of 53 cross-domain scenarios across these datasets (Cai et al., 28 Jul 2025).

The reported comparison includes 12 UDA methods, spanning adversarial, discrepancy-based, metric-learning, and self-supervised baselines. Across the 53 scenarios, DARSD achieves the best performance in 35 cases. The paper further reports that it ranks overall first across all four datasets. On WISDM and HHAR, which exhibit pronounced sensor and subject heterogeneity, DARSD achieves average ranks of 1.42 and 1.31, respectively. On MFD it remains competitive, with an average rank of 2.10 (Cai et al., 28 Jul 2025).

The ablation study is particularly central to the paper’s argument. Removing the LCIB module produces a sharp degradation; for example, on WISDM CLF(⋅)CLF(\cdot)4, Macro-F1 falls to 0.435, compared with 0.883 for the full model. Adding LCIB without the adversarial term raises the same scenario to 0.795, which the paper interprets as evidence that explicit decomposition is fundamental and that adversarial training improves the quality of the invariant basis. The full model also reports strong results on HAR CLF(⋅)CLF(\cdot)5 with Macro-F1 0.967 and on WISDM CLF(⋅)CLF(\cdot)6 with Macro-F1 0.822 (Cai et al., 28 Jul 2025).

Sensitivity analysis indicates that performance is strongest around an invariant basis dimension of CLF(⋅)CLF(\cdot)7 for CLF(⋅)CLF(\cdot)8, which the paper describes as roughly CLF(⋅)CLF(\cdot)9. If f=FE(x)∈Rdf = FE(x) \in \mathbb{R}^d0 is too small, invariant semantics are underrepresented; if it is too large, domain-specific information contaminates the invariant basis. The adversarial-loss weight is reported to be stable in the interval f=FE(x)∈Rdf = FE(x) \in \mathbb{R}^d1: too little adversarial pressure causes information loss in the reconstruction, whereas too much may reintroduce domain-specific content to satisfy the discriminator (Cai et al., 28 Jul 2025).

The qualitative analyses are consistent with the numerical results. t-SNE visualizations show that reconstructed invariant features are more structured than raw features, with tighter class clusters and greater overlap between source and target samples of the same class. Appendix convergence plots indicate that DARSD typically converges within 50–80 epochs, whereas CLUDA requires 100–150 epochs in the comparison reported by the paper (Cai et al., 28 Jul 2025).

6. Interpretation, limitations, and nomenclature

DARSD’s main conceptual contribution is the claim that UDA should be framed not merely as alignment, but as alignment after decomposition. The method is therefore interpretable at the representation level: the invariant basis provides an explicit mechanism for isolating transferable semantics, prototypes define class geometry in that invariant space, and the hybrid contrastive objective assigns distinct roles to confident and distrusted target samples. This suggests a view of UDA in which target adaptation quality depends on the structure of the latent space as much as on inter-domain discrepancy alone.

The paper also identifies several limitations. DARSD assumes a shared label space between source and target and does not address partial or open-set adaptation. It does not explicitly model label shift, only feature shift. The prototype mechanism may be sensitive to source imbalance or label noise. The overall training procedure introduces additional compute and memory overhead through multiple contrastive losses and adversarial optimization. Finally, LCIB is a linear invariant subspace model; the paper notes that this may be restrictive if invariance is highly nonlinear, although the nonlinear encoder is meant to mitigate that issue by mapping inputs into a feature space where linear decomposition is more appropriate (Cai et al., 28 Jul 2025).

The acronym should also be distinguished from several similarly named methods in other areas. DARS in requirements engineering denotes “Data Annotation Requirements Representation and Specification” (Peng et al., 15 Dec 2025). Another DARS designates “Dysarthria-Aware Rhythm-Style Synthesis for ASR Enhancement” (Wu et al., 2 Mar 2026). DAR refers to “Direction-Aware Diagonal Autoregressive Image Generation” (Xu et al., 14 Mar 2025), and DarSwin denotes the “Distortion Aware Radial Swin Transformer” (Athwale et al., 2023). In the time-series UDA literature, however, DARSD specifically denotes Domain Adaptation via Representation Space Decomposition (Cai et al., 28 Jul 2025).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to DARSD.