---
title: 'DARSD: Rep Space Decomposition for UDA'
url: https://www.emergentmind.com/topics/darsd
type: topic
---

# DARSD: Rep Space Decomposition for UDA

DARSD, short for **Domain Adaptation via Representation Space Decomposition**, is a framework for **unsupervised domain adaptation (UDA) on time series**. It addresses the setting in which a classifier is trained on a labeled **source** domain and transferred to an **unlabeled target** domain with a similar label space but a different data distribution. The central premise is that time-series representations should not be treated as indivisible features to be globally aligned; instead, they should be decomposed into **domain-invariant** and **domain-specific** components, with transfer operating primarily on the invariant part. DARSD operationalizes this view through an **Adversarial Learnable Common Invariant Basis**, a **prototypical pseudo-labeling mechanism**, and a **hybrid contrastive optimization strategy** [2507.20968].

## 1. Problem setting and motivation

DARSD is formulated for UDA problems with a labeled source dataset
$$
\mathcal{S} = \{(x_i^s, y_i^s)\}_{i=1}^{n_s} \sim \mathcal{D}_s
$$
and an unlabeled target dataset
$$
\mathcal{T} = \{x_i^t\}_{i=1}^{n_t} \sim \mathcal{D}_t,
$$
where each sample satisfies
$$
x_i^s, x_i^t \in \mathbb{R}^{T \times D}.
$$
The source and target share the same label set \(\mathcal{C}\), but \(\mathcal{D}_s \neq \mathcal{D}_t\). The goal is to learn a feature extractor \(FE(\cdot)\) and classifier \(CLF(\cdot)\) that generalize well on the target domain despite the absence of target labels [2507.20968].

The paper situates this problem in **human activity recognition (HAR)** and **industrial fault diagnosis**. In these settings, domain shift can arise from changes in device type, sensor placement, subject population, operating condition, or environment. The paper argues that time series are especially sensitive to such shifts because low-level signal statistics such as amplitude, noise, and sampling artifacts vary across domains even when higher-level temporal semantics remain comparable. A walking pattern, for example, may retain periodic structure while differing substantially in channel statistics due to a change from one wearable device to another [2507.20968].

A core critique of prior UDA methods is that they typically align entire feature distributions as though the representation were homogeneous. DARSD is motivated by the claim that such whole-vector alignment can either **over-align**, suppressing semantic content, or **under-align**, allowing domain-specific contamination to persist. This suggests a more granular view of the feature space, in which adaptation proceeds by explicitly isolating transferable information from domain-specific residue.

## 2. Representation space decomposition

DARSD models a feature vector \(f = FE(x) \in \mathbb{R}^d\) as the sum of an invariant component and a domain-specific component:
$$
f = f^{inv} + f^{spe}.
$$
The ambient space is assumed to admit an orthogonal decomposition
$$
\mathbb{R}^d = \mathcal{S}^{inv} \oplus \mathcal{S}^{spe},
$$
where \(\mathcal{S}^{inv}\) contains semantics that should transfer across domains and \(\mathcal{S}^{spe}\) contains domain-dependent artifacts [2507.20968].

Let \(B^{inv} \in \mathbb{R}^{d \times m}\) and \(B^{spe} \in \mathbb{R}^{d \times \bar m}\) denote orthonormal bases for the invariant and specific subspaces, respectively. Then
$$
f^{inv} = B^{inv} w^{inv}, \qquad
f^{spe} = B^{spe} w^{spe}.
$$
Because the true invariant basis is unknown, DARSD introduces a learnable approximation
$$
B^{inv}_{obs} \in \mathbb{R}^{d \times m}.
$$
Under the paper’s subspace assumption, projection onto this basis recovers the invariant coordinates:
$$
w^{inv} = (B^{inv}_{obs})^\top f,
$$
and the invariant reconstruction becomes
$$
\hat f = B^{inv}_{obs} w^{inv}.
$$
The theoretical argument in the paper states that if \(B^{inv}_{obs}\) spans the same subspace as \(B^{inv}\), then the domain-specific contribution is eliminated in the projection and the invariant coordinates are recovered exactly [2507.20968].

In practice, the method adds a softmax filtering step to suppress noisy directions:
$$
\hat w = \operatorname{Softmax}\big((B^{inv}_{obs})^\top FE(x)\big),
$$
$$
\hat f = B^{inv}_{obs}\hat w.
$$
The reconstructed feature \(\hat f\) is then treated as the operative domain-invariant representation for pseudo-labeling and contrastive learning. The paper’s interpretation is that the softmax emphasizes dominant invariant directions while attenuating leakage from domain-specific coordinates.

## 3. Constituent mechanisms

DARSD is organized around three interacting modules [2507.20968].

| Component | Function | Output |
|---|---|---|
| Adv-LCIB | Learns a shared invariant basis and reconstructs invariant features | \(\hat f\) |
| PPGCE | Assigns prototype-based target pseudo-labels with confidence partitioning | \(\hat F_t^{con}\), \(\hat F_t^{dis}\) |
| Hybrid contrastive optimization | Clusters, regularizes, and aligns source and target invariant features | \(\mathcal{L}_{sup}\), \(\mathcal{L}_{self}\), \(\mathcal{L}_{anti}\) |

The **Adversarial Learnable Common Invariant Basis (Adv-LCIB)** is the explicit subspace-learning stage. It projects source and target features into a low-dimensional invariant basis and reconstructs them there. To prevent arbitrary or degenerate reconstructions, DARSD introduces a discriminator \(D(\cdot)\) trained to distinguish original features \(f\) from reconstructed features \(\hat f\). The adversarial loss is
$$
\mathcal{L}^{adv} =
\mathbb{E}_{\bar f_i \in \{F \cup \hat F\}}
\Big[
z_i \log D(\bar f_i) + (1-z_i)\log(1-D(\bar f_i))
\Big],
$$
where \(z_i=1\) for original features and \(z_i=0\) for reconstructed features. The discriminator is trained to separate \(f\) from \(\hat f\), while the basis-learning mechanism is trained adversarially to make \(\hat f\) informationally close to \(f\). The intended effect is preservation of semantic content within the low-dimensional invariant subspace [2507.20968].

The **Prototypical Pseudo-label Generation with Confidence Evaluation (PPGCE)** module operates in that invariant space. For each class \(c\), DARSD maintains a prototype \(p_c\) updated from source invariant features by momentum:
$$
p_c^{t+1} = \mu p_c^t + (1-\mu)\frac{1}{|\mathcal{S}_c|}
\sum_{\hat f_i^s \in \mathcal{S}_c}\hat f_i^s.
$$
Each target invariant feature \(\hat f_i^t\) is pseudo-labeled by maximum cosine similarity to the prototypes:
$$
y_i^{psd} = \arg\max_c \cos(\hat f_i^t, p_c),
$$
with confidence
$$
\sigma_i = \max_c \cos(\hat f_i^t, p_c).
$$
Rather than trusting all pseudo-labels equally, DARSD partitions target samples into a **confident subset** and a **distrusted subset** using a time-varying confidence ratio. The paper states that the schedule begins conservatively and expands over training; in the reported implementation, the confidence ratio starts at \(0.1\) and increases by \(0.05\) every 15 batches [2507.20968].

The **hybrid contrastive optimization strategy** uses these partitions differently. Confident target samples are combined with labeled source samples for supervised contrastive aggregation. Distrusted target samples are not discarded; instead, they are optimized through self-supervised consistency and an anti-divergence regularizer that anchors them to source features. This design is meant to reduce pseudo-label error accumulation while still exploiting the full target set.

## 4. Objective function and training procedure

The supervised contrastive term \(\mathcal{L}_{sup}\) acts on the combined reliably labeled set
$$
\hat F^{lab} = \hat F_s \cup \hat F_t^{con}.
$$
For an anchor \(\hat f_i\), positives are features with the same class label and negatives are those with different labels. The loss is defined as
$$
\mathcal{L}_{sup}
=
\mathbb{E}_{\hat f_i \in \hat F_s \cup \hat F_t^{con}}
\Bigg[
\mathbb{E}_{\hat f_p \in \mathcal{P}(i)}
-\log
\frac{\exp(\cos(\hat f_i,\hat f_p)/\tau)}
{\sum_{\hat f_n \in \mathcal{N}(i)} \exp(\cos(\hat f_i,\hat f_n)/\tau)}
\Bigg].
$$
Its role is to pull same-class source and target invariant features together while separating different classes [2507.20968].

For the distrusted target subset \(\hat F_t^{dis}\), DARSD applies a self-supervised consistency loss. Each feature is paired with an augmented view produced through standard time-series augmentations such as jittering and scaling. The resulting objective is
$$
\mathcal{L}_{self}
=
\mathbb{E}_{\hat f_i^{dis} \in \hat F_t^{dis}}
\Bigg[
-\log
\frac{\exp(\cos(\hat f_i^{dis}, \hat f_i^{dis+})/\tau)}
{\sum_{j \neq i}\exp(\cos(\hat f_i^{dis}, \hat f_j^{dis})/\tau)}
\Bigg].
$$
This term is intended to improve target structure without relying on noisy pseudo-labels [2507.20968].

The paper also identifies a possible divergence between the supervised branch and the self-supervised branch. To counter this, DARSD adds an anti-divergence term. For each distrusted target feature \(\hat f_i^{dis}\), it finds the nearest source invariant feature by cosine similarity,
$$
PS(\hat f_i^{dis}, \hat F_s)
=
\arg\max_{\hat f_j^s \in \hat F_s} \cos(\hat f_i^{dis}, \hat f_j^s),
$$
and then applies
$$
\mathcal{L}_{anti}
=
\mathbb{E}_{\hat f_i^{dis} \in \hat F_t^{dis}}
\Bigg[
-\log
\frac{\exp(\cos(\hat f_i^{dis}, PS(\hat f_i^{dis}, \hat F_s))/\tau)}
{\sum_{\hat f_j^s \in \hat F_s}\exp(\cos(\hat f_i^{dis}, \hat f_j^s)/\tau)}
\Bigg].
$$
The stated purpose is to keep distrusted target features anchored to the source distribution in the invariant space and to avoid an “emerging divergence” between differently optimized target subsets [2507.20968].

The full objective is
$$
\mathcal{L}_{total}
=
\mathcal{L}_{sup}
+
\mathcal{L}_{self}
+
\lambda_1 \mathcal{L}_{anti}
+
\lambda_2 \mathcal{L}_{adv},
$$
with \(\lambda_1=\lambda_2=0.5\) in the reported experiments. The implementation uses a **4-layer TCN** as the shared feature extractor, with hidden size \(128\) and dropout \(0.2\). The feature dimension is \(d=128\), the invariant subspace dimension is \(m=24\), and the classifier is a **2-layer MLP** with hidden size \(128\) and dropout \(0.1\) [2507.20968].

The training loop proceeds by sampling source and target mini-batches, extracting features, reconstructing invariant representations through the learnable basis, updating the discriminator, refreshing class prototypes from source invariant features, partitioning target samples by pseudo-label confidence, and finally minimizing the combined contrastive and adversarial objective. The paper describes final classification as operating on invariant features after this pre-training stage [2507.20968].

## 5. Empirical evaluation and ablations

DARSD is evaluated on four benchmark datasets: **WISDM**, **HAR**, **HHAR**, and **MFD**. The first three are HAR benchmarks with varying device, subject, and placement heterogeneity; MFD is a machine fault diagnosis dataset based on vibration signals from electromechanical drive systems. The paper uses **Macro-F1** as the main metric because of class imbalance and evaluates a total of **53 cross-domain scenarios** across these datasets [2507.20968].

The reported comparison includes **12 UDA methods**, spanning adversarial, discrepancy-based, metric-learning, and self-supervised baselines. Across the 53 scenarios, DARSD achieves the best performance in **35** cases. The paper further reports that it ranks overall first across all four datasets. On WISDM and HHAR, which exhibit pronounced sensor and subject heterogeneity, DARSD achieves average ranks of **1.42** and **1.31**, respectively. On MFD it remains competitive, with an average rank of **2.10** [2507.20968].

The ablation study is particularly central to the paper’s argument. Removing the LCIB module produces a sharp degradation; for example, on **WISDM \(12 \to 19\)**, Macro-F1 falls to **0.435**, compared with **0.883** for the full model. Adding LCIB without the adversarial term raises the same scenario to **0.795**, which the paper interprets as evidence that explicit decomposition is fundamental and that adversarial training improves the quality of the invariant basis. The full model also reports strong results on **HAR \(20 \to 6\)** with Macro-F1 **0.967** and on **WISDM \(2 \to 28\)** with Macro-F1 **0.822** [2507.20968].

Sensitivity analysis indicates that performance is strongest around an invariant basis dimension of \(m=24\) for \(d=128\), which the paper describes as roughly \(0.2d\). If \(m\) is too small, invariant semantics are underrepresented; if it is too large, domain-specific information contaminates the invariant basis. The adversarial-loss weight is reported to be stable in the interval \([0.25, 0.50]\): too little adversarial pressure causes information loss in the reconstruction, whereas too much may reintroduce domain-specific content to satisfy the discriminator [2507.20968].

The qualitative analyses are consistent with the numerical results. t-SNE visualizations show that reconstructed invariant features are more structured than raw features, with tighter class clusters and greater overlap between source and target samples of the same class. Appendix convergence plots indicate that DARSD typically converges within **50–80 epochs**, whereas CLUDA requires **100–150 epochs** in the comparison reported by the paper [2507.20968].

## 6. Interpretation, limitations, and nomenclature

DARSD’s main conceptual contribution is the claim that UDA should be framed not merely as **alignment**, but as **alignment after decomposition**. The method is therefore interpretable at the representation level: the invariant basis provides an explicit mechanism for isolating transferable semantics, prototypes define class geometry in that invariant space, and the hybrid contrastive objective assigns distinct roles to confident and distrusted target samples. This suggests a view of UDA in which target adaptation quality depends on the structure of the latent space as much as on inter-domain discrepancy alone.

The paper also identifies several limitations. DARSD assumes a **shared label space** between source and target and does not address partial or open-set adaptation. It does not explicitly model **label shift**, only feature shift. The prototype mechanism may be sensitive to source imbalance or label noise. The overall training procedure introduces additional compute and memory overhead through multiple contrastive losses and adversarial optimization. Finally, LCIB is a **linear invariant subspace** model; the paper notes that this may be restrictive if invariance is highly nonlinear, although the nonlinear encoder is meant to mitigate that issue by mapping inputs into a feature space where linear decomposition is more appropriate [2507.20968].

The acronym should also be distinguished from several similarly named methods in other areas. **DARS** in requirements engineering denotes “Data Annotation Requirements Representation and Specification” [2512.13444]. Another **DARS** designates “Dysarthria-Aware Rhythm-Style Synthesis for ASR Enhancement” [2603.01369]. **DAR** refers to “Direction-Aware Diagonal Autoregressive Image Generation” [2503.11129], and **DarSwin** denotes the “Distortion Aware Radial Swin Transformer” [2304.09691]. In the time-series UDA literature, however, **DARSD** specifically denotes **Domain Adaptation via Representation Space Decomposition** [2507.20968].

Source: https://www.emergentmind.com/topics/darsd