Papers
Topics
Authors
Recent
Search
2000 character limit reached

RL-BioAug Framework for EEG Augmentation

Updated 23 January 2026
  • The paper demonstrates that RL-BioAug significantly improves EEG representation quality by employing a transformer-based RL agent to choose augmentation strategies using only 10% label guidance.
  • RL-BioAug is a framework that integrates a self-supervised encoder with context-aware augmentation, adapting to EEG non-stationarity and outperforming heuristic methods.
  • Experimental results show up to 9.7% absolute gain in Macro-F1 scores on Sleep-EDFX and CHB-MIT datasets, highlighting its effectiveness in sleep staging and seizure detection.

RL-BioAug is a label-efficient reinforcement learning (RL) framework that establishes an autonomous paradigm for data augmentation in self-supervised electroencephalography (EEG) representation learning. RL-BioAug employs an RL agent, guided by only a minimal fraction (10%) of labeled data, to determine context-appropriate augmentation strategies for each data sample. Its primary aim is to improve representation quality for downstream tasks under the inherent non-stationarity of EEG signals, outperforming heuristic or random composite augmentation techniques (Lee et al., 20 Jan 2026).

1. System Architecture

RL-BioAug integrates three major modules: a self-supervised encoder, a data-augmentation module, and a transformer-based RL agent. The encoder fθf_\theta is implemented as a 1D ResNet-18 mapping EEG segment xx to an embedding z=fθ(x)z = f_\theta(x). The data-augmentation module generates a "weak" view (xweakx_{\text{weak}}) using fixed small jitter and scaling, and a "strong" view (xstrongx_{\text{strong}}) by applying augmentations chosen by the RL agent. The RL agent πϕ\pi_\phi uses both the current embedding (state sts_t) and a history of past actions and rewards to generate a probability distribution across augmentation actions, sampling one per instance.

The training loop, executed per batch and sample, consists of:

  1. Computing sts_t using the frozen encoder fθf_\theta.
  2. Having πϕ\pi_\phi select xx0 (augmentation) given xx1.
  3. Augmenting xx2 into xx3 (using xx4) and xx5 (using fixed jitter+scaling).
  4. Updating xx6 via InfoNCE loss between the two views.
  5. Computing RL agent reward xx7 (using Soft-KNN metric and 10% labeled reference set), then updating xx8 via a policy-gradient step with entropy regularization.

2. Reinforcement-Learning Components

The RL agent operates with:

  • State space: xx9, representing the encoder's output for each EEG segment.
  • Action space: Discrete, z=fθ(x)z = f_\theta(x)0, each corresponding to a distinct "strong" augmentation:

    1. Time Masking: random interval zeroed out, length z=fθ(x)z = f_\theta(x)1.
    2. Time Permutation: segment sequence shuffled.
    3. Crop & Resize: contiguous window cropped (size z=fθ(x)z = f_\theta(x)2), resized back to length z=fθ(x)z = f_\theta(x)3.
    4. Time Flip: reverse time axis.
    5. Time Warp: random subinterval sped up or slowed down by z=fθ(x)z = f_\theta(x)4.
  • Reward: Soft-KNN consistency score calculated on a labeled reference set z=fθ(x)z = f_\theta(x)5, comparing whether the embedding z=fθ(x)z = f_\theta(x)6 of the current sample consistently shares labels with its z=fθ(x)z = f_\theta(x)7 nearest neighbors:

z=fθ(x)z = f_\theta(x)8

with z=fθ(x)z = f_\theta(x)9 denoting xweakx_{\text{weak}}0 nearest neighbors, xweakx_{\text{weak}}1 as cosine similarity, and xweakx_{\text{weak}}2.

3. Policy Network Details

The RL agent’s policy network is a transformer, receiving:

  • The current sample’s state embedding xweakx_{\text{weak}}3.
  • Embedded representations of the last xweakx_{\text{weak}}4 actions and rewards.
  • All inputs fused via fully connected layers to a shared dimension.
  • Learned positional encodings and xweakx_{\text{weak}}5 self-attention blocks model the temporal policy-reward dynamics.
  • Final logits over action space xweakx_{\text{weak}}6 produced by an output head followed by softmax; actions sampled via Top-K strategy.

Policy gradients use the REINFORCE++ update:

xweakx_{\text{weak}}7

xweakx_{\text{weak}}8

where entropy xweakx_{\text{weak}}9 fosters exploration, and xstrongx_{\text{strong}}0 is updated via gradient descent.

4. Contrastive Learning Objective

RL-BioAug employs SimCLR's InfoNCE objective:

xstrongx_{\text{strong}}1

Each sample produces two augmented views:

  • Weak augmentation: fixed jitter + scaling.
  • Strong augmentation: transformation selected by agent xstrongx_{\text{strong}}2.

5. Training Strategy and Label Utilization

Training proceeds in two phases: Phase 1 (agent pre-training):

  • 10% of data with labels forms both the reference set xstrongx_{\text{strong}}3 and the unlabeled pool.
  • RL agent xstrongx_{\text{strong}}4 learns augmentation policy using only reward signals, without propagating label supervision into the encoder.

Phase 2 (self-supervised encoder training):

  • xstrongx_{\text{strong}}5 frozen.
  • For each sample, agent suggests strongest augmentation; encoder xstrongx_{\text{strong}}6 trained using contrastive loss over all (unlabeled) data.

Pseudocode for the loop is:

xstrongx_{\text{strong}}7 This approach minimizes the need for expert intervention and labels, leveraging context-aware augmentation strategies to maximize discrimination in self-supervised spaces.

6. Experimental Validation

RL-BioAug was evaluated on Sleep-EDFX (5-class sleep staging) and CHB-MIT (binary seizure detection) datasets. Results reveal:

Dataset Random Composite (MF1) RL-BioAug (MF1) Absolute Gain
Sleep-EDFX ~59.86% 69.55% +9.69%
CHB-MIT ~62.70% 71.50% +8.80%

Analysis of learned augmentation strategies:

  • On Sleep-EDFX, Time Masking was selected with ~62% probability.
  • On CHB-MIT, Crop & Resize dominated at ~77% probability.

7. Context, Rationales, and Implications

EEG’s pronounced non-stationarity necessitates adaptive data augmentation: different neural states (e.g., REM, seizure) exhibit variable tolerance to distortions; fixed or random policies risk under- or over-distortion. RL-BioAug frames augmentation selection as a reinforcement learning problem, enabling per-sample policy adaptation that maximizes downstream representation quality (as measured by Soft-KNN consistency). Entropy regularization and Top-K sampling balance exploration and exploitation in policy selection.

Label-efficient RL guidance (only 10% of labels needed) allows the self-supervised encoder to function without direct label supervision, suggesting wider applicability to unlabeled biomedical contexts. The framework’s capability to supplant heuristic-based augmentation approaches and its demonstrated improvements in Macro-F1 scores indicate its significance for autonomous EEG data augmentation (Lee et al., 20 Jan 2026). A plausible implication is that similar RL-driven augmentation policies may benefit other non-stationary time-series domains with limited labeled data.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to RL-BioAug Framework.