Papers
Topics
Authors
Recent
Search
2000 character limit reached

FM-FoG: Real-Time FoG Wearable System

Updated 14 July 2026
  • The paper introduces FM-FoG, a real-time wearable system that uses a self-supervised transformer and sensor context integration to detect FoG without patient-specific training.
  • It employs a two-stage pipeline with a CNN-LSTM activity classifier gating a transformer-based FoG detector, achieving a 98.5% F1-score with sub-20 ms latency.
  • The system extends battery life by up to 72% through selective activation, offering practical improvements for fall risk mitigation in Parkinson’s patients.

FM-FoG is a real-time foundation model-based wearable system for Freezing-of-Gait mitigation in Parkinson’s disease. It targets Freezing-of-Gait (FoG), a debilitating motor symptom characterized by sudden episodes where walking cannot start or is interrupted despite the intention to move forward; it occurs exclusively during standing or walking, and never while sitting or lying down. The system is designed to detect FoG in previously unseen patients without patient-specific training by combining self-supervised pretraining on diverse Inertial Measurement Unit datasets, sensor context integration, and event-triggered inference on a smartphone. On the VCU FoG-IMU dataset with 23 Parkinson’s disease patients, FM-FoG achieves a 98.5% F1-score on unseen patients, operates with approximately 19.9 ms end-to-end intervention latency, and extends battery life by up to 72% through selective activation on a Google Pixel 8a (Chi et al., 29 Sep 2025).

1. Clinical scope and problem formulation

FoG affects over 50% of individuals with mid-to-late stage Parkinson’s disease and is associated with increased fall risk, hospitalizations, and a profound loss of movement independence, reducing quality of life for millions of patients worldwide. The symptom is episodic and activity-contingent: it appears during ambulatory states, particularly standing and walking, and not during sitting or lying down. That constraint is central to FM-FoG’s design, because it allows the system to avoid unnecessary inference outside FoG-relevant contexts (Chi et al., 29 Sep 2025).

The system is positioned against several limitations of existing FoG detection technologies. The reported shortcomings include dependence on patient-specific training, poor cross-subject generalization arising from high inter-subject variability in FoG manifestations, laborious manual labeling with possible expert disagreement, reliance on handcrafted features that do not transfer well across sensors and placements, and continuous inference regimes that drain battery and undermine practical mobile deployment. FM-FoG addresses these issues by using a foundation model tailored to IMU time-series, explicit sensor context in the model input, and a selective activation mechanism that only invokes FoG inference when the activity state is FoG-relevant. This suggests that the system treats generalization, labeling burden, and device energy as coupled design constraints rather than independent optimization targets.

2. Wearable pipeline and intervention loop

FM-FoG implements a two-stage pipeline. First, an activity trigger model continuously classifies activity in 1.28 s windows. Second, when walking or standing are detected, the FoG foundation model is executed; if FoG is detected, a Bluetooth haptic intervention is delivered through the PDVibe3 ankle-mounted device. The haptic pattern is rhythmic pulsed vibration with 2 seconds on and 1 second off. The entire pipeline runs on a Google Pixel 8a smartphone and communicates over Bluetooth Low Energy, with automatic reconnection handling (Chi et al., 29 Sep 2025).

The activity trigger exists because FoG occurs only during ambulatory states. FM-FoG therefore does not maintain always-on execution of the heavier model. Instead, the activity classifier gates the foundation model, activating it only for standing, walking, and “attempted walking,” while treating sitting and lying as FoG-irrelevant. The paper does not report explicit thresholds or hysteresis, and gating is event-triggered based on classifier outputs.

The deployment strategy also reflects a latency-aware systems choice: both the activity classifier and the FoG model are loaded permanently in memory to avoid latency from model swaps and I/O. The complete system uses a compact 1.2M-parameter backbone, with approximately 16% CPU usage and approximately 112 MB memory usage on the Pixel 8a. The reported end-to-end intervention latency is sub-20 ms, specifically approximately 19.9 ms for the triggered configuration. Such timing aligns with clinical evidence that prompt cue timing benefits gait initiation in Parkinson’s disease with FoG.

3. Sensing, datasets, and preprocessing

FM-FoG uses ankle-mounted UltiGesture devices capturing 3-axis accelerometer and 3-axis gyroscope data at 100 Hz. The ankle placement is justified in the paper by the presence of pronounced gait irregularities associated with FoG and by the fact that the device is non-intrusive under clothing. The evaluation dataset, VCU FoG-IMU, contains 23 Parkinson’s disease patients aged 57–82 years, with mean age 69.9±5.769.9 \pm 5.7, including 16 male and 7 female participants. Recording sessions were designed around five FoG-provoking conditions—dual-tasking, tight turns, narrow passages, visual targets, and time-pressured walking—plus 10 m straight-line walks (Chi et al., 29 Sep 2025).

The dataset was labeled using synchronized video annotations by movement disorder specialists. Signals were windowed into 1.28 s segments with 50% overlap. A window was labeled FoG if freezing activity occurred in the final 30% of the segment; otherwise it was labeled non-FoG. The dataset comprises 21,868 labeled windows, approximately 30% of which are FoG, spanning 528 FoG episodes across approximately 238 minutes.

FM-FoG’s self-supervised pretraining corpus spans multiple placements, modalities, and sampling rates. All datasets were harmonized to 100 Hz before pretraining.

Dataset Sensor setup Cohort / labels
tDCS FOG 128 Hz, lower back, accelerometer 67 PD patients; FoG start hesitation, turning, walking, non-FoG
DeFOG 100 Hz, lower back, accelerometer 66 PD patients; FoG vs. non-FoG
Daphnet 64 Hz, ankle, thigh, trunk, accelerometer 10 PD patients; FoG vs. non-FoG
PAMAP2 100 Hz, wrist/chest/ankle, accelerometer, gyroscope, magnetometer 9 healthy subjects; 12+ activities
VCU FoG-IMU 100 Hz, ankle, accelerometer and gyroscope 23 PD patients; FoG vs. non-FoG

Preprocessing standardizes both temporal and spatial heterogeneity. Sampling harmonization uses cubic spline interpolation for upsampling from 64 Hz and random sub-sampling for downsampling from 128 Hz. An ablation reports that this harmonization improves F1 from 91.2% to 98.5%, a gain of 7.3 percentage points. Sensor placement is encoded through learnable location embeddings, while coordinate transformations standardize axis orientation across devices and mounts. Multi-modal integration uses ordered channels for accelerometer, gyroscope, and magnetometer; missing modalities are zero-padded, with modality-specific normalization. The normalization stack includes per-axis, per-subject z-score normalization and foot-wise normalization, explicitly to preserve lateralized gait asymmetries common in Parkinson’s disease.

4. Context-aware foundation model and gated activity classifier

The FoG detector is an encoder-only Transformer tailored to IMU sequences. Its input is a 1.28 s sequence comprising 128 time steps at 100 Hz. Sensor channels are linearly projected into the model dimension, combined with positional encoding for temporal order, and fused with learnable location embeddings that represent placement such as ankle, thigh, or lower back. The Transformer contains 6 blocks with multi-head self-attention using 8 heads, model dimension dmodel=128d_{\text{model}} = 128, layer normalization, residual connections, and ReLU feed-forward networks. During self-supervised pretraining, the model reconstructs masked sensor values after randomly masking 30% of time positions; the objective is mean squared error on masked positions. Pretraining uses AdamW with learning rate 5×10−45 \times 10^{-4}, weight decay 0.01, cosine annealing, mixed precision, batch size 32, and 50 epochs over the combined corpus (Chi et al., 29 Sep 2025).

Sensor context integration is an architectural distinguishing feature. Each placement ID is mapped to a learned vector through an embedding layer, and the location embedding is fused with the IMU token representation at the input stage alongside positional encoding. The paper frames this as a response to spatial heterogeneity across datasets and deployments. A plausible implication is that the model is not merely learning gait dynamics, but learning them conditionally on sensor location.

Fine-tuning adds a FoG classification head to the pretrained backbone. The head consists of global average pooling across time, a linear classifier, and softmax for FoG versus non-FoG. Fine-tuning is end-to-end, with learning rate 1×10−41 \times 10^{-4} for the transformer layers and 1×10−31 \times 10^{-3} for the classification head. Temporal jittering and Gaussian noise injection are used as augmentations.

Selective activation is provided by a separate CNN-LSTM activity classifier of approximately 5 MB. It contains two 1D convolutional layers with 32 and 64 filters and kernel size 12, followed by ReLU activations, max pooling, and 50% dropout, then two LSTM layers with 64 hidden units each, and finally a fully connected layer outputting binary probabilities. Its input window is the same 1.28 s sequence used by the foundation model. Its output classes are FoG-relevant versus FoG-irrelevant activity, where the former includes standing, walking, and attempted walking.

5. Evaluation, generalization, and on-device efficiency

Evaluation on VCU FoG-IMU uses 16 training patients and 7 unseen test patients, repeated across 25 random partitions to estimate mean and variance. The unseen-patient setting ensures that test subjects are never used during fine-tuning. Performance is reported in terms of precision, recall, accuracy, and F1, with

P=TPTP+FP,R=TPTP+FN,F1=2PRP+R.P = \frac{\mathrm{TP}}{\mathrm{TP} + \mathrm{FP}}, \qquad R = \frac{\mathrm{TP}}{\mathrm{TP} + \mathrm{FN}}, \qquad \mathrm{F1} = \frac{2PR}{P+R}.

Under this protocol, FM-FoG substantially outperforms FoG-specific ML baselines, time-series foundation models, and general-purpose LLM baselines (Chi et al., 29 Sep 2025).

Method F1-score Notes
FM-FoG 98.5±0.7%98.5 \pm 0.7\% 1.2M parameters
Gait-Guard 90.9±1.3%90.9 \pm 1.3\% FoG-specific baseline
Dual-Level FoG recognition 71.9±7.5%71.9 \pm 7.5\% FoG-specific baseline
MOMENT 85.9±3.8%85.9 \pm 3.8\% Time-series foundation model
TimesFM-200M dmodel=128d_{\text{model}} = 1280 Time-series foundation model
TimesFM-500M dmodel=128d_{\text{model}} = 1281 Time-series foundation model
LLaMA2-7b dmodel=128d_{\text{model}} = 1282 Inference only
GPT-4o dmodel=128d_{\text{model}} = 1283 Inference only

The ablation studies identify two major contributors beyond the backbone itself. First, sampling harmonization improves F1 by dmodel=128d_{\text{model}} = 1284, from 91.2% to 98.5%. Second, sensor context integration improves F1, accuracy, precision, and recall across training sizes from 1 to 16 patients, while also reducing variance; gains are reported as most pronounced for small training cohorts, with approximately 5–8% F1 increase when fewer than 8 training subjects are available. This suggests that context conditioning is especially beneficial in low-data regimes, where uncontrolled placement variability would otherwise dominate.

Real-time deployment metrics indicate that triggering adds little latency while improving energy behavior. Reported inference and intervention times are 14.3 ms and 17.5 ms for FM-FoG in continuous mode, and 16.7 ms and 19.9 ms for triggered FM-FoG. CPU usage is approximately 16.1% for FM-FoG versus approximately 18.1% for Gait-Guard, and memory is approximately 107–112 MB. Energy was measured through ADB-over-WiFi power sampling every 100 ms using

dmodel=128d_{\text{model}} = 1285

with cumulative energy obtained by summing dmodel=128d_{\text{model}} = 1286. The paper notes that WiFi introduces measurement overhead, so actual model power would be lower without WiFi.

Operating mode Power Battery life
Idle phone dmodel=128d_{\text{model}} = 1287 W dmodel=128d_{\text{model}} = 1288 h
Continuous FM-FoG dmodel=128d_{\text{model}} = 1289 W 5×10−45 \times 10^{-4}0 h
Triggered FM-FoG, 10% active 1.5 W 11.5 h
Triggered FM-FoG, 30% active 1.8 W 9.6 h
Triggered FM-FoG, 40% active 2.1 W 8.2 h

At a 30% triggering rate, device lifetime improves from 6.7 h to 9.6 h, a 43% increase. At a 10% triggering rate, it improves to 11.5 h, approximately 72% longer than continuous monitoring. At 80–90% activity, power rises to approximately 2.9 W and battery life falls to approximately 6 h, indicating that the energy advantage is strongest when ambulatory activity is intermittent rather than continuous.

6. Robustness, limitations, and nomenclature

FM-FoG’s reported robustness rests on cross-subject transfer without patient-specific training. The paper attributes this to self-supervised pretraining across multiple IMU datasets, harmonized sampling, and sensor context embeddings that explicitly model placement and axis heterogeneity. The VCU FoG-IMU cohort also includes diverse motor comorbidities such as dyskinesia, dystonia, and ataxia, requiring discrimination of FoG from other abnormal movements. The system’s context-aware design and domain-specific pretraining are described as mechanisms that help preserve specificity under this heterogeneity (Chi et al., 29 Sep 2025).

Several limitations are also explicit. The transformer backbone operates as a black box, and the paper identifies attention visualization and explainable AI as future directions for improving clinical trust. The current dataset is based on controlled sessions rather than longitudinal, multi-institutional daily monitoring. The activity classifier’s gating logic does not report explicit hysteresis or thresholds. Validation is limited to the Pixel 8a, so broader device testing remains warranted. The paper also does not provide code or random seeds, although evaluation is standardized and repeated across 25 splits to estimate variance.

The term “FM-FoG” is not unique across the FoG literature. In “Freezing of Gait Detection Using Gramian Angular Fields and Federated Learning from Wearable Sensors” (Soumma et al., 2024), FM-FoG denotes a federated, privacy-preserving FOG detection framework instantiated through FOGSense, using Gramian Angular Field representations, a multichannel 2D CNN, and federated learning with patient-level personalization. By contrast, FM-FoG in “FM-FoG: A Real-Time Foundation Model-based Wearable System for Freezing-of-Gait Mitigation” denotes a context-aware IMU foundation model coupled to an event-triggered wearable-smartphone intervention stack. This shared acronym has created a nomenclature ambiguity, but the two systems are methodologically distinct: one centers on self-supervised foundation-model pretraining and selective activation, while the other centers on GAF-based representation learning and federated adaptation.

In clinical and practical terms, FM-FoG is presented as enabling immediate utility in new patients without weeks of per-patient data collection or fine-tuning, reducing healthcare costs and improving access where individualized calibration is infeasible. Its compact model and smartphone-based deployment are intended to support in-clinic assessment, at-home monitoring, and timely haptic cueing to help patients overcome FoG episodes and reduce fall risk. A plausible implication is that the system’s significance lies less in any single architectural component than in the integration of cross-patient generalization, real-time mobile inference, and energy-aware intervention delivery within a single wearable pipeline.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to FM-FoG.