SilentWear: Wearable Silent Speech Interfaces
- SilentWear is a set of wearable interfaces that decode speech-related biosignals from textile sensors and EMG, enabling silent and covert communications.
- These systems employ deep learning and on-device inference to achieve high accuracy and low latency for tasks ranging from command control to assistive communication for dysarthric patients.
- They exemplify the shift from bulky laboratory setups to commodity-grade wearables, integrating privacy-preserving techniques like encrypted gesture recognition.
Searching6 arXiv6^ for the referenced SilentWear-related papers to ground the article in the current literature. {"6query6 "6\6 arXiv6", "6max_results6 6\6query6} {"6query6 SilentWear denotes a set of wearable communication interfaces that seek to bypass, minimize, or protect the conventional acoustic speech channel by decoding speech-related biosignals or covert gestures from body-worn sensors. In the current literature, the name is used most directly for textile neck interfaces for EMG-based silent speech recognition, for a strain-sensor “intelligent throat” that reconstructs fluent speech in dysarthric stroke patients, and for a privacy-preserving smartwatch gesture system that performs encrypted recognition without exposing raw signals or predictions to third parties (&&&6query6&&&, &&&6\6&&&, &&&6 arXiv6&&&). Within the broader silent speech interface (SSI) literature, these systems exemplify the transition from bulky laboratory instrumentation toward “invisible interfaces” integrated into commodity-grade wearables, with increasing reliance on deep learning and, in some cases, LLMs to compensate for sparse and non-stationary biosignals (&&&6max_results6&&&).
6\6. Terminological scope and position within silent speech interfaces
SilentWear belongs to the broader class of SSIs, defined in the review literature as interfaces that decode speech-related intent from biosignals instead of acoustic speech. That literature organizes sensing around four physiological interception points: neural oscillations, neuromuscular activation, articulatory kinematics, and active probing via acoustic or radio-frequency sensing (&&&6max_results6&&&). SilentWear implementations in the narrow sense are concentrated mainly in the neuromuscular and articulatory layers: EMG neckbands, textile strain chokers, and throat-worn hybrids that infer intended speech before or without audible phonation (&&&6query6&&&, Tang et al., 2023, &&&6\6&&&).
The name is not used uniformly across papers. In one line of work, SilentWear is a fully wearable, textile-based neck interface for EMG acquisition and on-device command recognition (&&&6query6&&&). In another, the same name is effectively associated with an “intelligent throat” that combines throat muscle vibration sensing, carotid pulse sensing, and LLM-based reconstruction for stroke patients with dysarthria (&&&6\6&&&). A different paper applies the name to a privacy-preserving covert communication pipeline based on encrypted smartwatch gesture recognition rather than speech decoding (&&&6 arXiv6&&&). This suggests that “SilentWear” functions less as a single standardized architecture than as a recurring design motif: wearable, discreet, privacy-oriented communication without reliance on ordinary audible speech.
The review literature also places such systems within a larger field-wide transition. It states that SSIs are moving from tethered or invasive setups toward earables, smart glasses, headphones, masks, textiles, and neckbands, and that LLMs and deep generative models are increasingly used as high-level linguistic priors to resolve the “informational sparsity” of biosignals (&&&6max_results6&&&). SilentWear systems exemplify that shift in concrete engineering terms: textile sensing, dry electrodes, low-power edge hardware, and language-level correction.
6 arXiv6. Wearable architectures and sensing substrates
The direct SilentWear implementations differ substantially in sensing modality, form factor, and intended use.
| System | Form factor and sensing | Reported headline result |
|---|---|---|
| SilentWear choker | Single-channel graphene textile strain sensor in a choker | 96\6.6 arXiv6\6% accuracy on 6 arXiv6query6^ words; 6query6.6query6 G FLOPS (Tang et al., 2023) |
| Intelligent throat | Smart choker with throat-vibration and carotid-pulse channels | 6query6.6 arXiv6% WER; 6 arXiv6.9% SER; 6\6\6% satisfaction increase (&&&6\6&&&) |
| SilentWear neckband | 6\6query6-channel dry textile EMG neckband with BioGAP-Ultra | 86query6.8±6query6 vocalized; 77.6\6±6.6% silent; 6 arXiv6.6query67 ms on-device latency (&&&6query6&&&) |
| SilentWear covert communication | Smartwatch inertial sensing with encrypted gesture inference | over 96query6.6query6query6 plaintext NN accuracy; 96 arXiv6.6\69% HNN accuracy (&&&6 arXiv6&&&) |
The 6 arXiv6query6 arXiv6max_results6^ textile choker work emphasizes sensor physics. Its substrate is a textile composed of 96\6% bamboo fibers and 6\6% elastane, with a screen-printed graphene layer prestretched to 6\6% strain to induce ordered cracks aligned with the textile matrix. The reported sensing metrics include a gauge factor of 6max_results6\67 within 6\6% strain, a detection limit of 6query6.6query6\6 strain, stability through more than 6\6query6,6query6query6query6^ stretch–release cycles, and complete insensitivity to introduced 6\6query6query6^ dB acoustic noise (Tang et al., 2023). The central engineering claim is that ultrasensitive mechanical sensing can reduce downstream algorithmic complexity.
The “intelligent throat” system extends the choker concept into a dual-channel architecture. One textile strain-sensing channel is aligned with the center of the throat to capture extrinsic laryngeal muscle vibrations; the second is aligned with the carotid artery to capture pulse-related physiological signals. The sensing element is a screen-printed graphene strain sensor on elastic knitted textile, and a polyurethane acrylate strain-isolation layer surrounds each channel to reduce crosstalk and suppress wear-induced strain artifacts. The reported hardware characteristics include a response above 6\6query6% to subtle strains of 6query6.6\6 a gauge factor above 6\6query6query6^ under high-frequency stretching, a wireless PCB containing ADC, MCU, Bluetooth, op-amp conditioning, and reference-voltage circuitry, total power consumption of 76.6\6^ mW, and all-day operation from a 6\686query6query6^ mWh battery (&&&6\6&&&).
The 6 arXiv6query6 arXiv66^ SilentWear neckband shifts from strain to EMG and from cloud-side decoding to embedded edge inference. It uses a soft-fabric neckband with fully dry Datwyler SoftPulse electrodes connected through 6 arXiv67 sewn-in snap fasteners. The arrangement yields 6\6query6^ differential EMG channels: 6\6query6^ in the central overlapping differential array, 6query6^ lateral channels, and 6query6^ electrically shorted electrodes at the back for ground. The acquisition and processing platform is BioGAP-Ultra, integrating GAP9, Nordic nRF6\6max_results6query6query6^ with BLE, and two ADS6\6 arXiv698 analog front ends. The hardware dimensions are reported as PRESERVED_PLACEHOLDER_6query6, sampling is at 6\6query6query6^ Hz with PGA gain 6, and the total system power for acquisition, inference, and wireless result transmission is 6 arXiv6query6.6\6^ mW, enabling 6 arXiv67.6\6^ h of operation from a 6\6\6query6^ mAh Li-Po battery (&&&6query6&&&).
The covert-communication SilentWear is architecturally different. It uses a commodity Fossil Gen 6 smartwatch worn on the dominant wrist, with tri-axial gyroscope and tri-axial accelerometer sampled at 66query6^ Hz. Rather than decoding speech, it encodes messages as a gesture alphabet PRESERVED_PLACEHOLDER_6\6^ and performs classification directly over encrypted 96-dimensional motion features using CrypTen-based homomorphic and multi-party computation (&&&6 arXiv6&&&). Its inclusion under the same name underscores the breadth of the term’s usage.
Adjacent wearable systems help clarify what is distinctive about SilentWear. A headphone-integrated SSI embeds four graphene/PEDOT:PSS-coated towel-based textile EMG electrodes in earmuffs and streams 6query6-channel EMG via an ESP6max_results6 arXiv6-S6max_results6^ module at 6\6^ kHz, emphasizing discretion and adaptive robustness to skin-electrode coupling (&&&6 arXiv6query6&&&). SottoVoce uses a 6max_results6.6\6 convex ultrasound probe under the jaw to image internal oral motion and synthesize audio for existing smart speakers (&&&6 arXiv6\6&&&). NasoVoce mounts a microphone and vibration sensor at the nasal pads of smart glasses for low-audibility and whispered speech capture (&&&6 arXiv6 arXiv6&&&). These neighboring systems show that SilentWear is part of a broader migration toward socially acceptable, near-invisible wearables rather than a single hardware lineage.
6max_results6. Decoding pipelines, feature representations, and language reconstruction
A central distinction among SilentWear systems lies in how they represent time and context. Earlier wearable SSI designs often operated on fixed windows, but the “intelligent throat” paper identifies this as a major cause of fragmented interaction. It reports that traditional wearable silent-speech systems usually require fixed time windows of 6\6–6max_results6^ seconds, producing a “speak, stop, wait” rhythm. Its solution is token-level segmentation into approximately 6\6query6query6^ ms units, with each token labeled by the word to which it belongs and classified continuously rather than as isolated command windows. To restore temporal context without heavy recurrent or transformer models, the paper augments each current token with the previous PRESERVED_PLACEHOLDER_6 arXiv6^ tokens, using blanks for early padding, and reports an optimal context length of PRESERVED_PLACEHOLDER_6max_results6^ (&&&6\6&&&).
That token pipeline is coupled to a two-agent LLM layer. The Token Synthesis Agent (TSA), based on GPT-6query6o-mini, maps token labels into words and sentences. The Sentence Expansion Agent (SEA), also GPT-6query6o-mini-based, takes the TSA output together with emotion labels and objective context such as time and weather, then expands the utterance into a more coherent, personalized, emotionally appropriate sentence. The paper reports that TSA performance improved as prompt length increased up to about 6query6query6query6^ words and then degraded, and that including example label-to-word mappings and empirical token-count constraints improved decoding (&&&6\6&&&). In effect, linguistic priors are used not merely for post-processing but for semantic and affective restoration.
The 6 arXiv6query6 arXiv6max_results6^ textile choker adopts a markedly different principle: sensor quality substitutes for model size. Because the single-channel strain sensor is described as producing high-density one-dimensional signals, the system forgoes 6 arXiv6D transforms and heavy feature engineering. It uses an end-to-end 6\6D CNN with residual blocks over raw 6max_results6-second, 6\6\6query6query6-point waveforms sampled at 6\6query6query6^ Hz. The reported architecture includes a Conv6\6d layer with 66query6^ filters of size 7, later Conv6\6d stages at 6\6 arXiv68 and 6 arXiv6\66^ channels, residual blocks with paired kernel-size-6max_results6^ convolutions, AdaptiveAvgPool6\6d, and a linear layer to 6 arXiv6query6^ classes, for a total of 6query6\68,86max_results66^ parameters (Tang et al., 2023). Instead of online filtering, the training procedure uses random noise window injection: background noise collected while the user wears the choker silently is overlaid onto speech samples to create augmented examples.
The 6 arXiv6query6 arXiv66^ EMG SilentWear neckband emphasizes lightweight embedded inference. Its SpeechNet architecture, inspired by EpiDeNet, has 6\6\6,6query6 parameters and learns early temporal patterns per channel followed by later cross-channel spatial representations. The network comprises a sequence of Conv6 arXiv6D and MaxPool stages, then AdaptiveAvgPool and a 9-class dense output corresponding to 8 commands plus rest. Training uses Cross-Entropy, Adam, an initial learning rate of PRESERVED_PLACEHOLDER_6query6, Reduction on Plateau, and early stopping. The model was initially trained with 6\6query6query6query6^ ms EMG input windows, and the paper additionally studies 6query6query6query6–6\6query6query6query6^ ms windows in 6 arXiv6query6query6^ ms increments to characterize the latency–throughput–accuracy trade-off (&&&6query6&&&).
The headphone-based neighboring system reveals a parallel strategy for dealing with wearable instability. Its 6\6D SE-ResNet introduces squeeze-and-excitation blocks that dynamically reweight the four EMG channels according to coupling quality, suppressing noisy or weakly coupled channels. Inputs are segmented into 6max_results6-second windows of shape PRESERVED_PLACEHOLDER_6\6, bandpass filtered from 6 arXiv6query6^ to 6query6\6query6^ Hz with a 6query6th-order Butterworth filter, and augmented by time shift, Gaussian noise injection, and scale/offset perturbations (&&&6 arXiv6query6&&&). This is a decoder-level answer to the same problem that SilentWear neckbands confront at the hardware level: variable contact under everyday use.
The covert-communication SilentWear again diverges. After pause-based segmentation, it extracts 6max_results6 arXiv6^ features per gyroscope axis, yielding a 96-dimensional feature vector spanning temporal and spectral domains. Classification is performed by a 6max_results6-layer fully connected network , , , with Leaky ReLU and mean-squared error loss adapted to encrypted computation. Softmax is deferred until after decryption because it is non-polynomial (&&&6 arXiv6&&&). Here the modeling constraint is not biosignal ambiguity but secure arithmetic over ciphertexts.
6query6. Empirical performance and evaluation protocols
Reported performance varies strongly with task formulation. Word-level or small-vocabulary command recognition remains the most stable regime. The 6 arXiv6query6 arXiv6max_results6^ textile choker reports 96\6.6 arXiv6\6% accuracy on a 6 arXiv6query6-word lexicon, 96max_results6% on 6\6query6^ confusable words, and 96% on 6\6^ long words spoken at different speeds, while reducing computational load by 96query6% and operating at 6query6.6query6 G FLOPS per inference (Tang et al., 2023). Transfer experiments further report 86query6% accuracy for new users and 86query6% accuracy for new words with only 6\6\6–6 arXiv6query6^ samples per class, rising to 96query6% for both with 6max_results6query6^ samples per class (Tang et al., 2023).
The headphone-integrated EMG system reports 96% classification accuracy on 6\6query6^ commonly used voice-free control words from 6query6^ subjects, outperforming 6\6D ResNet, 6\6D VGG, SVM, Random Forest, MLP, and XGBoost. Its ablations are diagnostically important: removing bandpass filtering drops 6\6D SE-ResNet accuracy from 96% to 76.6\6%, and single-channel input reduces accuracy by nearly 6query6query6% (&&&6 arXiv6query6&&&). These results indicate that both low-frequency artifact suppression and channel redundancy are central in dry-electrode wearables.
The 6 arXiv6query6 arXiv66^ SilentWear neckband provides one of the most complete multi-day evaluations. Across four subjects and three sessions per subject, SpeechNet reaches 86query6.8±6query6 average accuracy for vocalized speech and 77.6\6±6.6% for silent speech in the global leave-one-batch-out setting. Under leave-one-session-out evaluation, which includes unseen day and neckband repositioning, performance drops to 76\6.6\6 and 6\69.6max_results6 arXiv6.6 arXiv6% respectively (&&&6query6&&&). The same paper reports window-size ablations showing the highest average vocalized accuracy at 6\6query6query6query6^ ms and the highest average silent accuracy at 6\6 arXiv6query6query6^ ms, but maximum information transfer rate at 86query6query6^ ms for both vocalized and silent speech, which the authors interpret as a practical trade-off (&&&6query6&&&).
A related fully dry EMG neckband paper reinforces the same robustness issue. Using 6\6query6^ fully differential channels and Random Forest classification, it reports 87±6max_results6% average accuracy for vocalized speech and 68±6max_results6% for silent articulation under 6\6-fold cross-validation, but only 66query6±6\6 and 6\6query6±7% under leave-one-session-out evaluation after repositioning (&&&6max_results6query6&&&). The qualitative conclusion is consistent across both neckband studies: session-to-session placement remains a major source of distribution shift.
The “intelligent throat” moves beyond command recognition toward sentence reconstruction and clinical utility. In tests with five stroke patients with dysarthria, after pretraining on healthy subjects and few-shot fine-tuning on patient data, token classification accuracy reached 96 arXiv6.6 arXiv6% after only 6 arXiv6\6^ repetitions per word, compared with 79.8% when training only on patient data. Response-based knowledge distillation from a 6\6D ResNet-6\6query6\6^ teacher to a 6\6D ResNet-6\68 student reduced computational load by 76\6.6% while retaining 96\6.6max_results6 accuracy, 6query6.9% below the teacher. At the language level, optimal prompting of the TSA yielded 6query6.6 arXiv6% word error rate and 6 arXiv6.9% sentence error rate, and the SEA increased user satisfaction by 6\6\6%, from “somewhat satisfied” to “fully satisfied” (&&&6\6&&&). The paper also reports an end-to-end delay from silent expression completion to speech playback of about 6\6^ second.
The covert-communication SilentWear reports best plaintext neural-network accuracy of 96query6.6query6query6 on Jetson Orin Nano and best homomorphic neural-network accuracy of 96 arXiv6.6\69% on RTX 6query6query696query6^ and Jetson Orin Nano. Weighted F6\6^ scores are reported as 6query6.96query6max_results6query6^ for the plaintext model and 6query6.96 arXiv6\6\6^ for the encrypted model on RTX 6query6query696query6 with micro- and macro-AUC values consistently above 6query6.96 (&&&6 arXiv6&&&). The latency overhead of encryption is substantial—6.6\6query66^ ms HNN latency on RTX 6query6query696query6 6query6\6.988 ms on Jetson Orin Nano, and 6\6max_results6 arXiv6.6query66query6^ ms on Jetson Nano 6 arXiv6GB—but the paper presents this as still practical on Orin-class edge devices (&&&6 arXiv6&&&).
6\6. Application domains and relation to neighboring wearable systems
SilentWear systems are used in at least three distinct application domains. The first is assistive communication. The “intelligent throat” is explicitly designed for stroke patients with dysarthria, using throat muscle vibration sensing and carotid pulse-based affect recognition to reconstruct fluent, emotionally expressive communication. Its speech corpus consists of 6query67 Chinese words commonly used in daily communication and 6 arXiv6query6^ sentences built from those words, and the authors state that the platform has potential for application across different neurological conditions and in multi-language support systems (&&&6\6&&&).
The second domain is human–machine interaction via command vocabularies. The 6 arXiv6query6 arXiv66^ EMG SilentWear neckband uses eight commands—up, down, left, right, forward, backward, start, and stop—together with a rest class, targeting representative HMI tasks. The 6 arXiv6query6 arXiv6\6^ headphone SSI likewise recognizes 6\6query6^ control words including Open, Close, Start, Stop, Yes, No, Next, Back, OK, and Cancel, and explicitly points to assistive communication, smart-device control, wearable human-computer interaction, and embodied AI or robotics scenarios such as teleoperation or exoskeleton control (&&&6query6&&&, &&&6 arXiv6query6&&&). These studies define “silent speech” pragmatically, as reliable biosignal-based command entry under wearable constraints.
The third domain is covert and privacy-preserving communication. In the smartwatch-based SilentWear, the central privacy claim is that no raw sensor signals, learned features, or classification outputs are exposed to any third party. Messages are conveyed through encrypted gesture recognition and delivered through haptic or low-salience visual feedback. The selected gesture set maps to semantic roles such as alert, request/action, acknowledge/confirm, and emergency/abort, and the finite-state communication process uses opening and closing pauses as implicit delimiters (&&&6 arXiv6&&&). Although not a speech decoder, it shares SilentWear’s recurring themes of wearability, discretion, and local control over sensitive signals.
Neighboring systems clarify the edge of the category. SpeechLess is not fully silent; rather, it provides speech-granularity control in wearable AR, allowing Full Utterance, Partial Utterance, or Zero Utterance based on personalized spatial memory. In a controlled study, it reports accuracy of 96\6.6query6 for Full, 86.7% for Partial, and 86max_results6.6max_results6 for Zero, with Partial reducing spoken word count by about 6query69.8% (&&&6query6 arXiv6&&&). NasoVoce is likewise not fully silent but supports whispered and low-volume speech using nose-bridge-mounted microphone and vibration sensors on smart glasses, with a dual-input D-DCCRN enhancement model and Whisper Large-v6 arXiv6-based evaluation (&&&6 arXiv6 arXiv6&&&). SottoVoce uses under-jaw ultrasound and a two-network pipeline to synthesize speech audio that can control unchanged Amazon Echo devices, reporting 66\6.6query6 average recognition success with both networks and 6max_results6max_results6.6\66 WER for Network 6 arXiv6^ on Google speech-to-text (&&&6 arXiv6\6&&&). LipLearner offers customizable mobile lipreading with one-shot adaptation, on-device fine-tuning, and visual keyword spotting, achieving PRESERVED_PLACEHOLDER_6\6query6^ for 6 arXiv6\6-command classification with one shot (&&&6query6\6&&&). These systems are not all called SilentWear, but they map the broader ecosystem in which the term operates.
6. Limitations, misconceptions, and future directions
A common misconception is that SilentWear already denotes a mature, standardized everyday speech replacement. The literature does not support that reading. The strongest clinical study involves only five stroke patients with dysarthria (&&&6\6&&&). The 6 arXiv6query6 arXiv66^ EMG SilentWear dataset includes four subjects across three sessions (&&&6query6&&&). A related fully dry neckband study is a single-subject evaluation (&&&6max_results6query6&&&). The headphone-integrated EMG system uses four subjects (&&&6 arXiv6query6&&&). These are substantial engineering demonstrations, but not large-cohort deployment studies.
Another misconception is that wearability alone resolves inter-session robustness. Both EMG neckband papers show measurable performance degradation after removing and repositioning the device between sessions. The 6 arXiv6query6 arXiv66^ SilentWear paper explicitly interprets the global-to-inter-session accuracy gap as evidence that multi-day use induces a distribution shift in EMG space (&&&6query6&&&). It proposes incremental fine-tuning as mitigation and reports more than 6\6query6% accuracy recovery with less than 6\6query6^ minutes of additional user data, with even one fine-tuning round producing large gains (&&&6query6&&&). This suggests that practical SilentWear systems may require lightweight continual adaptation rather than fixed once-trained models.
The role of LLMs also warrants a precise reading. In the review literature, LLMs are presented as part of a broader shift toward “Latent Semantic Alignment,” where fragmented physiological evidence is mapped into structured semantic latent spaces (&&&6max_results6&&&). In the “intelligent throat,” the LLM agents do not merely format output; they are explicitly used to correct token errors, restore logical coherence, and enrich emotional appropriateness (&&&6\6&&&). A plausible implication is that future SilentWear systems may increasingly separate low-level biosignal decoding from high-level semantic reconstruction.
The literature also raises unresolved ethical and privacy questions. The SSI review introduces “neuro-security” and the protection of cognitive liberty as emerging design constraints for increasingly invisible interfaces (&&&6max_results6&&&). The encrypted gesture-based SilentWear addresses this directly by ensuring that raw motion, learned features, intermediate representations, and predictions remain encrypted during computation (&&&6 arXiv6&&&). By contrast, silent speech systems that rely on cloud-side LLMs or external inference services would have to solve privacy in different ways; the reviewed papers do not provide a unified answer.
Several concrete future directions recur across the papers. The “intelligent throat” identifies larger clinical validation, multilingual support, broader neurological conditions, improved demographic diversity, and miniaturization into an edge-computing architecture as next steps (&&&6\6&&&). The EMG neckband studies point to greater robustness to placement variation, larger and more diverse cohorts, and expanded vocabularies (&&&6query6&&&, &&&6max_results6query6&&&). The SSI review frames the broader agenda as self-supervised foundation models, on-device continual learning, multimodal fusion, and low-latency edge deployment, while also stating that end-to-end delay should remain below 6\6query6^ ms for practical closed-loop use and that a WER below 6\6\6% is generally deemed essential for functional parity with traditional ASR (&&&6max_results6&&&).
Taken together, the literature presents SilentWear not as a single finished product but as an emerging class of wearable, privacy-oriented, non-acoustic communication systems. Its unifying features are textile or commodity-grade wearability, direct interception of speech-related or intent-related body signals, and increasingly sophisticated inference stacks that range from compact CNNs on microcontrollers to encrypted neural networks and LLM-based semantic restoration.