---
title: 'EchoForce: Acoustic Force-Sensing Interfaces'
url: https://www.emergentmind.com/topics/echoforce
type: topic
---

# EchoForce: Acoustic Force-Sensing Interfaces

EchoForce denotes an acoustic force-sensing paradigm in which force is inferred from how a mechanical interaction perturbs an acoustic field. In its narrowest and most explicit usage, it is a wrist-worn system that emits inaudible FMCW ultrasound and estimates grip force from echoes modulated by forearm skin deformation [2507.20437]. In a broader synthesis-level sense, the same logic extends to soft tactile skins that learn force from deformation-induced transfer-function changes in embedded channels and to resonance-based taxels whose frequencies shift under load [2303.17355], [2307.09730]. Across these variants, the shared structure is an acoustic excitation or self-excited acoustic response, a deformation-dependent acoustic observable, and a mapping from that observable to force, contact location, or another mechanically meaningful state.

## 1. Conceptual scope and transduction logic

EchoForce is best understood as a family of acoustically mediated force interfaces rather than a single transducer topology. The observable may be an ultrasonic echo image, a small set of FFT amplitudes at driven tones, a resonance frequency, or an edge-triggered acoustic pulse. The force-dependent physics likewise varies: skin curvature changes multipath delay and phase in air-coupled ultrasound; soft-channel deformation changes acoustic impedance and transmission; compliant end-caps alter boundary conditions and effective cavity length; electrostatic adhesion transitions radiate transient acoustic pressure [2507.20437], [2303.17355], [2307.09730].

This diversity matters because “acoustic sensing” is sometimes treated as synonymous with time-of-flight ranging. The literature here is broader. One line uses matched-filtered FMCW echoes and differential echo images; another uses amplitude modulation of known tones at fixed frequencies; a third uses resonance shifts in audible-band pneumatic chambers. A plausible implication is that EchoForce is more accurately defined by force-conditioned acoustic state estimation than by any single waveform family.

| Instantiation | Acoustic observable | Representative reported result |
|---|---|---|
| Wristband EchoForce | Differential FMCW echo profile from 20–29 kHz chirps | Mean error rate 9.08% with fine-tuned user-dependent training; 12.29% for the user-independent foundation model [2507.20437] |
| AST Skin | FFT amplitudes at 300, 500, 700, and 900 Hz through deformable channels | More than 93% of estimates within ±1.5 N over 0–30\(^{+1}\) N; contact-location accuracy typically exceeds 96% [2303.17355] |
| AcousTac-derived EchoForce | Resonance-frequency shifts of pneumatically driven taxels | Sensitivities of approximately 3.2–19 Hz/N over approximately 0.2–15 N, depending on geometry [2307.09730] |

## 2. Wrist-worn EchoForce for continuous grip-force estimation

The EchoForce wristband is built from an off-the-shelf silicone wristband, two custom PCBs, and a 3D-printed sensing bracket carrying an OWR-0504T-16 miniature speaker and an SPH0641LU4H-1 digital microphone. The sensing board is tilted at \(45^\circ\) relative to the skin surface and faces the anterior forearm flexor region. A pilot ablation found that the \(45^\circ\) bracket outperformed \(90^\circ\) by 4.19%, while a flat \(0^\circ\) placement suffered from inadvertent skin-contact shifts. Compute is provided by a Teensy 4.0 Development Board on the wrist, with streaming to a PC for data logging and training; inference is identified as portable on-device in future work [2507.20437].

Its excitation is an inaudible FMCW chirp from \(20\) kHz to \(29\) kHz. With microphone sampling rate \(f_s = 96\) kHz and frame length \(600\) samples, the sweep rate is \(160\) frames/s. The transmitted signal is modeled as
$$
s(t) = A \sin\left(2\pi\left(f_0 t + \frac{k}{2} t^2\right)\right), \quad 0 \le t < T_{\text{frame}},
$$
with \(f_0 = 20\) kHz, \(f_1 = 29\) kHz, and \(k = (f_1-f_0)/T_{\text{frame}}\). The received signal is represented as
$$
r(t) = \sum_i \alpha_i(t)\, s(t-\tau_i(t)) + n(t),
$$
where \(\alpha_i(t)\) captures reflection-gain changes and \(\tau_i(t)\) captures time-of-flight changes across multiple paths. Grip-induced deformation of the forearm, especially around flexor digitorum superficialis and palmaris longus, changes amplitude, delay, phase, and de-chirped energy distribution. At the reported settings, the vertical resolution is approximately \(1.79\) mm/pixel.

The continuous waveform is segmented into \(600\)-sample frames and cross-correlated with the transmitted chirp to obtain a 1D range profile. Stacking profiles over time yields a 2D echo profile, and static reflections are suppressed by frame differencing,
$$
E_{\text{diff}}(:, t) = E(:, t) - E(:, t-1).
$$
Model input is a moving \(2\) s window of size \(320 \times 78\) in time-by-range coordinates, corresponding to \(320\) frames and a \(78\)-pixel spatial window covering approximately \(13.93\) cm. Ground truth is provided by a CAMRY EH101 dynamometer; pounds displayed on-screen are converted to kilograms, “hold” frames are discarded, and the resulting labels are linearly interpolated to uniform \(100\) Hz. EchoForce reports force in kilograms, with Newton conversion given by \(F[\mathrm{N}] = 9.81 \times f[\mathrm{kg}]\).

## 3. Learning formulation, evaluation protocol, and empirical results

The wristband system formulates grip-force estimation as continuous regression on differential echo images of size \(320 \times 78\). The model family is described as a user-independent “foundation model” trained on a multi-user corpus and optionally fine-tuned per user. With \(f(t)\) denoting ground-truth force and \(\hat{g}_\theta(t)\) the estimate, the training objective is mean squared error,
$$
L(\theta) = \frac{1}{N}\sum_i \left(\hat{g}_\theta(t_i)-f(t_i)\right)^2.
$$
The principal training regimes are leave-one-user-out user-independent training for \(10\) epochs, user-dependent training for \(40\) epochs, fine-tuning from the user-independent backbone, and leave-one-orientation-out cross-orientation testing [2507.20437].

The experimental corpus comprises \(N=11\) participants, ages \(19\)–\(33\) years, with \(15\) recorded sessions per participant and three wrist orientations: supinated, neutral, and pronated. Each session lasts \(2\) minutes, producing \(30\) minutes per participant and \(330\) minutes overall. At \(160\) fps, this corresponds to \(19{,}200\) frames per session, approximately \(288{,}000\) frames per participant, and approximately \(3.17\) million acoustic frames in total. Trials target \(5\%\), \(25\%\), \(50\%\), \(75\%\), and \(100\%\) MVC, with grip-and-release duration of approximately \(2\) s per repetition.

The primary reported metric is
$$
\text{Error Rate}(\% \text{MVC}) = \frac{\text{RMSE}}{\text{MVC}} \times 100.
$$
Aggregate results are: mean error rate \(9.08\%\) and RMSE \(2.31\) kg for fine-tuned user-dependent training; \(10.10\%\) and \(2.56\) kg for user-dependent training without fine-tuning; \(12.29\%\) and \(3.11\) kg for the user-independent foundation model; and \(12.94\%\) and \(3.32\) kg for cross-orientation evaluation. Supinated orientation gives the best cross-orientation accuracy, with mean error \(11.77\%\) and RMSE \(3.02\) kg. Per-participant RMSE ranges from approximately \(1.66\) to \(3.5\) kg in the user-dependent setting and approximately \(2.43\) to \(4.04\) kg in the user-independent setting. Comfort was rated \(4.63/5\), and participants reported that the ultrasound was inaudible.

The reported comparisons place EchoForce against several wearable alternatives. Keir and Mogk reported approximately \(11.4\%\) error under similar MVC protocols for sEMG-based grip estimation, whereas EchoForce reports \(9.08\%\) in fine-tuned user-dependent testing and \(12.29\%\) without user-specific calibration. Static EMG calibration in Hoozemans et al. is summarized as approximately \(6.1\) kg error. HIPPO reflectivity rises to approximately \(9.72\) kg RMSE when user-independent, capacitive wrist topography reported \(27.9\%\) regression error, and vision methods are summarized at approximately \(10.1\%\) error for grip estimation. EchoForce’s central claim is therefore not only absolute accuracy but cross-session, cross-orientation, and cross-user robustness without per-session recalibration.

## 4. EchoForce as deformable acoustic transfer sensing: AST Skin

The paper on Acoustic Soft Tactile Skin does not use the name EchoForce, but the supplied interpretation treats it as an echo or transfer-function approach in which known tones are injected into soft acoustic channels and the resulting spectral amplitudes are mapped to force and contact location. A speaker drives four sinusoids at \(300\), \(500\), \(700\), and \(900\) Hz, each with normalized amplitude \(0.6\), through hollow channels embedded in a soft silicone membrane. Under normal force \(F\), local deformation changes the channel cross-section \(A(x)\), hence the acoustic impedance
$$
Z(x) = \frac{\rho c}{A(x)},
$$
and modifies transmission and reflection. In a simplified discontinuity model,
$$
R = \frac{Z_2-Z_1}{Z_2+Z_1}, \qquad T = \frac{2Z_2}{Z_2+Z_1}.
$$
The channel is therefore treated as a distributed acoustic filter with transfer function \(H(\omega;F)\), and the measured spectral amplitudes at the microphone satisfy
$$
A_k(F)=|H(\omega_k;F)|\cdot A_{k,\mathrm{ref}},
$$
for \(\omega_k \in \{2\pi\cdot 300, 2\pi\cdot 500, 2\pi\cdot 700, 2\pi\cdot 900\}\). The learned feature vector is \(x=[A_3,A_5,A_7,A_9]\), used for both force regression and location classification [2303.17355].

The flat membrane is a \(35 \text{ mm} \times 60 \text{ mm}\) silicone skin cast in a 3D-printed PLA casing, using Polycraft Silskin with Shore A \(\approx 13\) and \(1{:}1\) catalyst ratio. Tested geometries include a single cylindrical channel (AST 1), dual cylinders (AST 2a/2b), dual cones (AST 3a/3b), and mixed cone-plus-cylinder layouts (AST 4a–4d). A frame-less variant, f-AST, integrates transducers in a base while extending a flexible membrane and cylindrical channel to conform to curved surfaces such as a gripper finger. The sensing pipeline uses FFT amplitudes at the four driven frequencies, with \(5100\) calibration points for the flat AST and \(6350\) points for f-AST. Calibration is performed at three discrete locations \(A,B,C\), with a 6-DoF robot and inline load cell pressing in \(0.2\) mm increments until approximately \(30^{+1}\) N; \(50\) waveform samples are recorded at each increment.

Training uses a \(90{:}10\) train:test split and \(10\)-fold cross-validation in MATLAB’s Regression and Classification Learner tools. Regressors and classifiers include Gaussian Process Regression with several kernels, Bagged Ensemble Trees, Weighted kNN, SVM with Gaussian kernel, and a bilayered neural network. Across flat configurations, force RMSE ranges from \(0.72\) N in AST 1 to \(3.60\) N in AST 4b, and contact-location accuracy ranges from \(91.9\%\) in AST 3b to \(98.2\%\) in AST 2b. Over the full operating range, more than \(93\%\) of force estimates fall within \(\pm 1.5\) N over \(0\)–\(30^{+1}\) N. For AST 1, \(82.74\%\) of estimates lie within \(\pm 0.5\) N and \(93.5\%\) within \(\pm 1.0\) N. Contact-location classification typically exceeds \(96\%\); AST 2b reached \(99.01\%\), AST 4c \(97.84\%\), and AST 4d \(97.45\%\).

The f-AST curved-surface variant achieved its best performance with a Bagged Trees Ensemble, with cross-validation error \(1.77\) N for force and \(99.2\%\) accuracy for location. Accuracy bands were \(70.39\%\) within \(\pm 0.5\) N, \(81.73\%\) within \(\pm 1\) N, \(87.40\%\) within \(\pm 1.5\) N, and \(89.60\%\) within \(\pm 2\) N. The system was further demonstrated in real-time gripping-force control on an SMC LEZH gripper mounted to a Franka Emika arm. The gripper reduced width in \(1\) mm steps until the f-AST estimate reached a target of \(2\), \(10\), or \(20\) N, then maintained that grip during lift, move, and place operations. Reported MAE ranges were \(0.02\)–\(0.05\) N at \(2\) N with noise off and \(0.05\)–\(0.09\) N with \(100\) dB white noise on; \(0.14\)–\(0.51\) N at \(10\) N with noise off and \(0.22\)–\(0.91\) N with noise on; and \(0.19\)–\(1.20\) N at \(20\) N with noise off and \(0.21\)–\(1.10\) N with noise on. Robustness tests also included contact scratching with a stiff brush under zero load, where readings remained near-constant for smooth strokes.

## 5. Resonance-based and electronics-free EchoForce variants

A second branch of the EchoForce idea is realized by AcousTac and the supplied EchoForce interpretation built on it: compliant silicone caps and short plastic tubes form air-driven resonant chambers that emit audible tones under a steady airflow, while a remote microphone reads force from resonance shifts. Each taxel is a 3D-printed PLA tube with inner diameter \(6\) mm and an edge-orifice inlet supplied at \(4\)–\(5\) L/min. A silicone hemispherical shell of Smooth-On Dragon Skin 30 forms the soft end-cap. A smartphone microphone placed approximately \(0.5\) m away samples at \(44.1\) kHz and records the tone. Because the transduction is pneumatic-acoustic, there are no wires, ICs, or embedded transducers at the contact point [2307.09730].

The core model is 1D pipe resonance. For tube length \(L\), the open-closed and open-open fundamentals are
$$
f_{oc} = \frac{c}{4L}, \qquad f_{oo} = \frac{c}{2L}.
$$
A practical fit reported for the study is
$$
f = \frac{b_1}{b_2 L} + b_3,
$$
with representative constants given for the theoretical open-closed case and measured fits at \(5\) N and \(10\) N. Cap deformation \(\delta\) changes effective length and therefore resonance. Force-deformation behavior is nonlinear; an empirical fit is
$$
F = \beta_1\, \delta^{\beta_2} + \beta_3.
$$
Sensitivity follows
$$
S_f = \frac{df}{dF} \approx \left(\frac{c}{4L^2}\right)\frac{d\delta}{dF},
$$
which makes the \(1/L^2\) dependence operationally important: shorter tubes are both higher in base frequency and more sensitive.

The geometry is explicitly tunable. Tube lengths of \(41\), \(47\), \(59\), and \(65\) mm were used to create distinct base frequencies. End-cap wall thickness \(t=1\)–\(5\) mm trades sensitivity against force range. With a \(3\) mm hole enforcing monotonic decoding by amplitude thresholding, measured sensitivities were approximately \(19\) Hz/N for \(t=1\) mm over about \(0.5\)–\(2.6\) N, \(8.8\) Hz/N for \(t=2\) mm over about \(0.6\)–\(6.1\) N, \(5.1\) Hz/N for \(t=3\) mm over about \(0.5\)–\(10\) N, \(3.6\) Hz/N for \(t=4\) mm over about \(1.3\)–\(15\) N, and \(3.2\) Hz/N for \(t=5\) mm over about \(1.5\)–\(15\) N. Hole size \(h\) also changes minimum detectable force: for \(t=3\) mm, \(h=5\) mm yielded approximately \(0.2\) N \(F_{\min}\), whereas no hole yielded approximately \(1.6\) N. Small distal masses of \(50\)–\(200\) mg provide an alternative way to eliminate the initial nonmonotonic boundary-condition transition.

Signal processing is lightweight: `spectrogram()` and `tfridge()` track peak frequency at \(25\) Hz update rate, with \(2.5\) Hz frequency bins. For hole-equipped taxels, the amplitude envelope is computed every \(22.7\) ms and smoothed every \(113\) ms before downsampling to \(25\) Hz. Combined with the measured sensitivities, this gives force resolution from approximately \(0.13\) N to approximately \(0.78\) N. A four-taxel array with \(15\) mm center-to-center spacing and lengths \(41\), \(47\), \(59\), and \(65\) mm was demonstrated, as was a three-taxel astrictive gripper using \(L=41\), \(47\), and \(59\) mm tubes with \(2.5\) mm caps. The gripper tracked approximately \(4\) Hz oscillations during hefting and approximately \(1\) Hz during compression. The reported decoding was monotonic with no hysteresis beyond the eliminated transition region, though RMSE and repeatability statistics were not reported.

## 6. Related extensions and broader acoustic-force interpretations

EchoForce also intersects with multimodal wearable force estimation. Wrist2Finger combines a thumb ring carrying an ICM-20948 IMU with a smartwatch-based single-channel EMG sensor, and uses a dual-branch transformer with bidirectional cross-modal attention to output \(21\) 3D hand joints and five fingertip forces continuously. The model uses hidden size \(d=128\), two transformer layers per branch, four heads per layer, and losses that include pose error, force RMSE, smoothness, saturation, and kinematic constraints. Reported performance across \(20\) participants is average MPJPE \(0.57\) cm, fingertip-force RMSE \(0.213\) in normalized units, and Pearson \(r=0.759\), with \(50\) Hz streaming to Unity and \(8\) ms average inference latency on desktop. The paper explicitly includes a section on integrating Wrist2Finger into EchoForce for grasp detection, haptic mapping, safety limits, and personalization [2510.04122].

A different non-contact line uses acoustic pressure emitted by electrostatic adhesion systems. When an EA pad is driven by a bipolar square wave, polarity reversals generate short acoustic pulses captured by a nearby microphone. Peak acoustic pressure increases with object mass, contact area, and drive voltage, and grows roughly with \(\log_{10}(f)\) over \(1\)–\(1000\) Hz. At \(\pm 600\) V, \(5\) Hz, and \(50 \times 50\) mm\(^2\) contact, the paper fits the empirical mass mapping
$$
P = 0.0016\, m + 0.044,
$$
where \(P\) is normalized peak acoustic pressure and \(m\) is mass in grams. Two identical EA pads were also monitored simultaneously with a \(90^\circ\) phase offset, enabling non-overlapping impulse trains for separate decoding. In the reported mass-estimation demo on five objects, RMSE was approximately \(2.3\) g [2505.16609].

At a more theoretical level, acoustokinetics treats acoustic force design itself as an inverse problem. The pressure field is decomposed as
$$
p(\mathbf{r},t)=\Re\{P(\mathbf{r})e^{i\omega t}\}, \qquad P(\mathbf{r})=A(\mathbf{r})e^{i\phi(\mathbf{r})},
$$
with amplitude gradients associated with conservative components and phase gradients with nonconservative components. The framework develops explicit forms for standing waves, pseudo-standing waves, tractor beams, and purely nonconservative fields, and poses inverse design as minimization of \(\int_\Omega \|\mathbf{F}(A,\phi)-\mathbf{F}_{\text{target}}\|^2 d\Omega + R(A,\phi)\) subject to the Helmholtz equation [1910.03670]. Relatedly, under restrictive conditions, the secondary Bjerknes force between oscillating bubbles becomes always attractive and proportional to the product of the bubbles’ virtual masses, \(m_v = 4\pi \rho R_0^3\), with inverse-square distance dependence. In the most compact form summarized in the paper, this yields a gravitational-type structure \(F \propto (m_{v1} m_{v2})/r^2\) [1905.03622]. These works are not EchoForce sensors in the narrow wearable sense, but they broaden the conceptual neighborhood from acoustic force estimation to acoustic force generation and indirect acoustic monitoring of force-bearing states.

## 7. Comparative position, common misconceptions, and open problems

Relative to other tactile and wearable modalities, EchoForce occupies a distinctive trade space. The wristband EchoForce is positioned against sEMG, strain and pressure sensors, IMUs, and vision; its stated advantages are non-contact air-coupled sensing, robustness across sessions and orientations, and user-independent continuous estimation without per-session recalibration [2507.20437]. AST Skin is positioned against capacitive, resistive, piezoelectric, magnetic, optical, and fluidic skins; its distinguishing attributes are compliance, easy affixation to robot parts, and low-cost off-the-shelf transducers located off the contact surface [2303.17355]. AcousTac-derived EchoForce goes further by removing electronics from the contact point entirely, which is described as advantageous in environments hostile to wiring or electronics, including MRI, wet or dirty settings, and volatile handling [2307.09730].

Several misconceptions recur in this space. First, EchoForce is not a single sensing primitive. The wristband uses FMCW matched filtering and differential echo imaging; AST Skin explicitly uses amplitude modulation of driven tones and not frequency, phase, or time-of-flight as the primary signal; AcousTac uses resonance shifts; EA monitoring uses self-generated acoustic pulses. Second, the systems are not universally calibration-free. EchoForce’s user-independent model is trained on a multi-user corpus and can be fine-tuned; AST Skin requires geometry-specific calibration datasets for force and location; AcousTac requires per-taxel calibration of frequency-versus-force, band planning, and often amplitude thresholds. Third, robustness claims are domain-specific. EchoForce reports remount robustness across sessions and orientations, but dynamic wrist motion and gait-induced artifacts were not evaluated. AST Skin was robust to \(100\) dB white noise and incidental scratching, but hysteresis, repeatability, drift, and temperature sensitivity were not explicitly characterized. AcousTac reports monotonic operation without hysteresis once the transition region is managed, yet larger arrays still require careful band separation and microphone placement.

The open problems are correspondingly clear. Multi-touch disambiguation and shear-force estimation remain unresolved in AST-style channel sensors. Large-area arrays raise cross-talk, identifiability, and calibration-scaling issues in both channel-based and resonance-based skins. The wristband EchoForce paper does not report Bland–Altman analysis, Pearson correlation, confidence intervals, or direct real-time latency figures, and code or data were not publicly released. Across the family, long-term drift, temperature dependence, daily-life motion robustness, and on-device deployment remain active engineering and research questions. A plausible implication is that the enduring value of EchoForce lies less in any one sensor embodiment than in a general recipe: engineer a mechanically informative acoustic path, encode deformation into a stable acoustic statistic, and learn or model the inverse map with enough invariance to survive real use.

Source: https://www.emergentmind.com/topics/echoforce