- The paper introduces UMEDA, a federated framework that aligns heterogeneous sensor representations in operator space through spectral-gated linear attention and diffusion-based graph neural operator aggregation.
- UMEDA achieves 86.0 mm pose MPJPE, 89.6% action accuracy, and 10.5 cm localization RMSE at ε=2, outperforming evaluated non-private and differentially private federated baselines.
- The method improves cross-modality transfer, communication efficiency, and privacy protection, reaching 24.8 cm RMSE on RELI11D and reducing gradient-inversion SSIM from 0.41 to 0.12 versus isotropic DP.
Problem setting and motivation
Device-free localization (DFL) infers position, pose, and activity from ambient sensor streams—Wi-Fi channel state information, millimeter-wave radar, LiDAR point clouds—without wearable devices. Fusing these modalities is attractive because each sensor compensates for the others' blind spots, but the raw streams are privacy-sensitive and bandwidth-heavy, making federated learning (FL) the natural training paradigm. The authors identify three obstacles that standard FL does not address. First, clients hold different sensors at different resolutions, so weight-space aggregation methods such as FedAvg, FedProx, and SCAFFOLD—which assume identical architectures and input shapes—are inapplicable, and existing graph FL (GFL) methods build inter-client graphs from heuristic similarity scores that become unstable when one client produces Wi-Fi spectrograms and another produces sparse point clouds. Second, client distributions drift across rooms, body sizes, and sensor placements; optimizer-level fixes constrain weight updates but do not align representations. Third, standard differential-privacy (DP) mechanisms inject isotropic Gaussian noise that destroys the low-rank structural signal on which localization depends.
The paper's central hypothesis is that while sensors discretize the world differently, the underlying physical kernel is shared, so federation should occur in operator space rather than weight space. UMEDA instantiates this with three coupled components: a Spectral-Gated Linear Transformer (SGLT) for local encoding, a Diffusion-based Graph Neural Operator (Diff-GNO) for server aggregation, and Subspace-Projected Differential Privacy (SP-DP).
Method
SGLT maps heterogeneous inputs to a latent sequence via linearized attention with a feature map ϕ(⋅) (FAVOR+ random features). The semantic matrix M=ϕ(K)⊤V∈Rd×d has dimension determined solely by the hidden width d, independent of input sequence length—a discretization-invariance property the paper proves formally (Lemma: the inner dimension L contracts away). Modality-specific noise manifests as high-frequency perturbations of M's spectrum, so SGLT applies a spectral gate g(σi​)—a differentiable sigmoid relaxation during training, or a hard threshold recovering truncated SVD for analysis—that retains dominant singular directions carrying shared semantics and attenuates high-frequency residuals. A corollary of Eckart–Young–Mirsky shows the hard-gated kernel equals the optimal rank-r approximation with error σr+1​.
SP-DP privatizes the vectorized kernel update by clipping to norm bound C, then adding anisotropic Gaussian noise: scale σsig​ calibrated by the Gaussian mechanism (M=ϕ(K)⊤V∈Rd×d0) in the signal subspace defined by the top-M=ϕ(K)⊤V∈Rd×d1 eigenspace of the previous round's global operator, and larger noise (M=ϕ(K)⊤V∈Rd×d2, M=ϕ(K)⊤V∈Rd×d3) in its null space. Because the projector derives strictly from public broadcast information and M=ϕ(K)⊤V∈Rd×d4, the noise covariance dominates M=ϕ(K)⊤V∈Rd×d5 in every direction, so the mechanism inherits formal M=ϕ(K)⊤V∈Rd×d6-DP. This "safe projection" design avoids leakage through the subspace choice itself—an important subtlety, since projecting with a private basis would break the guarantee.
Diff-GNO replaces weight averaging for the semantic block. Each privatized update is projected onto the global basis to yield an M=ϕ(K)⊤V∈Rd×d7-dimensional spectral coefficient vector (with M=ϕ(K)⊤V∈Rd×d8, M=ϕ(K)⊤V∈Rd×d9, dimensionality drops from 65,536 to 256), which mitigates both the curse of dimensionality and overfitting when few clients participate. A score network trained by denoising score matching under a variance-exploding SDE models the distribution of client updates; reverse-time sampling produces a consensus operator update. Non-kernel parameters are aggregated by FedAvg. Notably, the empirical "d0 privacy amplification" claim—UMEDA at d1 matching DP-FedAvg at d2—is explicitly derived empirically from the frontier's linear region, not analytically; the anisotropic allocation does not formally tighten the d3 accounting.
Experimental results
The evaluation simulates 100 clients on MM-Fi with Dirichlet label skew and three client types spanning coarse/sparse, dense/rich, and mixed-modality configurations, sweeping d4 at d5. Key numbers:
| Method |
Pose MPJPE (mm) ↓ |
Action Top-1 (%) ↑ |
Loc RMSE (cm) ↓ |
| Centralized (oracle) |
62.4 |
93.8 |
8.2 |
| Best non-DP baseline (MOON) |
88.9 |
87.6 |
12.2 |
| DP-FedAvg (isotropic, d6) |
104.5 |
82.1 |
15.8 |
| TA-DPFL |
98.8 |
83.9 |
14.3 |
| UMEDA (d7) |
86.0 |
89.6 |
10.5 |
A striking result is that UMEDA under SP-DP outperforms every non-private federated baseline—including X-Fi + FedAvg at full modality access (91.3 MPJPE)—despite operating under a privacy budget, indicating operator-space alignment more than compensates for DP noise. Ablations attribute gains to both components: removing SGLT raises RMSE to 12.2 and replacing Diff-GNO with FedAvg on the semantic block raises it to 12.9.
On zero-shot transfer from MM-Fi to RELI11D—a strict cross-modality test since the two datasets share no sensor—UMEDA achieves target RMSE 24.8 cm versus 48.5 for FedAvg and 34.1 for FedRod, with a transfer gap of 14.3 versus 20.7+ for all baselines. Communication efficiency improves comparably: UMEDA reaches Loc RMSE ≤ 12.5 in 360 rounds / 90 GB versus 540 rounds / 135 GB for the second-best linear-attention baseline. Robustness sweeps show UMEDA degrades least under increasing discretization heterogeneity, strongest non-IID skew, and test-time missing modalities. Scalability experiments show per-round server cost is invariant to total population d8 (constant from d9 to L0), since score-model training operates on fixed-dimension L1 states. Empirically, gradient-inversion attacks recover recognizable silhouettes under isotropic DP at L2 (mean SSIM 0.41) but only structureless noise under SP-DP (SSIM 0.12), complementing the formal guarantee with operational evidence.
Limitations and open questions
The paper concedes several constraints. SGLT requires SVD of the L3 semantic matrix and Diff-GNO adds server-side score-model training and sampling, costs that may be non-trivial at scale. The strongest invariance holds only in the SGLT operator-update space; if performance is dominated by other parameter blocks (e.g., modality-specific front-ends), additional alignment mechanisms would be needed. Diffusion schedules and solver discretization affect stability, particularly when update distributions are strongly multi-modal or heavily perturbed by DP noise. The DP analysis relies on bounded L4 sensitivity via clipping and does not yet incorporate round-wise composition (RDP/moments accountant with subsampling amplification), secure-aggregation assumptions, or long-horizon participation heterogeneity. Finally, treating L5 as resolution-agnostic presumes modality-invariant semantics are representable in a shared kernel; extreme sensing gaps may produce operator distributions that resist alignment without stronger conditioning. Fairness across demographics is not evaluated. Open questions include adaptive (learnable) spectral gates, conditioning the score model on modality descriptors, and robustness under asynchronous partial participation.
Conclusion
UMEDA reformulates multi-modal federated learning as spectral signal processing over a shared continuous integral operator, combining spectral-gated linear attention, diffusion-based generative aggregation of operator updates, and subspace-projected differential privacy. The empirical case is strong: state-of-the-art accuracy under matched privacy budgets, an L6 effective privacy gain at matched utility, substantially reduced cross-dataset transfer gap, and population-invariant server cost. The framework's validity rests on the assumption that shared semantics live in a compact common kernel subspace—an assumption the ablations support within tested regimes but whose boundary under extreme modality shifts remains unresolved.