Papers
Topics
Authors
Recent
Search
2000 character limit reached

UMEDA: Unified Multi-modal Efficient Data Fusion for Privacy-Preserving Graph Federated Learning via Spectral-Gated Attention and Diffusion-Based Operator Alignment

Published 8 May 2026 in cs.LG, cs.AI, cs.CR, and cs.DC | (2605.08288v1)

Abstract: Device-free localization trains models from heterogeneous wireless and visual sensors (e.g., Wi-Fi, LiDAR) distributed across edge devices. Federated learning offers a privacy-respecting framework, but is brittle when clients differ in sensor modality and resolution, when their data distributions drift, and when privacy noise destroys the structural signal needed for localization. We propose UMEDA, a graph federated learning framework in which clients form nodes of a global graph that share a continuous integral operator, and aggregation is reformulated as spectral signal processing on this operator. Each client encodes its local sensors with a linear-attention layer whose kernel spectrum is low-rank filtered, suppressing modality-specific residuals so clients with different sensors align in a common low-rank subspace. The server then aggregates client updates via a diffusion model over the kernel's spectral coefficients, treating updates as discretizations of a shared operator rather than topology-bound weights -- this absorbs varying graph sizes and missing modalities without node-wise correspondence. To balance privacy and utility, we add an anisotropic differential-privacy mechanism that projects noise preferentially into the null space of the signal subspace, preserving dominant eigendirections while ensuring formal (ε,δ)(ε, δ)-DP under gradient clipping. On MM-Fi and the RELI11D out-of-distribution benchmark, UMEDA outperforms state-of-the-art federated baselines in accuracy, convergence, and communication efficiency, particularly under high modality heterogeneity and tight privacy budgets.

Summary

  • The paper introduces UMEDA, a federated framework that aligns heterogeneous sensor representations in operator space through spectral-gated linear attention and diffusion-based graph neural operator aggregation.
  • UMEDA achieves 86.0 mm pose MPJPE, 89.6% action accuracy, and 10.5 cm localization RMSE at ε=2, outperforming evaluated non-private and differentially private federated baselines.
  • The method improves cross-modality transfer, communication efficiency, and privacy protection, reaching 24.8 cm RMSE on RELI11D and reducing gradient-inversion SSIM from 0.41 to 0.12 versus isotropic DP.

Problem setting and motivation

Device-free localization (DFL) infers position, pose, and activity from ambient sensor streams—Wi-Fi channel state information, millimeter-wave radar, LiDAR point clouds—without wearable devices. Fusing these modalities is attractive because each sensor compensates for the others' blind spots, but the raw streams are privacy-sensitive and bandwidth-heavy, making federated learning (FL) the natural training paradigm. The authors identify three obstacles that standard FL does not address. First, clients hold different sensors at different resolutions, so weight-space aggregation methods such as FedAvg, FedProx, and SCAFFOLD—which assume identical architectures and input shapes—are inapplicable, and existing graph FL (GFL) methods build inter-client graphs from heuristic similarity scores that become unstable when one client produces Wi-Fi spectrograms and another produces sparse point clouds. Second, client distributions drift across rooms, body sizes, and sensor placements; optimizer-level fixes constrain weight updates but do not align representations. Third, standard differential-privacy (DP) mechanisms inject isotropic Gaussian noise that destroys the low-rank structural signal on which localization depends.

The paper's central hypothesis is that while sensors discretize the world differently, the underlying physical kernel is shared, so federation should occur in operator space rather than weight space. UMEDA instantiates this with three coupled components: a Spectral-Gated Linear Transformer (SGLT) for local encoding, a Diffusion-based Graph Neural Operator (Diff-GNO) for server aggregation, and Subspace-Projected Differential Privacy (SP-DP).

Method

SGLT maps heterogeneous inputs to a latent sequence via linearized attention with a feature map ϕ(⋅)\phi(\cdot) (FAVOR+ random features). The semantic matrix M=ϕ(K)⊤V∈Rd×d\mathbf{M} = \phi(\mathbf{K})^\top\mathbf{V} \in \mathbb{R}^{d\times d} has dimension determined solely by the hidden width dd, independent of input sequence length—a discretization-invariance property the paper proves formally (Lemma: the inner dimension LL contracts away). Modality-specific noise manifests as high-frequency perturbations of M\mathbf{M}'s spectrum, so SGLT applies a spectral gate g(σi)g(\sigma_i)—a differentiable sigmoid relaxation during training, or a hard threshold recovering truncated SVD for analysis—that retains dominant singular directions carrying shared semantics and attenuates high-frequency residuals. A corollary of Eckart–Young–Mirsky shows the hard-gated kernel equals the optimal rank-rr approximation with error σr+1\sigma_{r+1}.

SP-DP privatizes the vectorized kernel update by clipping to norm bound CC, then adding anisotropic Gaussian noise: scale σsig\sigma_{\mathrm{sig}} calibrated by the Gaussian mechanism (M=ϕ(K)⊤V∈Rd×d\mathbf{M} = \phi(\mathbf{K})^\top\mathbf{V} \in \mathbb{R}^{d\times d}0) in the signal subspace defined by the top-M=ϕ(K)⊤V∈Rd×d\mathbf{M} = \phi(\mathbf{K})^\top\mathbf{V} \in \mathbb{R}^{d\times d}1 eigenspace of the previous round's global operator, and larger noise (M=ϕ(K)⊤V∈Rd×d\mathbf{M} = \phi(\mathbf{K})^\top\mathbf{V} \in \mathbb{R}^{d\times d}2, M=ϕ(K)⊤V∈Rd×d\mathbf{M} = \phi(\mathbf{K})^\top\mathbf{V} \in \mathbb{R}^{d\times d}3) in its null space. Because the projector derives strictly from public broadcast information and M=ϕ(K)⊤V∈Rd×d\mathbf{M} = \phi(\mathbf{K})^\top\mathbf{V} \in \mathbb{R}^{d\times d}4, the noise covariance dominates M=ϕ(K)⊤V∈Rd×d\mathbf{M} = \phi(\mathbf{K})^\top\mathbf{V} \in \mathbb{R}^{d\times d}5 in every direction, so the mechanism inherits formal M=ϕ(K)⊤V∈Rd×d\mathbf{M} = \phi(\mathbf{K})^\top\mathbf{V} \in \mathbb{R}^{d\times d}6-DP. This "safe projection" design avoids leakage through the subspace choice itself—an important subtlety, since projecting with a private basis would break the guarantee.

Diff-GNO replaces weight averaging for the semantic block. Each privatized update is projected onto the global basis to yield an M=ϕ(K)⊤V∈Rd×d\mathbf{M} = \phi(\mathbf{K})^\top\mathbf{V} \in \mathbb{R}^{d\times d}7-dimensional spectral coefficient vector (with M=ϕ(K)⊤V∈Rd×d\mathbf{M} = \phi(\mathbf{K})^\top\mathbf{V} \in \mathbb{R}^{d\times d}8, M=ϕ(K)⊤V∈Rd×d\mathbf{M} = \phi(\mathbf{K})^\top\mathbf{V} \in \mathbb{R}^{d\times d}9, dimensionality drops from 65,536 to 256), which mitigates both the curse of dimensionality and overfitting when few clients participate. A score network trained by denoising score matching under a variance-exploding SDE models the distribution of client updates; reverse-time sampling produces a consensus operator update. Non-kernel parameters are aggregated by FedAvg. Notably, the empirical "dd0 privacy amplification" claim—UMEDA at dd1 matching DP-FedAvg at dd2—is explicitly derived empirically from the frontier's linear region, not analytically; the anisotropic allocation does not formally tighten the dd3 accounting.

Experimental results

The evaluation simulates 100 clients on MM-Fi with Dirichlet label skew and three client types spanning coarse/sparse, dense/rich, and mixed-modality configurations, sweeping dd4 at dd5. Key numbers:

Method Pose MPJPE (mm) ↓ Action Top-1 (%) ↑ Loc RMSE (cm) ↓
Centralized (oracle) 62.4 93.8 8.2
Best non-DP baseline (MOON) 88.9 87.6 12.2
DP-FedAvg (isotropic, dd6) 104.5 82.1 15.8
TA-DPFL 98.8 83.9 14.3
UMEDA (dd7) 86.0 89.6 10.5

A striking result is that UMEDA under SP-DP outperforms every non-private federated baseline—including X-Fi + FedAvg at full modality access (91.3 MPJPE)—despite operating under a privacy budget, indicating operator-space alignment more than compensates for DP noise. Ablations attribute gains to both components: removing SGLT raises RMSE to 12.2 and replacing Diff-GNO with FedAvg on the semantic block raises it to 12.9.

On zero-shot transfer from MM-Fi to RELI11D—a strict cross-modality test since the two datasets share no sensor—UMEDA achieves target RMSE 24.8 cm versus 48.5 for FedAvg and 34.1 for FedRod, with a transfer gap of 14.3 versus 20.7+ for all baselines. Communication efficiency improves comparably: UMEDA reaches Loc RMSE ≤ 12.5 in 360 rounds / 90 GB versus 540 rounds / 135 GB for the second-best linear-attention baseline. Robustness sweeps show UMEDA degrades least under increasing discretization heterogeneity, strongest non-IID skew, and test-time missing modalities. Scalability experiments show per-round server cost is invariant to total population dd8 (constant from dd9 to LL0), since score-model training operates on fixed-dimension LL1 states. Empirically, gradient-inversion attacks recover recognizable silhouettes under isotropic DP at LL2 (mean SSIM 0.41) but only structureless noise under SP-DP (SSIM 0.12), complementing the formal guarantee with operational evidence.

Limitations and open questions

The paper concedes several constraints. SGLT requires SVD of the LL3 semantic matrix and Diff-GNO adds server-side score-model training and sampling, costs that may be non-trivial at scale. The strongest invariance holds only in the SGLT operator-update space; if performance is dominated by other parameter blocks (e.g., modality-specific front-ends), additional alignment mechanisms would be needed. Diffusion schedules and solver discretization affect stability, particularly when update distributions are strongly multi-modal or heavily perturbed by DP noise. The DP analysis relies on bounded LL4 sensitivity via clipping and does not yet incorporate round-wise composition (RDP/moments accountant with subsampling amplification), secure-aggregation assumptions, or long-horizon participation heterogeneity. Finally, treating LL5 as resolution-agnostic presumes modality-invariant semantics are representable in a shared kernel; extreme sensing gaps may produce operator distributions that resist alignment without stronger conditioning. Fairness across demographics is not evaluated. Open questions include adaptive (learnable) spectral gates, conditioning the score model on modality descriptors, and robustness under asynchronous partial participation.

Conclusion

UMEDA reformulates multi-modal federated learning as spectral signal processing over a shared continuous integral operator, combining spectral-gated linear attention, diffusion-based generative aggregation of operator updates, and subspace-projected differential privacy. The empirical case is strong: state-of-the-art accuracy under matched privacy budgets, an LL6 effective privacy gain at matched utility, substantially reduced cross-dataset transfer gap, and population-invariant server cost. The framework's validity rests on the assumption that shared semantics live in a compact common kernel subspace—an assumption the ablations support within tested regimes but whose boundary under extreme modality shifts remains unresolved.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.