Migration-aware adaptive expert placement

Develop migration-aware adaptive expert-placement strategies that balance the long-term inference benefits of updating expert-device associations against the communication, storage, and service-interruption costs of expert migration.

Background

AirMoE determines expert-device associations on a slow timescale and assumes that the assignment remains fixed during online inference. The paper identifies adaptive reassignment as an unresolved problem: changing placements could improve inference reliability as workloads or wireless conditions evolve, but expert migration incurs communication, storage, and service-interruption costs that must be incorporated into the placement decision.

References

From a broader perspective, AirMoE highlights several open problems in wireless MoE serving systems. Robust AirMoE aggregation can be revisited from an inference-aware perspective under imperfect CSI, synchronization errors, interference, and device mobility. Beyond empirical layer-sensitivity calibration, analytical models are needed to characterize how over-the-air aggregation distortion propagates across MoE layers and affects E2E inference performance. Another important direction is migration-aware adaptive expert placement, which balances the long-term inference benefit of updating expert-device associations against the communication, storage, and service-interruption costs incurred by expert migration.

AirMoE: Realizing Over-the-Air Distributed Mixture-of-Experts Inference at the Wireless Edge  (2608.22932 - Yang et al., 24 Aug 2026) in Concluding Remarks, Section VI