Load Agent in Multi-Domain Systems
- Load agent is an autonomous entity that monitors and responds to load as a state variable across diverse domains such as computing, networks, and robotics.
- It employs decentralized learning and control methods, including DQN, multi-agent actor-critic, and consensus-based estimation, to optimize performance metrics like throughput, latency, and efficiency.
- Applications range from distributed computing and fog offloading to energy management, radio systems, and cognitive load estimation, emphasizing feasibility and local autonomy in control loops.
Searching arXiv for recent and representative papers on “load agent” and closely related multi-agent load-balancing/load-estimation formulations. “Load agent” denotes an autonomous computational or physical entity whose decision loop is organized around sensing, estimating, balancing, or reacting to some notion of load. In the cited literature, that notion ranges from computational demand and queueing pressure to radio-cell occupancy, household electrical demand, suspended payload dynamics, passenger occupancy, and human cognitive load. Accordingly, a load agent may be a grid node that joins clusters to match CPU demand, a base station that learns handoff parameters, a household energy controller, a robot estimating the state of a manipulated object, or an EEG- or gaze-based software agent inferring cognitive workload (Banerjee et al., 2015, Kang et al., 2023, Qin et al., 2021, Franchi et al., 2016, Xu et al., 25 Jun 2026, Wang et al., 7 May 2026).
1. Load as a state variable across domains
The literature uses “load” in several technically distinct but structurally related senses. In distributed computing, load is defined at three levels: per-task load as , per-cluster load as , and global load as the task queue plus system utilization; the control problem is to allocate tasks from a global queue and form clusters such that while minimizing and and maximizing utilization (Banerjee et al., 2015). In cellular networks, load is the number of User Equipments served in each channel, sector, and base station, together with bandwidth utilization and throughput, and the learning objective is to redistribute UE traffic while improving average throughput, minimum throughput, and fairness (Kang et al., 2023). In fog computing, the central observable is queue length at candidate fog nodes, and the agent’s reward is the negative queue length at the previously chosen fog node (Ebrahim et al., 2024). In residential energy scheduling, each household agent controls AC and EV power under a balance equation , so load is both an electrical demand and a scheduling variable constrained by comfort and battery dynamics (Qin et al., 2021).
Other domains use the same term for latent or embodied state. In transit estimation, passenger load is the onboard occupancy state after stop , recursively updated from boarding and alighting flows and constrained to remain in 0 (Xu et al., 19 May 2026). In gaze-based smart-glasses systems, cognitive load is a discrete variable 1 inferred from temporal gaze features (Wang et al., 7 May 2026). In the NeuraDock framework, visual cognitive load is operationalized through within-subject posterior Alpha suppression between Rest and Task, with workload metrics computed only after preprocessing and quality-control gating (Xu et al., 25 Jun 2026). This suggests a family of architectures rather than a single canonical definition: a load agent is identified less by domain than by the fact that “load” is the state variable around which sensing, inference, and action are organized.
2. Distributed computing and digital infrastructure
In distributed computing, the most literal use of the term appears in the dRAP system, where each computer in the grid is explicitly a “load agent” that senses its own load and neighborhood, reads a global task queue, forms or dissolves clusters, and transitions among four modes: free/unclustered, single-node executing, clustered idle, and clustered executing (Banerjee et al., 2015). The local selection heuristic is to minimize 2 for a free node and 3 for an idle cluster, while overloaded single nodes recruit neighbors until 4. On the simulator, dRAP outperforms FIFO scheduling with 5 versus 6, 7 versus 8, cluster utilization 9 by design versus about 0 under FIFO, and node utilization around 1–2 versus 3–4 (Banerjee et al., 2015). Its main global cost is queue traversal, bounded by per-timestep worst-case 5 and whole-run 6, although empirical traversal is reported as only about 7 of worst case.
A related but more explicitly learned formulation appears in fog computing. There, each Access Point hosts one Double DQN load agent that observes workload category and candidate fog-node queues, selects an offloading destination, and is trained from the reward 8 (Ebrahim et al., 2024). The design is fully distributed: no parameter sharing, no centralized critic, and no direct inter-agent communication. State dissemination is made realistic through an interval-based Gossip-like protocol with a 3-second period rather than idealized per-decision perfect observability. The reported outcome is lower waiting time than centralized RL and heuristic baselines, together with better fairness in delay and utilization, even when observations are stale (Ebrahim et al., 2024).
Data-center load balancing has also been formalized with the load balancers themselves as agents. In the Markov-potential-game approach, each load balancer learns from local observations while the reward is constructed from variance-based fairness of server completion times, which is used as a potential function aligning local improvement with global fairness (Yao et al., 2022). The distributed design avoids centralized-training communication overhead; the reported control-traffic burden for CTDE in the large-scale setup is about 9 MiB per episode, with probing experiments showing per-packet RTT increases of up to 0 under added control traffic (Yao et al., 2022). This use of fairness as a potential function is important because it turns independent local learning into a coordinated equilibrium-seeking process.
A different infrastructure interpretation appears in residential microgrids, where each household HEMS is a load agent with decentralized actors and distributed critics. The DADC architecture preserves privacy because the global critic is a learned function of individual critics computed only from local observations. In the reported experiments, DADC reduces average total cost from 1 under independent actor-critic to 2, and adjustment cost from 3 to 4 (Qin et al., 2021). In performance testing, RELOAD uses the same pattern in a test-engineering setting: the agent adapts per-transaction workload increments until the objective 5 ms or 6 is reached, achieving about 7–8 test-cost savings versus a standard baseline and 9–0 versus random testing, with 1 savings retained in transfer learning (Moghadam et al., 2021).
3. Communication networks and radio systems
In radio networks, a load agent is typically a base station or radio unit that adjusts association and resource-allocation parameters to redistribute traffic. In the robust multi-agent attention actor-critic framework for cellular networks, each base station is an agent in a Markov game 2, with local state components 3, 4, and 5, and actions consisting of AULB thresholds 6 and IULB reselection parameters 7 (Kang et al., 2023). The reward is 8, combining average throughput, minimum throughput, and inverse throughput variance. A centralized critic with self-attention and an adversarial nature agent is used during training, while execution is decentralized. Relative to Robust-MADDPG, Robust-MA3C improves average throughput by up to 9, minimum throughput by up to 0, and the standard-deviation metric by up to 1; the ablation without the nature agent already improves average throughput by up to 2, minimum throughput by up to 3, and fairness by up to 4 (Kang et al., 2023).
A more explicitly competitive formulation appears in wireless SON load balancing. There, each base station observes 5, where 6 is attached-UE ratio, 7 is resource-block utilization, and 8 is the Cell Individual Offset; the action is the continuous CIO value 9 dB (Rivera et al., 2021). MADDPG-AP adds multiple sub-policies per agent and an LSTM policy predictor trained on a ranked buffer. In the reported simulations, it reduces delay by about 0–1, packet loss ratio by about 2–3, and average convergence from about 4 episodes to about 5, a 6 reduction (Rivera et al., 2021). The result is not higher total throughput so much as better latency and loss under competitive interaction.
O-RAN introduces another agentic partition. In mmLBRA, each O-RU is a multi-armed-bandit agent operating at two timescales: mmLBRA-LB in the Non-RT RIC for mobility load balancing, and mmLBRA-RA in the Near-RT RIC for subchannel and power allocation (Lai et al., 2023). Load is defined as subchannel utilization 7, imbalance as 8, the LB reward as 9, and the RA reward as rate. UCB is used instead of 0-greedy. In the simulation setting with 7 O-RUs, 80 UEs, and 80 subchannels, the reported outcome is lower standard deviation of subchannel utilization and higher effective sum-rate than default, rule-based, and 1-greedy alternatives (Lai et al., 2023). Across these radio papers, load agents are distinguished by local observation, decentralized or near-decentralized action, and explicit handling of coupling through either critics, bandits, or handoff logic.
4. Embodied systems and manipulation of physical loads
In embodied systems, “load agent” may refer not to an abstract controller but to a robot physically coupled to a shared object. In the distributed-estimation framework of Franchi, Petitti, and Rizzo, each agent manipulates an unknown planar rigid body, applies a 2D wrench, measures its contact-point velocity, and communicates over a connected graph to estimate the load’s kinematic state and parameters such as 2, 3, 4, 5, and 6 (Franchi et al., 2016). The core rigid-body relation is 7, while one safe stopping controller is 8. A Lyapunov function 9 is used to show asymptotic stopping under that law. The technical report emphasizes nonlinear observability conditions, consensus-based estimation, and Monte Carlo robustness rather than centralized fusion (Franchi et al., 2016). In this sense, the load agent is simultaneously an actuator and a distributed observer.
Aerial manipulation extends the idea to fully decentralized control of a cable-suspended payload. Each MAV agent observes the load pose, goal pose, its own state, and a one-hot identifier, but does not observe other MAVs and does not communicate with them; coordination is implicit through the shared load (Zeng et al., 2 Aug 2025). The action space is six-dimensional, consisting of desired linear acceleration and body rates. Real-world experiments with three MAVs report position RMSE 0 m and attitude RMSE 1, compared with 2 m and 3 for a centralized NMPC baseline (Zeng et al., 2 Aug 2025). The same policy class is also shown to tolerate heterogeneous controllers and complete in-flight loss of one MAV. This suggests that, for some tightly coupled physical systems, the load itself can function as a communication channel: agents infer the team state from the object’s motion rather than from explicit message passing.
5. Cognitive, passenger, and human-state load agents
A distinct branch of the literature uses the term for systems that estimate, rather than redistribute, load. NeuraDock is an EEG-based visual cognitive-load agent that reads 7-channel EEG at 250 Hz, preprocesses it, applies quality control, extracts posterior Alpha features in the 8–12 Hz band, and exposes a real-time workload index through a dashboard and HTTP API (Xu et al., 25 Jun 2026). Its workflow is quality-gated: downstream Alpha and workload metrics are computed only after preprocessing and QC gating. In the tutorial mini-dataset, the agent processed 18 recordings, produced 10 within-subject comparisons, observed task-related posterior Alpha suppression in 7 of 10 contrasts, and reported baseline repeatability of Pearson 4 and ICC(C,1)5 for median posterior log Alpha (Xu et al., 25 Jun 2026). Local replay benchmarks reported median core processing latency 6 ms and median HTTP endpoint latency 7 ms (Xu et al., 25 Jun 2026). The agent therefore couples state estimation with explicit confidence reporting.
GazeMind performs an analogous role for smart glasses. It converts gaze into Temporal Gaze Encoding, augments it with Task-Guidance Reasoning, Adaptive User Profile Calibration, and Cognitive Retrieval-Augmented Generation, and predicts 8 along with a textual rationale (Wang et al., 7 May 2026). The associated CogLoad-Bench contains 152 participants, 40+ hours of multimodal data, and 10K+ real-time annotations. On the cross-user evaluation split, GazeMind reaches 9 accuracy and 0 macro-F1, compared with 1 and 2 for zero-shot GPT-4o and 3 and 4 for the best supervised baseline LSTM (Wang et al., 7 May 2026). The ablation shows cumulative gains from TGR, AUP, and CogRAG, culminating in the full framework. Here the load agent is not a controller in the usual systems sense but an interpretable reasoning system over physiological signals.
Transit load estimation pushes this pattern into operational state estimation. The state-centric multi-agent framework models passenger occupancy after stop 5 as 6, uses a Perception module 7 to predict 8, a Physical module 9 to enforce 00, 01, and 02, and then fuses this with a Wi-Fi-derived anchor 03 using a trust weight 04 derived from disagreement and physics residuals (Xu et al., 19 May 2026). On all test trips, the proposed rule-based trust fusion without macro-shift achieves RMSE 05, MAE 06, and trip-end absolute error 07, compared with RMSE 08 and MAE 09 for Perception-only open loop (Xu et al., 19 May 2026). On APC-inconsistent trips, RMSE drops from 10 for Perception-only to 11 for the proposed rule-based fusion (Xu et al., 19 May 2026). In this domain, the load agent is a closed-loop state estimator that enforces feasibility and dynamically allocates trust among heterogeneous evidence sources.
6. Recurrent architectural themes, controversies, and limits
Several architectural motifs recur. First, many systems implement decentralized execution with either centralized training or no central coordination at all. Robust-MA3C uses centralized critics with decentralized actors (Kang et al., 2023); DADC uses distributed critics to preserve privacy (Qin et al., 2021); fog offloading is fully independent (Ebrahim et al., 2024); data-center load balancing avoids CTDE overhead through a Markov potential game (Yao et al., 2022). Second, many papers separate proposal from feasibility. The transit estimator projects 12 onto physically admissible flows (Xu et al., 19 May 2026); the distribution-restoration framework uses invalid action masking to ensure zero constraint violations while also shrinking the feasible action space (Vu et al., 2023); NeuraDock withholds workload metrics until QC passes (Xu et al., 25 Jun 2026). Third, many load agents do not rebalance by migrating already-running work. dRAP balances at assignment time, with no explicit task migration (Banerjee et al., 2015); the task-allocation framework with load management encourages idling and penalizes unnecessary reassignment rather than reactive reallocation after overload (Wu et al., 2022).
The literature also shows that “load agent” should not be reduced to computational balancing alone. In task allocation, the HTLM Dec-POMDP makes idling a first-class action and uses 13 to preserve reserve capacity; with moderate idle incentive, unused capability rises from about 14 to about 15 while completion steps remain almost unchanged (Wu et al., 2022). In LLM systems, CoThinker reframes overload through Cognitive Load Theory and distributes intrinsic load via specialization, transactive memory, and a small-world communication graph; with default 16, 17, 18, and 19, it improves normalized LiveBench averages over IO, CoT, MAD, DMAD, and Self-Refine on high-load tasks, while also showing that extra agents can hurt low-intrinsic-load instruction-following tasks (Shang et al., 7 Jun 2025). A plausible implication is that the most stable definition of a load agent is not “an entity that balances servers,” but “an entity whose policy is organized around a load-bearing latent variable and a feasibility-aware control loop.”
Common limits are equally consistent. dRAP still pays queue-traversal overhead and lacks preemptive migration (Banerjee et al., 2015). Fog MARL assumes synchronous observation intervals and does not provide convergence guarantees for independent DDQL (Ebrahim et al., 2024). Robust-MA3C provides empirical convergence but no formal proof of Nash-equilibrium convergence (Kang et al., 2023). The distributed load-restoration framework depends on centralized environment evaluation during training, and its gains hinge on action masking (Vu et al., 2023). NeuraDock is explicitly within-subject and non-diagnostic (Xu et al., 25 Jun 2026), GazeMind depends on self-reported cognitive-load labels and task-specific rule synthesis (Wang et al., 7 May 2026), and the transit macro-correction module can over-correct and is therefore left off by default in the reported experiments (Xu et al., 19 May 2026). These limits are not incidental: they reflect a common tension between local autonomy, partial observability, feasibility guarantees, and the cost of maintaining trustworthy load estimates or load-balancing actions.