---
title: Load Agent in Multi-Domain Systems
url: https://www.emergentmind.com/topics/load-agent
type: topic
---

# Load Agent in Multi-Domain Systems

Searching arXiv for recent and representative papers on “load agent” and closely related multi-agent load-balancing/load-estimation formulations.
“Load agent” denotes an autonomous computational or physical entity whose decision loop is organized around sensing, estimating, balancing, or reacting to some notion of load. In the cited literature, that notion ranges from computational demand and queueing pressure to radio-cell occupancy, household electrical demand, suspended payload dynamics, passenger occupancy, and human cognitive load. Accordingly, a load agent may be a grid node that joins clusters to match CPU demand, a base station that learns handoff parameters, a household energy controller, a robot estimating the state of a manipulated object, or an EEG- or gaze-based software agent inferring cognitive workload [1509.06420] [2303.08003] [2110.02784] [1602.01891] [2606.26518] [2605.05790].

## 1. Load as a state variable across domains

The literature uses “load” in several technically distinct but structurally related senses. In distributed computing, load is defined at three levels: per-task load as \((CPU_{req}, time_{rem})\), per-cluster load as \(CPU_{cluster}\), and global load as the task queue plus system utilization; the control problem is to allocate tasks from a global queue \(Q\) and form clusters such that \(CPU_{cluster} \approx CPU_{req}\) while minimizing \(T_{complete}\) and \(T_{wait}\) and maximizing utilization [1509.06420]. In cellular networks, load is the number of User Equipments served in each channel, sector, and base station, together with bandwidth utilization and throughput, and the learning objective is to redistribute UE traffic while improving average throughput, minimum throughput, and fairness [2303.08003]. In fog computing, the central observable is queue length \(\mathbb{Q}_n\) at candidate fog nodes, and the agent’s reward is the negative queue length at the previously chosen fog node [2405.12236]. In residential energy scheduling, each household agent controls AC and EV power under a balance equation \(P_t^{DG}=\sum_i(P^{BL}_{i,t}+P^{AC}_{i,t}+P^{EV}_{i,t})\), so load is both an electrical demand and a scheduling variable constrained by comfort and battery dynamics [2110.02784].

Other domains use the same term for latent or embodied state. In transit estimation, passenger load is the onboard occupancy state \(L_k\) after stop \(k\), recursively updated from boarding and alighting flows and constrained to remain in \([0,C]\) [2605.19834]. In gaze-based smart-glasses systems, cognitive load is a discrete variable \(y \in \{\text{low}, \text{moderate}, \text{high}\}\) inferred from temporal gaze features [2605.05790]. In the NeuraDock framework, visual cognitive load is operationalized through within-subject posterior Alpha suppression between Rest and Task, with workload metrics computed only after preprocessing and quality-control gating [2606.26518]. This suggests a family of architectures rather than a single canonical definition: a load agent is identified less by domain than by the fact that “load” is the state variable around which sensing, inference, and action are organized.

## 2. Distributed computing and digital infrastructure

In distributed computing, the most literal use of the term appears in the dRAP system, where each computer in the grid is explicitly a “load agent” that senses its own load and neighborhood, reads a global task queue, forms or dissolves clusters, and transitions among four modes: free/unclustered, single-node executing, clustered idle, and clustered executing [1509.06420]. The local selection heuristic is to minimize \(|CPU_{req}-1|\) for a free node and \(|CPU_{req}-CPU_{cluster}|\) for an idle cluster, while overloaded single nodes recruit neighbors until \(CPU_{cluster}=CPU_{req}\). On the simulator, dRAP outperforms FIFO scheduling with \(T_{complete} \approx 845.6\) versus \(1071.2\), \(T_{wait} \approx 342.54\) versus \(475.31\), cluster utilization \(\mu_{cluster}=1\) by design versus about \(56\%\) under FIFO, and node utilization around \(90\)–\(95\%\) versus \(70\)–\(75\%\) [1509.06420]. Its main global cost is queue traversal, bounded by per-timestep worst-case \(O(nm)\) and whole-run \(O(n^2m)\), although empirical traversal is reported as only about \(10\%\) of worst case.

A related but more explicitly learned formulation appears in fog computing. There, each Access Point hosts one Double DQN load agent that observes workload category and candidate fog-node queues, selects an offloading destination, and is trained from the reward \(r_t=-\mathbb{Q}_{a_{t-1},t}\) [2405.12236]. The design is fully distributed: no parameter sharing, no centralized critic, and no direct inter-agent communication. State dissemination is made realistic through an interval-based Gossip-like protocol with a 3-second period rather than idealized per-decision perfect observability. The reported outcome is lower waiting time than centralized RL and heuristic baselines, together with better fairness in delay and utilization, even when observations are stale [2405.12236].

Data-center load balancing has also been formalized with the load balancers themselves as agents. In the Markov-potential-game approach, each load balancer learns from local observations while the reward is constructed from variance-based fairness of server completion times, which is used as a potential function aligning local improvement with global fairness [2206.01451]. The distributed design avoids centralized-training communication overhead; the reported control-traffic burden for CTDE in the large-scale setup is about \(221\) MiB per episode, with probing experiments showing per-packet RTT increases of up to \(10\times\) under added control traffic [2206.01451]. This use of fairness as a potential function is important because it turns independent local learning into a coordinated equilibrium-seeking process.

A different infrastructure interpretation appears in residential microgrids, where each household HEMS is a load agent with decentralized actors and distributed critics. The DADC architecture preserves privacy because the global critic is a learned function of individual critics computed only from local observations. In the reported experiments, DADC reduces average total cost from \(65.39 \pm 1.20\) under independent actor-critic to \(58.20 \pm 0.95\), and adjustment cost from \(4.86 \pm 0.78\) to \(2.41 \pm 0.33\) [2110.02784]. In performance testing, RELOAD uses the same pattern in a test-engineering setting: the agent adapts per-transaction workload increments until the objective \(RT>1500\) ms or \(ER>20\%\) is reached, achieving about \(30\%\)–\(34\%\) test-cost savings versus a standard baseline and \(17\%\)–\(20\%\) versus random testing, with \(25\%\) savings retained in transfer learning [2104.12893].

## 3. Communication networks and radio systems

In radio networks, a load agent is typically a base station or radio unit that adjusts association and resource-allocation parameters to redistribute traffic. In the robust multi-agent attention actor-critic framework for cellular networks, each base station is an agent in a Markov game \(G=(\mathcal{N},\mathcal{S},\mathcal{A},\mathcal{P},\mathcal{R})\), with local state components \(\boldsymbol{s}_{ue}\), \(\boldsymbol{s}_{band}\), and \(\boldsymbol{s}_{tput}\), and actions consisting of AULB thresholds \(\boldsymbol{\alpha}^k \in [-2\,\mathrm{dB},2\,\mathrm{dB}]\) and IULB reselection parameters \(\boldsymbol{\beta}^k,\boldsymbol{\gamma}^k \in [-20\,\mathrm{dB},20\,\mathrm{dB}]\) [2303.08003]. The reward is \(r=G_{\text{aver}}+G_{\text{min}}-G_{\text{sd}}\), combining average throughput, minimum throughput, and inverse throughput variance. A centralized critic with self-attention and an adversarial nature agent is used during training, while execution is decentralized. Relative to Robust-MADDPG, Robust-MA3C improves average throughput by up to \(18\%\), minimum throughput by up to \(37\%\), and the standard-deviation metric by up to \(45\%\); the ablation without the nature agent already improves average throughput by up to \(12\%\), minimum throughput by up to \(31\%\), and fairness by up to \(42\%\) [2303.08003].

A more explicitly competitive formulation appears in wireless SON load balancing. There, each base station observes \(S_j(t)=[u_j(t)\; p_j(t)\; \omega_j(t)]\), where \(u_j\) is attached-UE ratio, \(p_j\) is resource-block utilization, and \(\omega_j\) is the Cell Individual Offset; the action is the continuous CIO value \(\omega_j(t)\in[-9,9]\) dB [2110.07050]. MADDPG-AP adds multiple sub-policies per agent and an LSTM policy predictor trained on a ranked buffer. In the reported simulations, it reduces delay by about \(20.9\)–\(25.7\%\), packet loss ratio by about \(28.1\)–\(32.1\%\), and average convergence from about \(61\) episodes to about \(18\), a \(70.49\%\) reduction [2110.07050]. The result is not higher total throughput so much as better latency and loss under competitive interaction.

O-RAN introduces another agentic partition. In mmLBRA, each O-RU is a multi-armed-bandit agent operating at two timescales: mmLBRA-LB in the Non-RT RIC for mobility load balancing, and mmLBRA-RA in the Near-RT RIC for subchannel and power allocation [2303.14355]. Load is defined as subchannel utilization \(\Omega_{g,s}\), imbalance as \(\eta_{g,s}\), the LB reward as \(-\eta_{g,s}\), and the RA reward as rate. UCB is used instead of \(\epsilon\)-greedy. In the simulation setting with 7 O-RUs, 80 UEs, and 80 subchannels, the reported outcome is lower standard deviation of subchannel utilization and higher effective sum-rate than default, rule-based, and \(\epsilon\)-greedy alternatives [2303.14355]. Across these radio papers, load agents are distinguished by local observation, decentralized or near-decentralized action, and explicit handling of coupling through either critics, bandits, or handoff logic.

## 4. Embodied systems and manipulation of physical loads

In embodied systems, “load agent” may refer not to an abstract controller but to a robot physically coupled to a shared object. In the distributed-estimation framework of Franchi, Petitti, and Rizzo, each agent manipulates an unknown planar rigid body, applies a 2D wrench, measures its contact-point velocity, and communicates over a connected graph to estimate the load’s kinematic state and parameters such as \(\mathbf{v}_C\), \(\omega\), \(\mathbf{z}_C\), \(m\), and \(J\) [1602.01891]. The core rigid-body relation is \(\mathbf{v}_{C_i}=\mathbf{v}_C+\omega(\mathbf{z}_C+\mathbf{z}_i)^\perp\), while one safe stopping controller is \(\mathbf{f}_i=-b\mathbf{v}_{C_i}\). A Lyapunov function \(V=\tfrac{1}{2}(J\omega^2+m\|\mathbf{v}_C\|^2)\) is used to show asymptotic stopping under that law. The technical report emphasizes nonlinear observability conditions, consensus-based estimation, and Monte Carlo robustness rather than centralized fusion [1602.01891]. In this sense, the load agent is simultaneously an actuator and a distributed observer.

Aerial manipulation extends the idea to fully decentralized control of a cable-suspended payload. Each MAV agent observes the load pose, goal pose, its own state, and a one-hot identifier, but does not observe other MAVs and does not communicate with them; coordination is implicit through the shared load [2508.01522]. The action space is six-dimensional, consisting of desired linear acceleration and body rates. Real-world experiments with three MAVs report position RMSE \(0.52\) m and attitude RMSE \(22.93^\circ\), compared with \(0.45\) m and \(16.24^\circ\) for a centralized NMPC baseline [2508.01522]. The same policy class is also shown to tolerate heterogeneous controllers and complete in-flight loss of one MAV. This suggests that, for some tightly coupled physical systems, the load itself can function as a communication channel: agents infer the team state from the object’s motion rather than from explicit message passing.

## 5. Cognitive, passenger, and human-state load agents

A distinct branch of the literature uses the term for systems that estimate, rather than redistribute, load. NeuraDock is an EEG-based visual cognitive-load agent that reads 7-channel EEG at 250 Hz, preprocesses it, applies quality control, extracts posterior Alpha features in the 8–12 Hz band, and exposes a real-time workload index through a dashboard and HTTP API [2606.26518]. Its workflow is quality-gated: downstream Alpha and workload metrics are computed only after preprocessing and QC gating. In the tutorial mini-dataset, the agent processed 18 recordings, produced 10 within-subject comparisons, observed task-related posterior Alpha suppression in 7 of 10 contrasts, and reported baseline repeatability of Pearson \(r=0.803\) and ICC(C,1)\(=0.765\) for median posterior log Alpha [2606.26518]. Local replay benchmarks reported median core processing latency \(1.89\) ms and median HTTP endpoint latency \(15.15\) ms [2606.26518]. The agent therefore couples state estimation with explicit confidence reporting.

GazeMind performs an analogous role for smart glasses. It converts gaze into Temporal Gaze Encoding, augments it with Task-Guidance Reasoning, Adaptive User Profile Calibration, and Cognitive Retrieval-Augmented Generation, and predicts \(y\in\{\text{low},\text{moderate},\text{high}\}\) along with a textual rationale [2605.05790]. The associated CogLoad-Bench contains 152 participants, 40+ hours of multimodal data, and 10K+ real-time annotations. On the cross-user evaluation split, GazeMind reaches \(62.73\%\) accuracy and \(62.11\%\) macro-F1, compared with \(39.62\%\) and \(38.98\%\) for zero-shot GPT-4o and \(37.96\%\) and \(35.48\%\) for the best supervised baseline LSTM [2605.05790]. The ablation shows cumulative gains from TGR, AUP, and CogRAG, culminating in the full framework. Here the load agent is not a controller in the usual systems sense but an interpretable reasoning system over physiological signals.

Transit load estimation pushes this pattern into operational state estimation. The state-centric multi-agent framework models passenger occupancy after stop \(k\) as \(L_k\), uses a Perception module \(\mathcal{R}\) to predict \((\hat{B}_k,\hat{A}_k)\), a Physical module \(\mathcal{P}_C\) to enforce \(A_k^*=\min(\hat{A}_k,L_{k-1})\), \(B_k^*=\min(\hat{B}_k,C-(L_{k-1}-A_k^*))\), and \(L_k^{phys}=L_{k-1}-A_k^*+B_k^*\), and then fuses this with a Wi-Fi-derived anchor \(y_k^{load}=\rho_{h(k)}w_k\) using a trust weight \(\alpha_k\) derived from disagreement and physics residuals [2605.19834]. On all test trips, the proposed rule-based trust fusion without macro-shift achieves RMSE \(9.13 \pm 0.90\), MAE \(7.04 \pm 0.62\), and trip-end absolute error \(6.96 \pm 0.68\), compared with RMSE \(20.57 \pm 5.29\) and MAE \(13.72 \pm 3.37\) for Perception-only open loop [2605.19834]. On APC-inconsistent trips, RMSE drops from \(37.67 \pm 5.58\) for Perception-only to \(10.57 \pm 2.25\) for the proposed rule-based fusion [2605.19834]. In this domain, the load agent is a closed-loop state estimator that enforces feasibility and dynamically allocates trust among heterogeneous evidence sources.

## 6. Recurrent architectural themes, controversies, and limits

Several architectural motifs recur. First, many systems implement decentralized execution with either centralized training or no central coordination at all. Robust-MA3C uses centralized critics with decentralized actors [2303.08003]; DADC uses distributed critics to preserve privacy [2110.02784]; fog offloading is fully independent [2405.12236]; data-center load balancing avoids CTDE overhead through a Markov potential game [2206.01451]. Second, many papers separate proposal from feasibility. The transit estimator projects \((\hat{B}_k,\hat{A}_k)\) onto physically admissible flows [2605.19834]; the distribution-restoration framework uses invalid action masking to ensure zero constraint violations while also shrinking the feasible action space [2306.14018]; NeuraDock withholds workload metrics until QC passes [2606.26518]. Third, many load agents do not rebalance by migrating already-running work. dRAP balances at assignment time, with no explicit task migration [1509.06420]; the task-allocation framework with load management encourages idling and penalizes unnecessary reassignment rather than reactive reallocation after overload [2207.08279].

The literature also shows that “load agent” should not be reduced to computational balancing alone. In task allocation, the HTLM Dec-POMDP makes idling a first-class action and uses \(R_{LM}=R_{II}(b(s))+T_{RP}(a_0,a)\) to preserve reserve capacity; with moderate idle incentive, unused capability rises from about \(2.5\%\) to about \(71\%\) while completion steps remain almost unchanged [2207.08279]. In LLM systems, CoThinker reframes overload through Cognitive Load Theory and distributes intrinsic load via specialization, transactive memory, and a small-world communication graph; with default \(M=6\), \(T=3\), \(N=3\), and \(\beta=0.3\), it improves normalized LiveBench averages over IO, CoT, MAD, DMAD, and Self-Refine on high-load tasks, while also showing that extra agents can hurt low-intrinsic-load instruction-following tasks [2506.06843]. A plausible implication is that the most stable definition of a load agent is not “an entity that balances servers,” but “an entity whose policy is organized around a load-bearing latent variable and a feasibility-aware control loop.”

Common limits are equally consistent. dRAP still pays queue-traversal overhead and lacks preemptive migration [1509.06420]. Fog MARL assumes synchronous observation intervals and does not provide convergence guarantees for independent DDQL [2405.12236]. Robust-MA3C provides empirical convergence but no formal proof of Nash-equilibrium convergence [2303.08003]. The distributed load-restoration framework depends on centralized environment evaluation during training, and its gains hinge on action masking [2306.14018]. NeuraDock is explicitly within-subject and non-diagnostic [2606.26518], GazeMind depends on self-reported cognitive-load labels and task-specific rule synthesis [2605.05790], and the transit macro-correction module can over-correct and is therefore left off by default in the reported experiments [2605.19834]. These limits are not incidental: they reflect a common tension between local autonomy, partial observability, feasibility guarantees, and the cost of maintaining trustworthy load estimates or load-balancing actions.

Source: https://www.emergentmind.com/topics/load-agent