---
title: 'DiffMAS: Dual Systems for Control & Communication'
url: https://www.emergentmind.com/topics/diffmas
type: topic
---

# DiffMAS: Dual Systems for Control & Communication

DiffMAS is a label used in the arXiv record for two distinct systems. In "Multi-UAV-based Optimal Crop-dusting of Anomalously Diffusing Infestation of Crops" [1411.2880], FO-DiffMAS-2D is introduced as a simulation platform for measurement scheduling and controls in fractional order distributed parameter systems, motivated by real-time pest management with networked unmanned cropdusters. In "Learning to Communicate: Toward End-to-End Optimization of Multi-Agent Language Systems" [2604.21794], DiffMAS denotes a training framework that treats latent communication as a learnable component of multi-agent systems by replacing text-based message passing with a differentiable Key–Value cache latent trace.

## 1. Terminological scope

The two arXiv usages of DiffMAS differ in domain, mathematical formalism, and system objective.

| Usage | Domain | Core mechanism |
|---|---|---|
| FO-DiffMAS-2D | Multi-UAV crop-dusting and anomalous diffusion control | Fractional PDE simulation, CVT-based actuator placement, measurement scheduling |
| DiffMAS | Multi-agent language systems | End-to-end optimization of latent communication through KV-cache trajectories |

FO-DiffMAS-2D appears in the 2014 crop-dusting work [1411.2880]. The 2026 language-systems paper uses DiffMAS for a framework in which agents sequentially build and consume a continuous latent trace rather than exchange text messages [2604.21794]. A separate acronym, MAS, is also used in "Measurement-aligned Flow for Inverse Problem" [2506.11893] for Measurement-Aligned Sampling, a diffusion-model-based solver for linear inverse problems; that method is distinct from both DiffMAS usages.

## 2. FO-DiffMAS-2D: fractional-order distributed sensing and actuation

FO-DiffMAS-2D is implemented in MATLAB/Simulink and integrates in time the discretized fractional PDE, either time-fractional or space-fractional, over a rectangular 2D mesh [1411.2880]. Spatial derivatives are discretized by standard finite differences for integer-order terms and by the fractional-central-difference scheme for Riesz derivatives. The platform contains three principal classes of entities: static sensors at mesh points that measure local pest density $\rho(x,y,t)$ at each PDE time-step; mobile actuators, modeled as UAV crop-dusters; and disturbance sources represented as localized point terms $f_d(x,y,t)$ that inject pests at moving or fixed locations.

The actuator dynamics follow a second-order damped control law,
$$
\ddot p_i = -\,k_p\bigl(p_i-\bar p_i(t)\bigr)-k_v\,\dot p_i,
$$
where $p_i\in\mathbb R^2$ is the current UAV position and $\bar p_i$ is the density-weighted centroid of the current Voronoi cell. The environment is a domain $\Omega\subset\mathbb R^2$, usually a unit square, with either Dirichlet or Neumann boundary conditions:
$$
\rho|_{\partial\Omega}=C
$$
or
$$
\frac{\partial \rho}{\partial n}=C_1+C_2\rho.
$$
The mesh is typically uniform; the summary gives 29$\times$29 sensors as a representative configuration.

Two anomalous diffusion models govern the infestation dynamics. The time-fractional model is
$$
_{C}D_{0,t}^{\alpha}\,\rho(x,y,t)
=
k_{\alpha}\bigl(\rho_{xx}+\rho_{yy}\bigr)
+
f_d\bigl(\rho,x,y,t\bigr)
+
f_c\bigl(\tilde\rho,x,y,t\bigr),
\qquad 0<\alpha\le1,
$$
with initial condition $\rho(x,y,0)=\rho_0(x,y)$ and the Caputo derivative
$$
_{C}D_{0,t}^{\alpha}\rho
=
\frac{1}{\Gamma(1-\alpha)}
\int_{0}^{t}(t-\tau)^{-\alpha}
\frac{\partial\rho(x,y,\tau)}{\partial\tau}\,d\tau.
$$
The space-fractional model is
$$
\rho_t
=
k_{\beta}\left(
\frac{\partial^{\beta}\rho}{\partial|x|^{\beta}}
+
\frac{\partial^{\beta}\rho}{\partial|y|^{\beta}}
\right)
+
f_d(\rho,x,y,t)
+
f_c(\tilde\rho,x,y,t),
\qquad 1<\beta\le2,
$$
with the Riesz derivative defined through left and right Riemann–Liouville derivatives and the coefficient
$$
c_{\beta}=\frac1{2\cos(\pi\beta/2)}.
$$

The communication layer is adaptive. Each actuator has a communication or observation radius $R_i(t)$. At every CVT-update step, the actuator polls all sensors within $R_i$, constructs its local Voronoi cell $V_i$, and adapts $R_i$ until all owned sensors lie inside $V_i$. Actuators recompute centroids and issue motion commands every $T_c$ seconds, with $T_c=0.1\,\mathrm s$ given as an example.

## 3. CVT formulation, scheduling loop, and control metrics in FO-DiffMAS-2D

The crop-dusting formulation uses Centroidal Voronoi Tessellations to compute optimal dynamic actuator locations [1411.2880]. For UAV positions $P=\{p_i\}_{i=1}^n$ and Voronoi cells $\{V_i\}$, the density-weighted coverage cost is
$$
\mathcal K(P,\{V_i\})
=
\sum_{i=1}^{n}
\int_{V_i}\rho(q)\,\|q-p_i\|^2\,dq.
$$
Minimizing this functional yields a CVT, with necessary condition
$$
p_i=\bar p_i=
\frac{\int_{V_i}q\,\rho(q)\,dq}
{\int_{V_i}\rho(q)\,dq},
\qquad i=1,\dots,n.
$$
The discrete solution uses Lloyd’s algorithm, also described as the MacQueen probabilistic method: initialize actuator positions, compute Voronoi cells, calculate density-weighted centroids from sampled sensor values, test centroid-position discrepancies against a tolerance, and iterate until convergence.

The runtime loop operates at every Simulink time-step $\Delta t$. Each sensor measures $\rho_j$ at its grid node; each actuator collects measurements from sensors within its current radius; every $T_c$ seconds the actuator performs the discrete Lloyd update to obtain $\bar p_i$ and applies the motion law $\ddot p_i=-k_p(p_i-\bar p_i)-k_v\dot p_i$; and simultaneously the control input $f_c(\tilde\rho,x,y,t)$ is applied to the PDE at the UAV’s current location, often implemented as a localized sink in the finite-difference stencil.

Radius adaptation follows the distributed rule attributed to Chen et al. (2007): start with a small $R_i$, detect neighbors inside $R_i$, form a preliminary $V_i$, compute
$$
d_i=\max_{q_j\in V_i}\|q_j-p_i\|,
$$
and if $R_i>2d_i$ fix $R_i$; otherwise set $R_i\leftarrow 2R_i$ and repeat. Performance is summarized by three metrics: the global infestation level
$$
I(t)=\sum_j \rho_j(t),
$$
the confinement area defined as a measure of the support of $\rho$ above a threshold, and coverage efficiency measured by the rate of decrease $-dI/dt$ and the final reduction $I(t_f)/\max I$.

## 4. FO-DiffMAS-2D simulation results and the role of fractional order

The reported simulations examine both time-fractional and space-fractional infestations [1411.2880]. In the time-fractional example, the domain is $(0,1)^2$ with homogeneous Neumann boundaries and a point source $f_d(t)=20e^{-t}$ at $(0.8,0.2)$. The setup uses 29$\times$29 sensors, 4 UAVs with second-order dynamics, simulation interval $t\in[0,6]\,\mathrm s$, time step $\Delta t=0.002\,\mathrm s$, and update interval $T_c=0.1\,\mathrm s$. The order parameter is varied over $\alpha\in\{0.7,0.8,0.9,1.0\}$.

In that experiment, the sum $I(t)$ exhibits subdiffusive flattening as $\alpha$ decreases, and the best worst-case peak reduction occurs at $\alpha=0.7$. The summary also states that for $\alpha<0.7$ the uncontrolled PDE becomes divergent. Under closed-loop control with $k_p=6$ and $k_v=1$, the final infestation satisfies $I(t_f)\approx 0.19\,\max I$, with monotonic decrease.

The space-fractional example uses the same geometry, a Dirichlet boundary $\rho=0$, and a point source at $(0.75,0.35)$, while varying $\beta\in\{1.5,1.6,1.7,1.9,2.0\}$. Optimal control performance, characterized as fastest suppression and lowest peak, occurs at $\beta=1.7$. The experiments are summarized as showing that fractional orders $\alpha<1$ or $\beta<2$ can better model anomalous subdiffusion, that FO-DiffMAS-2D integrates CVT placement, measurement scheduling, and fractional-order PDE simulation in real time, and that actuator placement with simple local control laws achieves over 80% infestation reduction in seconds.

## 5. DiffMAS for end-to-end optimization of multi-agent latent communication

The 2026 DiffMAS framework addresses a different problem: multi-agent reasoning with large language models [2604.21794]. It replaces text-based message passing in a $K$-stage multi-agent pipeline with a continuous, differentiable Key–Value cache latent trace that agents sequentially build and consume. Each agent $j$, equipped with prompt $p_j$, takes as input the problem $x$ and the accumulated trace $\mathbf Z_{1:N_{j-1}}$, emits $T$ new latent blocks, and appends them to the trace. Only the final agent performs autoregressive token decoding conditioned on the full trace, so gradients can flow back through all stages.

The latent block space is $\mathcal Z\subset\mathbb R^d$, with each block $z_t\in\mathbb R^d$ and $N_j=jT$ after $j$ agents. The stage operator is written as
$$
A^{(j)}_\theta:\mathcal T_{j-1}\to\mathcal T_j,
\qquad
\mathbf Z_{1:N_j}=A^{(j)}_\theta\bigl(\mathbf Z_{1:N_{j-1}},x,p_j\bigr),
$$
and the macro-view composes these operators across all stages. Within each stage, the micro-dynamics use an internal state $s_t^{(j)}\in\mathbb R^m$:
$$
s_0^{(j)}=\eta_\theta\bigl(x,p_j,\mathbf Z_{1:N_{j-1}}\bigr),
$$
$$
z_t^{(j)}=
g_\theta\bigl(
s_{t-1}^{(j)},
x,p_j,
\mathbf Z_{1:N_{j-1}},
\mathbf Z^{(j)}_{1:t-1}
\bigr),
$$
$$
s_t^{(j)}=
s_{t-1}^{(j)}+
f_\theta\bigl(
s_{t-1}^{(j)},
x,p_j,
\mathbf Z_{1:N_{j-1}},
\mathbf Z^{(j)}_{1:t}
\bigr).
$$
All of $\eta_\theta$, $g_\theta$, and $f_\theta$ are implemented by a shared pretrained transformer with stage-specific prompts.

After $K$ stages, the final agent decodes
$$
p_\theta(y\mid x,\{p_j\})=\mathrm{Dec}_\theta\bigl(x,p_K,\mathbf Z_{1:N_K}\bigr),
$$
and training minimizes the supervised cross-entropy loss
$$
\mathcal L(\theta)
=
-\sum_{t=1}^{T}
\log p_\theta\bigl(
y_t^\star\mid x,y_{<t}^\star,\{p_j\}
\bigr).
$$
Because $\mathbf Z_{1:N_K}$ is constructed by differentiable operators, $\nabla_\theta \mathcal L$ flows through all stages and micro-steps.

The parameter-efficient fine-tuning protocol freezes the backbone transformer and injects LoRA adapters into query, key, value, and projection matrices of the final agent’s cross-attention, and in some tasks into all agents. The LoRA hyperparameters are rank $r=8$, scaling $\alpha=16$, and dropout $0.05$. Optimization uses AdamW with cosine annealing, linear warmup $3\%$, learning rate $5\times 10^{-5}$ on LoRA weights only, gradient clipping norm $1.0$, and 64 micro-batches accumulation. Stage I, comprising agents $1\dots K-1$, runs in inference mode with no gradient; Stage II uses teacher-forced decoding on agent $K$ and updates the LoRA parameters.

## 6. Benchmarks, ablations, and disambiguation of the 2026 DiffMAS

The 2026 experiments cover mathematical reasoning, scientific QA, code generation, and commonsense benchmarks [2604.21794]. The tasks and datasets are AIME 2024, AIME 2025, GPQA-Diamond, OpenBookQA, HumanEval-Plus, and MBPP-Plus. The evaluated models include Qwen3 (4B, 8B, 14B), Mistral3-8B, and DeepSeek-R1-Distill-Qwen-32B. Baselines are Single, TextMAS, LatentMAS, and C2C. LoRA fine-tuning uses 210 math samples from Hendrycks Math, 50 code samples from HumanEval, and 700 commonsense samples from CommonsenseQA with correct reasoning traces generated by Gemini 3 Flash. Inference uses temperature $0.6$, top-$p=0.95$, and task-specific maximum lengths.

The principal Qwen3-8B results are reported as follows.

| Task | Single | DiffMAS |
|---|---:|---:|
| AIME24 | 50.0 | 76.7 |
| AIME25 | 46.7 | 56.7 |
| GPQA | 39.9 | 60.1 |
| HumanEval+ | 74.5 | 81.5 |
| MBPP+ | 64.8 | 74.8 |
| OpenBookQA | 83.6 | 85.8 |

The same table gives explicit gains over Single of $+26.7$ on AIME24, $+10.0$ on AIME25, $+20.2$ on GPQA, $+7.0$ on HumanEval+, $+10.0$ on MBPP+, and $+2.2$ on OpenBookQA. A second table states that large-scale models—Mistral3-8B, Qwen3-14B, and DeepSeek-32B—all show consistent 2–10 point gains for DiffMAS over the best baseline.

Decoding stability is measured on AIME 2024 with Qwen3-4B. DiffMAS attains mean perplexity $1.24$ versus $1.31$ for LatentMAS, with outlier perplexity greater than $2.0$ virtually eliminated. Under self-consistency with four stochastic samples per problem, DiffMAS yields mostly 3–4 correct per instance, whereas LatentMAS and TextMAS concentrate at 0–1. Token-level entropy of the final agent shows smooth entropy growth for DiffMAS and large spikes for LatentMAS, which the paper interprets as downstream uncertainty from untrained KV injection.

The ablations emphasize that communication is beneficial but not monotone in trace length. On AIME 2024 with Qwen3-8B, accuracy across latent-block counts is reported as $50$ for $0$ blocks, $76.7$ for $10$, $63.3$ for $40$, $73.3$ for $100$, $66.7$ for $150$, and $63.3$ for $200$, leading to the stated conclusion that a small number of latent blocks, specifically $10$, suffices while too many blocks introduce noise. Further comparisons show DiffMAS and TextMAS+SFT tied at $76.7\%$ on AIME24, with DiffMAS higher on AIME25, GPQA-Diamond, and OpenBookQA; and DiffMAS higher than StitchMAS on GPQA-Diamond, AIME24, and AIME25. The qualitative case study characterizes TextMAS as stable but lossy, LatentMAS as high-capacity but chaotic, and DiffMAS as high-capacity plus stable; only DiffMAS yields a correct, coherent AIME 2024 solution in that example.

A recurrent source of confusion is acronym overlap. The 2026 DiffMAS framework is a latent-communication training method for multi-agent language systems, whereas FO-DiffMAS-2D is a fractional-order simulation and control platform for crop-dusting, and MAS in Measurement-Aligned Sampling is a diffusion-model-based method for inverse problems [2506.11893]. The arXiv record therefore uses closely related names for different technical objects. A plausible implication is that citations to DiffMAS benefit from explicit paper titles or arXiv identifiers when disambiguation is important.

Source: https://www.emergentmind.com/topics/diffmas