Papers
Topics
Authors
Recent
Search
2000 character limit reached

DiffMAS: Dual Systems for Control & Communication

Updated 10 July 2026
  • DiffMAS refers to two distinct systems: FO-DiffMAS-2D for fractional-order crop-dusting control and a multi-agent framework for optimizing latent communication.
  • FO-DiffMAS-2D integrates fractional PDE simulation, CVT-based actuator placement, and adaptive sensor scheduling to achieve rapid pest infestation reduction.
  • In the multi-agent framework, replacing text-based messaging with a differentiable Key–Value cache enables end-to-end optimization and improved benchmark performance.

DiffMAS is a label used in the arXiv record for two distinct systems. In "Multi-UAV-based Optimal Crop-dusting of Anomalously Diffusing Infestation of Crops" (Cao et al., 2014), FO-DiffMAS-2D is introduced as a simulation platform for measurement scheduling and controls in fractional order distributed parameter systems, motivated by real-time pest management with networked unmanned cropdusters. In "Learning to Communicate: Toward End-to-End Optimization of Multi-Agent Language Systems" (Yu et al., 23 Apr 2026), DiffMAS denotes a training framework that treats latent communication as a learnable component of multi-agent systems by replacing text-based message passing with a differentiable Key–Value cache latent trace.

1. Terminological scope

The two arXiv usages of DiffMAS differ in domain, mathematical formalism, and system objective.

Usage Domain Core mechanism
FO-DiffMAS-2D Multi-UAV crop-dusting and anomalous diffusion control Fractional PDE simulation, CVT-based actuator placement, measurement scheduling
DiffMAS Multi-agent language systems End-to-end optimization of latent communication through KV-cache trajectories

FO-DiffMAS-2D appears in the 2014 crop-dusting work (Cao et al., 2014). The 2026 language-systems paper uses DiffMAS for a framework in which agents sequentially build and consume a continuous latent trace rather than exchange text messages (Yu et al., 23 Apr 2026). A separate acronym, MAS, is also used in "Measurement-aligned Flow for Inverse Problem" (Zhang et al., 13 Jun 2025) for Measurement-Aligned Sampling, a diffusion-model-based solver for linear inverse problems; that method is distinct from both DiffMAS usages.

2. FO-DiffMAS-2D: fractional-order distributed sensing and actuation

FO-DiffMAS-2D is implemented in MATLAB/Simulink and integrates in time the discretized fractional PDE, either time-fractional or space-fractional, over a rectangular 2D mesh (Cao et al., 2014). Spatial derivatives are discretized by standard finite differences for integer-order terms and by the fractional-central-difference scheme for Riesz derivatives. The platform contains three principal classes of entities: static sensors at mesh points that measure local pest density ρ(x,y,t)\rho(x,y,t) at each PDE time-step; mobile actuators, modeled as UAV crop-dusters; and disturbance sources represented as localized point terms fd(x,y,t)f_d(x,y,t) that inject pests at moving or fixed locations.

The actuator dynamics follow a second-order damped control law,

p¨i=− kp(pi−pˉi(t))−kv p˙i,\ddot p_i = -\,k_p\bigl(p_i-\bar p_i(t)\bigr)-k_v\,\dot p_i,

where pi∈R2p_i\in\mathbb R^2 is the current UAV position and pˉi\bar p_i is the density-weighted centroid of the current Voronoi cell. The environment is a domain Ω⊂R2\Omega\subset\mathbb R^2, usually a unit square, with either Dirichlet or Neumann boundary conditions:

ρ∣∂Ω=C\rho|_{\partial\Omega}=C

or

∂ρ∂n=C1+C2ρ.\frac{\partial \rho}{\partial n}=C_1+C_2\rho.

The mesh is typically uniform; the summary gives 29×\times29 sensors as a representative configuration.

Two anomalous diffusion models govern the infestation dynamics. The time-fractional model is

CD0,tα ρ(x,y,t)=kα(ρxx+ρyy)+fd(ρ,x,y,t)+fc(ρ~,x,y,t),0<α≤1,_{C}D_{0,t}^{\alpha}\,\rho(x,y,t) = k_{\alpha}\bigl(\rho_{xx}+\rho_{yy}\bigr) + f_d\bigl(\rho,x,y,t\bigr) + f_c\bigl(\tilde\rho,x,y,t\bigr), \qquad 0<\alpha\le1,

with initial condition fd(x,y,t)f_d(x,y,t)0 and the Caputo derivative

fd(x,y,t)f_d(x,y,t)1

The space-fractional model is

fd(x,y,t)f_d(x,y,t)2

with the Riesz derivative defined through left and right Riemann–Liouville derivatives and the coefficient

fd(x,y,t)f_d(x,y,t)3

The communication layer is adaptive. Each actuator has a communication or observation radius fd(x,y,t)f_d(x,y,t)4. At every CVT-update step, the actuator polls all sensors within fd(x,y,t)f_d(x,y,t)5, constructs its local Voronoi cell fd(x,y,t)f_d(x,y,t)6, and adapts fd(x,y,t)f_d(x,y,t)7 until all owned sensors lie inside fd(x,y,t)f_d(x,y,t)8. Actuators recompute centroids and issue motion commands every fd(x,y,t)f_d(x,y,t)9 seconds, with p¨i=− kp(pi−pˉi(t))−kv p˙i,\ddot p_i = -\,k_p\bigl(p_i-\bar p_i(t)\bigr)-k_v\,\dot p_i,0 given as an example.

3. CVT formulation, scheduling loop, and control metrics in FO-DiffMAS-2D

The crop-dusting formulation uses Centroidal Voronoi Tessellations to compute optimal dynamic actuator locations (Cao et al., 2014). For UAV positions p¨i=− kp(pi−pˉi(t))−kv p˙i,\ddot p_i = -\,k_p\bigl(p_i-\bar p_i(t)\bigr)-k_v\,\dot p_i,1 and Voronoi cells p¨i=− kp(pi−pˉi(t))−kv p˙i,\ddot p_i = -\,k_p\bigl(p_i-\bar p_i(t)\bigr)-k_v\,\dot p_i,2, the density-weighted coverage cost is

p¨i=− kp(pi−pˉi(t))−kv p˙i,\ddot p_i = -\,k_p\bigl(p_i-\bar p_i(t)\bigr)-k_v\,\dot p_i,3

Minimizing this functional yields a CVT, with necessary condition

p¨i=− kp(pi−pˉi(t))−kv p˙i,\ddot p_i = -\,k_p\bigl(p_i-\bar p_i(t)\bigr)-k_v\,\dot p_i,4

The discrete solution uses Lloyd’s algorithm, also described as the MacQueen probabilistic method: initialize actuator positions, compute Voronoi cells, calculate density-weighted centroids from sampled sensor values, test centroid-position discrepancies against a tolerance, and iterate until convergence.

The runtime loop operates at every Simulink time-step p¨i=− kp(pi−pˉi(t))−kv p˙i,\ddot p_i = -\,k_p\bigl(p_i-\bar p_i(t)\bigr)-k_v\,\dot p_i,5. Each sensor measures p¨i=− kp(pi−pˉi(t))−kv p˙i,\ddot p_i = -\,k_p\bigl(p_i-\bar p_i(t)\bigr)-k_v\,\dot p_i,6 at its grid node; each actuator collects measurements from sensors within its current radius; every p¨i=− kp(pi−pˉi(t))−kv p˙i,\ddot p_i = -\,k_p\bigl(p_i-\bar p_i(t)\bigr)-k_v\,\dot p_i,7 seconds the actuator performs the discrete Lloyd update to obtain p¨i=− kp(pi−pˉi(t))−kv p˙i,\ddot p_i = -\,k_p\bigl(p_i-\bar p_i(t)\bigr)-k_v\,\dot p_i,8 and applies the motion law p¨i=− kp(pi−pˉi(t))−kv p˙i,\ddot p_i = -\,k_p\bigl(p_i-\bar p_i(t)\bigr)-k_v\,\dot p_i,9; and simultaneously the control input pi∈R2p_i\in\mathbb R^20 is applied to the PDE at the UAV’s current location, often implemented as a localized sink in the finite-difference stencil.

Radius adaptation follows the distributed rule attributed to Chen et al. (2007): start with a small pi∈R2p_i\in\mathbb R^21, detect neighbors inside pi∈R2p_i\in\mathbb R^22, form a preliminary pi∈R2p_i\in\mathbb R^23, compute

pi∈R2p_i\in\mathbb R^24

and if pi∈R2p_i\in\mathbb R^25 fix pi∈R2p_i\in\mathbb R^26; otherwise set pi∈R2p_i\in\mathbb R^27 and repeat. Performance is summarized by three metrics: the global infestation level

pi∈R2p_i\in\mathbb R^28

the confinement area defined as a measure of the support of pi∈R2p_i\in\mathbb R^29 above a threshold, and coverage efficiency measured by the rate of decrease pˉi\bar p_i0 and the final reduction pˉi\bar p_i1.

4. FO-DiffMAS-2D simulation results and the role of fractional order

The reported simulations examine both time-fractional and space-fractional infestations (Cao et al., 2014). In the time-fractional example, the domain is pˉi\bar p_i2 with homogeneous Neumann boundaries and a point source pˉi\bar p_i3 at pˉi\bar p_i4. The setup uses 29pˉi\bar p_i529 sensors, 4 UAVs with second-order dynamics, simulation interval pˉi\bar p_i6, time step pˉi\bar p_i7, and update interval pˉi\bar p_i8. The order parameter is varied over pˉi\bar p_i9.

In that experiment, the sum Ω⊂R2\Omega\subset\mathbb R^20 exhibits subdiffusive flattening as Ω⊂R2\Omega\subset\mathbb R^21 decreases, and the best worst-case peak reduction occurs at Ω⊂R2\Omega\subset\mathbb R^22. The summary also states that for Ω⊂R2\Omega\subset\mathbb R^23 the uncontrolled PDE becomes divergent. Under closed-loop control with Ω⊂R2\Omega\subset\mathbb R^24 and Ω⊂R2\Omega\subset\mathbb R^25, the final infestation satisfies Ω⊂R2\Omega\subset\mathbb R^26, with monotonic decrease.

The space-fractional example uses the same geometry, a Dirichlet boundary Ω⊂R2\Omega\subset\mathbb R^27, and a point source at Ω⊂R2\Omega\subset\mathbb R^28, while varying Ω⊂R2\Omega\subset\mathbb R^29. Optimal control performance, characterized as fastest suppression and lowest peak, occurs at ρ∣∂Ω=C\rho|_{\partial\Omega}=C0. The experiments are summarized as showing that fractional orders ρ∣∂Ω=C\rho|_{\partial\Omega}=C1 or ρ∣∂Ω=C\rho|_{\partial\Omega}=C2 can better model anomalous subdiffusion, that FO-DiffMAS-2D integrates CVT placement, measurement scheduling, and fractional-order PDE simulation in real time, and that actuator placement with simple local control laws achieves over 80% infestation reduction in seconds.

5. DiffMAS for end-to-end optimization of multi-agent latent communication

The 2026 DiffMAS framework addresses a different problem: multi-agent reasoning with LLMs (Yu et al., 23 Apr 2026). It replaces text-based message passing in a ρ∣∂Ω=C\rho|_{\partial\Omega}=C3-stage multi-agent pipeline with a continuous, differentiable Key–Value cache latent trace that agents sequentially build and consume. Each agent ρ∣∂Ω=C\rho|_{\partial\Omega}=C4, equipped with prompt ρ∣∂Ω=C\rho|_{\partial\Omega}=C5, takes as input the problem ρ∣∂Ω=C\rho|_{\partial\Omega}=C6 and the accumulated trace ρ∣∂Ω=C\rho|_{\partial\Omega}=C7, emits ρ∣∂Ω=C\rho|_{\partial\Omega}=C8 new latent blocks, and appends them to the trace. Only the final agent performs autoregressive token decoding conditioned on the full trace, so gradients can flow back through all stages.

The latent block space is ρ∣∂Ω=C\rho|_{\partial\Omega}=C9, with each block ∂ρ∂n=C1+C2ρ.\frac{\partial \rho}{\partial n}=C_1+C_2\rho.0 and ∂ρ∂n=C1+C2ρ.\frac{\partial \rho}{\partial n}=C_1+C_2\rho.1 after ∂ρ∂n=C1+C2ρ.\frac{\partial \rho}{\partial n}=C_1+C_2\rho.2 agents. The stage operator is written as

∂ρ∂n=C1+C2ρ.\frac{\partial \rho}{\partial n}=C_1+C_2\rho.3

and the macro-view composes these operators across all stages. Within each stage, the micro-dynamics use an internal state ∂ρ∂n=C1+C2ρ.\frac{\partial \rho}{\partial n}=C_1+C_2\rho.4:

∂ρ∂n=C1+C2ρ.\frac{\partial \rho}{\partial n}=C_1+C_2\rho.5

∂ρ∂n=C1+C2ρ.\frac{\partial \rho}{\partial n}=C_1+C_2\rho.6

∂ρ∂n=C1+C2ρ.\frac{\partial \rho}{\partial n}=C_1+C_2\rho.7

All of ∂ρ∂n=C1+C2ρ.\frac{\partial \rho}{\partial n}=C_1+C_2\rho.8, ∂ρ∂n=C1+C2ρ.\frac{\partial \rho}{\partial n}=C_1+C_2\rho.9, and ×\times0 are implemented by a shared pretrained transformer with stage-specific prompts.

After ×\times1 stages, the final agent decodes

×\times2

and training minimizes the supervised cross-entropy loss

×\times3

Because ×\times4 is constructed by differentiable operators, ×\times5 flows through all stages and micro-steps.

The parameter-efficient fine-tuning protocol freezes the backbone transformer and injects LoRA adapters into query, key, value, and projection matrices of the final agent’s cross-attention, and in some tasks into all agents. The LoRA hyperparameters are rank ×\times6, scaling ×\times7, and dropout ×\times8. Optimization uses AdamW with cosine annealing, linear warmup ×\times9, learning rate CD0,tα ρ(x,y,t)=kα(ρxx+ρyy)+fd(ρ,x,y,t)+fc(ρ~,x,y,t),0<α≤1,_{C}D_{0,t}^{\alpha}\,\rho(x,y,t) = k_{\alpha}\bigl(\rho_{xx}+\rho_{yy}\bigr) + f_d\bigl(\rho,x,y,t\bigr) + f_c\bigl(\tilde\rho,x,y,t\bigr), \qquad 0<\alpha\le1,0 on LoRA weights only, gradient clipping norm CD0,tα ρ(x,y,t)=kα(ρxx+ρyy)+fd(ρ,x,y,t)+fc(ρ~,x,y,t),0<α≤1,_{C}D_{0,t}^{\alpha}\,\rho(x,y,t) = k_{\alpha}\bigl(\rho_{xx}+\rho_{yy}\bigr) + f_d\bigl(\rho,x,y,t\bigr) + f_c\bigl(\tilde\rho,x,y,t\bigr), \qquad 0<\alpha\le1,1, and 64 micro-batches accumulation. Stage I, comprising agents CD0,tα ρ(x,y,t)=kα(ρxx+ρyy)+fd(ρ,x,y,t)+fc(ρ~,x,y,t),0<α≤1,_{C}D_{0,t}^{\alpha}\,\rho(x,y,t) = k_{\alpha}\bigl(\rho_{xx}+\rho_{yy}\bigr) + f_d\bigl(\rho,x,y,t\bigr) + f_c\bigl(\tilde\rho,x,y,t\bigr), \qquad 0<\alpha\le1,2, runs in inference mode with no gradient; Stage II uses teacher-forced decoding on agent CD0,tα ρ(x,y,t)=kα(ρxx+ρyy)+fd(ρ,x,y,t)+fc(ρ~,x,y,t),0<α≤1,_{C}D_{0,t}^{\alpha}\,\rho(x,y,t) = k_{\alpha}\bigl(\rho_{xx}+\rho_{yy}\bigr) + f_d\bigl(\rho,x,y,t\bigr) + f_c\bigl(\tilde\rho,x,y,t\bigr), \qquad 0<\alpha\le1,3 and updates the LoRA parameters.

6. Benchmarks, ablations, and disambiguation of the 2026 DiffMAS

The 2026 experiments cover mathematical reasoning, scientific QA, code generation, and commonsense benchmarks (Yu et al., 23 Apr 2026). The tasks and datasets are AIME 2024, AIME 2025, GPQA-Diamond, OpenBookQA, HumanEval-Plus, and MBPP-Plus. The evaluated models include Qwen3 (4B, 8B, 14B), Mistral3-8B, and DeepSeek-R1-Distill-Qwen-32B. Baselines are Single, TextMAS, LatentMAS, and C2C. LoRA fine-tuning uses 210 math samples from Hendrycks Math, 50 code samples from HumanEval, and 700 commonsense samples from CommonsenseQA with correct reasoning traces generated by Gemini 3 Flash. Inference uses temperature CD0,tα ρ(x,y,t)=kα(ρxx+ρyy)+fd(ρ,x,y,t)+fc(ρ~,x,y,t),0<α≤1,_{C}D_{0,t}^{\alpha}\,\rho(x,y,t) = k_{\alpha}\bigl(\rho_{xx}+\rho_{yy}\bigr) + f_d\bigl(\rho,x,y,t\bigr) + f_c\bigl(\tilde\rho,x,y,t\bigr), \qquad 0<\alpha\le1,4, top-CD0,tα ρ(x,y,t)=kα(ρxx+ρyy)+fd(ρ,x,y,t)+fc(ρ~,x,y,t),0<α≤1,_{C}D_{0,t}^{\alpha}\,\rho(x,y,t) = k_{\alpha}\bigl(\rho_{xx}+\rho_{yy}\bigr) + f_d\bigl(\rho,x,y,t\bigr) + f_c\bigl(\tilde\rho,x,y,t\bigr), \qquad 0<\alpha\le1,5, and task-specific maximum lengths.

The principal Qwen3-8B results are reported as follows.

Task Single DiffMAS
AIME24 50.0 76.7
AIME25 46.7 56.7
GPQA 39.9 60.1
HumanEval+ 74.5 81.5
MBPP+ 64.8 74.8
OpenBookQA 83.6 85.8

The same table gives explicit gains over Single of CD0,tα ρ(x,y,t)=kα(ρxx+ρyy)+fd(ρ,x,y,t)+fc(ρ~,x,y,t),0<α≤1,_{C}D_{0,t}^{\alpha}\,\rho(x,y,t) = k_{\alpha}\bigl(\rho_{xx}+\rho_{yy}\bigr) + f_d\bigl(\rho,x,y,t\bigr) + f_c\bigl(\tilde\rho,x,y,t\bigr), \qquad 0<\alpha\le1,6 on AIME24, CD0,tα ρ(x,y,t)=kα(ρxx+ρyy)+fd(ρ,x,y,t)+fc(ρ~,x,y,t),0<α≤1,_{C}D_{0,t}^{\alpha}\,\rho(x,y,t) = k_{\alpha}\bigl(\rho_{xx}+\rho_{yy}\bigr) + f_d\bigl(\rho,x,y,t\bigr) + f_c\bigl(\tilde\rho,x,y,t\bigr), \qquad 0<\alpha\le1,7 on AIME25, CD0,tα ρ(x,y,t)=kα(ρxx+ρyy)+fd(ρ,x,y,t)+fc(ρ~,x,y,t),0<α≤1,_{C}D_{0,t}^{\alpha}\,\rho(x,y,t) = k_{\alpha}\bigl(\rho_{xx}+\rho_{yy}\bigr) + f_d\bigl(\rho,x,y,t\bigr) + f_c\bigl(\tilde\rho,x,y,t\bigr), \qquad 0<\alpha\le1,8 on GPQA, CD0,tα ρ(x,y,t)=kα(ρxx+ρyy)+fd(ρ,x,y,t)+fc(ρ~,x,y,t),0<α≤1,_{C}D_{0,t}^{\alpha}\,\rho(x,y,t) = k_{\alpha}\bigl(\rho_{xx}+\rho_{yy}\bigr) + f_d\bigl(\rho,x,y,t\bigr) + f_c\bigl(\tilde\rho,x,y,t\bigr), \qquad 0<\alpha\le1,9 on HumanEval+, fd(x,y,t)f_d(x,y,t)00 on MBPP+, and fd(x,y,t)f_d(x,y,t)01 on OpenBookQA. A second table states that large-scale models—Mistral3-8B, Qwen3-14B, and DeepSeek-32B—all show consistent 2–10 point gains for DiffMAS over the best baseline.

Decoding stability is measured on AIME 2024 with Qwen3-4B. DiffMAS attains mean perplexity fd(x,y,t)f_d(x,y,t)02 versus fd(x,y,t)f_d(x,y,t)03 for LatentMAS, with outlier perplexity greater than fd(x,y,t)f_d(x,y,t)04 virtually eliminated. Under self-consistency with four stochastic samples per problem, DiffMAS yields mostly 3–4 correct per instance, whereas LatentMAS and TextMAS concentrate at 0–1. Token-level entropy of the final agent shows smooth entropy growth for DiffMAS and large spikes for LatentMAS, which the paper interprets as downstream uncertainty from untrained KV injection.

The ablations emphasize that communication is beneficial but not monotone in trace length. On AIME 2024 with Qwen3-8B, accuracy across latent-block counts is reported as fd(x,y,t)f_d(x,y,t)05 for fd(x,y,t)f_d(x,y,t)06 blocks, fd(x,y,t)f_d(x,y,t)07 for fd(x,y,t)f_d(x,y,t)08, fd(x,y,t)f_d(x,y,t)09 for fd(x,y,t)f_d(x,y,t)10, fd(x,y,t)f_d(x,y,t)11 for fd(x,y,t)f_d(x,y,t)12, fd(x,y,t)f_d(x,y,t)13 for fd(x,y,t)f_d(x,y,t)14, and fd(x,y,t)f_d(x,y,t)15 for fd(x,y,t)f_d(x,y,t)16, leading to the stated conclusion that a small number of latent blocks, specifically fd(x,y,t)f_d(x,y,t)17, suffices while too many blocks introduce noise. Further comparisons show DiffMAS and TextMAS+SFT tied at fd(x,y,t)f_d(x,y,t)18 on AIME24, with DiffMAS higher on AIME25, GPQA-Diamond, and OpenBookQA; and DiffMAS higher than StitchMAS on GPQA-Diamond, AIME24, and AIME25. The qualitative case study characterizes TextMAS as stable but lossy, LatentMAS as high-capacity but chaotic, and DiffMAS as high-capacity plus stable; only DiffMAS yields a correct, coherent AIME 2024 solution in that example.

A recurrent source of confusion is acronym overlap. The 2026 DiffMAS framework is a latent-communication training method for multi-agent language systems, whereas FO-DiffMAS-2D is a fractional-order simulation and control platform for crop-dusting, and MAS in Measurement-Aligned Sampling is a diffusion-model-based method for inverse problems (Zhang et al., 13 Jun 2025). The arXiv record therefore uses closely related names for different technical objects. A plausible implication is that citations to DiffMAS benefit from explicit paper titles or arXiv identifiers when disambiguation is important.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to DiffMAS.