---
title: 'CrazyMARL: Multi-Agent RL for Aerial Robotics'
url: https://www.emergentmind.com/topics/crazymarl
type: topic
---

# CrazyMARL: Multi-Agent RL for Aerial Robotics

Searching arXiv for CrazyMARL-related papers to ground the article in current literature.
CrazyMARL is the name used for two closely related multi-agent reinforcement learning systems for cooperative aerial robotics, both implemented on Crazyflie platforms but targeting different control problems. In one usage, CrazyMARL denotes a Proximal Policy Optimization-based framework for collision-free, centimeter-accurate landing of drone swarms on static and moving platforms [2406.04159]. In another, it denotes a decentralized direct motor-control policy for cooperative transport of cable-suspended payloads, explicitly modeling slack–taut cable transitions and emphasizing zero-shot sim-to-real transfer under disturbances [2509.14126]. Across both usages, the common thread is decentralized execution with shared policy structure, deployment on resource-constrained Crazyflie hardware, and an attempt to replace or reduce reliance on analytical centralized control.

## 1. Terminological scope and research context

The term CrazyMARL does not designate a single invariant algorithmic artifact; rather, it identifies two research systems built around Crazyflie quadrotors and multi-agent RL.

| Paper | Task | Core setup |
|---|---|---|
| "MARLander: A Local Path Planning for Drone Swarms using Multiagent Deep Reinforcement Learning" [2406.04159] | Precise landing of a drone swarm at relocated target locations | PPO, CTDE, parameter sharing, continuous velocity commands |
| "CrazyMARL: Decentralized Direct Motor Control Policies for Cooperative Aerial Transport of Cable-Suspended Payloads" [2509.14126] | Cooperative transport of cable-suspended payloads | decentralized parameter-shared IPPO under CTDE, direct normalized motor commands |

The first system is presented as a local path-planning and landing method for drone swarms trained in a realistic simulated environment and deployed with Crazyflie drones and a Vicon indoor localization system [2406.04159]. The second system is formulated as a Dec-POMDP for teams of quadrotors transporting a point-mass payload via massless cables of fixed rest-length $L$, with the dynamics simulated in MuJoCo using tendons [2509.14126].

This suggests that CrazyMARL is best understood as a research lineage or naming convention rather than a single standardized framework. A plausible implication is that the name identifies a family of decentralized or CTDE MARL controllers for aerial multi-robot coordination on Crazyflie-class platforms.

## 2. Swarm landing formulation on Crazyflie hardware

In the landing setting, CrazyMARL is built atop Proximal Policy Optimization in a centralized-training, decentralized-execution paradigm with parameter sharing [2406.04159]. A single neural network hosts both an actor head, representing the policy $\pi_\theta$, and a critic head, representing the value function $V_\theta$, trained jointly. During training, the centralized critic receives the concatenated observations of all $N$ agents, while at execution each agent queries only its own policy. No explicit message passing is required at runtime; coordination is described as emerging through shared rewards and the centralized critic.

The shared actor-critic architecture takes as input a flattened vector of length $4N$ for the experimental case $N=2$, comprising each agent’s local state. The hidden layers are fully connected with ReLU activations and sizes $512$, $256$, and $128$. The actor head predicts $N\times 3$ continuous velocity commands, $\{u_x^i,u_y^i,u_z^i\}_{i=1\ldots N}$, while the critic head emits a scalar value estimate $V_\theta(s)$ [2406.04159].

Each Crazyflie agent $i$ at time $t$ observes
$$
O_t^i=[p_t^i,q_t^i,v_t^i,\omega_t^i],
$$
where $p_t^i\in\mathbb{R}^3$ is the 3D position of drone $i$ relative to its assigned landing target, $q_t^i\in\mathbb{R}^4$ is the unit quaternion orientation, $v_t^i\in\mathbb{R}^3$ is the linear velocity in the inertial frame, and $\omega_t^i\in\mathbb{R}^3$ is the angular velocity. Neighbor information is implicitly encoded through the centralized critic, which sees all $O_t^j$ for $j\neq i$; no additional message passing occurs during execution. The action for each agent is a continuous velocity set-point,
$$
A_t^i=[u_{x,t}^i,u_{y,t}^i,u_{z,t}^i]\in[-3,3]^3 \text{ (m/s)}.
$$

The simulation environment is a custom multi-agent Gym-style environment in PyBullet. The workspace is a cube $[-2,2]^3$ m, with agents and platform spawned uniformly at random at each episode start. Maximum commanded velocity is clipped at $3$ m/s. No explicit sensor or dynamics noise is injected, but the PyBullet solver is described as replicating aerodynamic drag and ground effects. In moving-platform episodes, the UR10-equipped platform follows a linear trajectory with speed sampled in $[0.2,0.5]$ m/s [2406.04159].

## 3. Reward design, optimization, and coordination in the landing system

The landing system uses a composite per-step reward designed to encourage precise, smooth, and collision-free landings [2406.04159]:
$$
R_t=\sum_{i=1}^N r_t^i+r_t^c+K\cdot I_{\mathrm{all\_landed}},
$$
with
$$
r_t^i=\alpha(\|p_{t-1}^i\|_2-\|p_t^i\|_2)-\beta\|v_t^i\|_2
$$
for each agent, and
$$
r_t^c=
\begin{cases}
-\alpha_c, & \text{if }\exists(i,j)\text{ s.t. }\|p_t^i-p_t^j\|_2<d_{\min} \\
0, & \text{otherwise.}
\end{cases}
$$
The terminal bonus $K>0$ is awarded only when all $N$ agents are simultaneously within a small radius $\epsilon$ of their targets. The distance-difference term shapes approach versus retreat, the velocity penalty favors smooth descents, and the hard negative term discourages inter-drone collisions.

Training is reported for $20$ M time-steps using PPO with discount factor $\gamma=0.99$, GAE $\lambda=0.95$, learning rate $3\times 10^{-4}$, clip range $\epsilon=0.2$, value loss coefficient $c_1=0.5$, entropy coefficient $c_2=0.01$, batch size $64$, and $10$ epochs per update. The combined PPO loss at update $k$ is given as
$$
L(\theta_k)=E_t\Big[
-\min(r_t(\theta_k)\cdot A_t,\;\mathrm{clip}(r_t(\theta_k),1-\epsilon,1+\epsilon)\cdot A_t)
+ c_1(V_\theta(s_t)-R_t^{\mathrm{adv}})^2
- c_2 H(\pi_\theta(\cdot|s_t))
\Big],
$$
where
$$
r_t(\theta)=\frac{\pi_\theta(a_t|s_t)}{\pi_{\theta_{\mathrm{old}}}(a_t|s_t)}.
$$

Coordination is attributed to shared network parameters and the centralized critic during training rather than to explicit inter-agent communication. Because agents do not exchange real-time messages, inference is stated to scale in $O(N)$. The only quadratic term in computation is within the reward penalty for collision checks, which can be pruned via locality by considering only near neighbors [2406.04159]. This suggests a particular conception of decentralization: coordination is learned implicitly through training-time access to joint state and a shared reward, while deployment remains communication-free.

## 4. Cooperative cable-suspended payload transport

In the payload-transport setting, CrazyMARL is formulated for $Q$ quadrotors jointly transporting a point-mass payload $P$ of mass $m_P$ via massless cables of fixed rest-length $L$ [2509.14126]. The state space at time $t$ is
$$
s_t=\{p_t^i\in\mathbb{R}^3,v_t^i\in\mathbb{R}^3,R_t^i\in SO(3),\omega_t^i\in\mathbb{R}^3\}_{i=1\ldots Q}\cup\{p_t^P\in\mathbb{R}^3,v_t^P\in\mathbb{R}^3\}.
$$
Each agent $i$ takes action $a_t^i\in[-1,1]^4$ based on its local observation $o_t^i$, and the agents share a common reward $r(s_t,a_t)$ and discount $\gamma$.

The cable dynamics are expressed through a hybrid slack–taut law. Cable length is
$$
\ell_t^i=\|p_t^i-p_t^P\|,
$$
and tension $T_t^i$ along
$$
u_t^i=\frac{p_t^i-p_t^P}{\ell_t^i}
$$
obeys
$$
T_t^i=\max\{0,\;k_c(\ell_t^i-L)\}.
$$
Thus $\ell_t^i>L$ implies a taut cable, while $\ell_t^i\leq L$ implies slack and $T_t^i=0$. The payload equations of motion are
$$
m_P\ddot x_t^P=\sum_{i=1}^Q T_t^i u_t^i + m_P g + F_t^{P,\mathrm{ext}},
$$
and each UAV satisfies
$$
m_i\ddot x_t^i=R_t^i f_t^i - T_t^i u_t^i + m_i g + F_t^{i,\mathrm{ext}}.
$$
The paper notes that exact kinematics are handled by MuJoCo’s physics engine, but that these expressions capture the slack/taut switch.

The RL framework is described as decentralized, parameter-shared IPPO under a CTDE paradigm, with no runtime communication required. The local observation for agent $i$ is
$$
o_t^i=[
e_t^P=p^P_{\mathrm{des},t}-p_t^P,\;
v_t^P,\;
\delta_t^i=p_t^i-p_t^P,\;
\mathrm{vec}(R_t^i),\;
v_t^i,\;
\omega_t^i,\;
a_{t-1}^i,\;
\{\delta_t^j\}_{j\neq i}
].
$$
Actions are mapped from $a_t^i\in[-1,1]^4$ to normalized thrusts $u_t^i=(a_t^i+1)/2\in[0,1]^4$, then to commanded thrust $f_{\mathrm{cmd}}=u_t^i\cdot f_{\max}$. A first-order motor lag is included per rotor:
$$
\nu_t^{i,j}=\sqrt{f_{\mathrm{cmd}}^{i,j}},
$$
$$
\tilde\nu_{t+1}^{i,j}=\tilde\nu_t^{i,j}+\alpha^i(\nu_t^{i,j}-\tilde\nu_t^{i,j}),\qquad \alpha^i=\Delta t/\tau^i,
$$
$$
f_t^{i,j}=\mathrm{clip}((\tilde\nu_{t+1}^{i,j})^2,0,f_{\max}).
$$

The actor is an MLP with layers $[64,64,64]$ and tanh activations, with Gaussian mean head and learned log-standard deviation. The critic is an MLP $[128,128,128]$ with tanh and scalar output $V(o)$. Total trainable parameters are reported as on the order of $20$–$30$ k for the actor and about $100$ k for the critic, with exact counts depending on input dimensions [2509.14126].

## 5. Learning objective, randomization strategy, and robustness in transport

The transport system uses Independent Proximal Policy Optimization with shared parameters [2509.14126]. For each agent,
$$
L^{\mathrm{CLIP}}(\theta)=E_t[\min(r_t(\theta)\hat A_t,\mathrm{clip}(r_t(\theta),1-\epsilon,1+\epsilon)\hat A_t)],
$$
where
$$
r_t(\theta)=\frac{\pi_\theta(a_t|o_t)}{\pi_{\theta_{\mathrm{old}}}(a_t|o_t)},
$$
with $\epsilon=0.2$. The value loss is
$$
L^V(\phi)=\tfrac12 E_t[(V_\phi(o_t)-R_t)^2],
$$
and an entropy bonus $c_H E_t[H(\pi_\theta(\cdot|o_t))]$ is included. Updates use $8$ PPO epochs per batch, $256$ minibatches, and GAE with $\lambda=0.95$.

The reward is factorized as
$$
r=r_{\mathrm{track}}\,r_{\mathrm{stable}}+r_{\mathrm{safe}}.
$$
The tracking term is
$$
r_{\mathrm{track}}=\tfrac12(r_{\mathrm{pos}}+r_{\mathrm{dir}}),
$$
with
$$
r_{\mathrm{pos}}=\exp(-s\|e^P\|),\qquad s=2,
$$
and
$$
r_{\mathrm{dir}}=\exp[-s_{\mathrm{align}}(1-v_{\mathrm{dir}}\cdot e_{\mathrm{dir}})],
$$
where
$$
s_{\mathrm{align}}=\min(c_g\|e^P\|,c_s),\qquad c_g=40,\; c_s=2.
$$

The stability term is
$$
r_{\mathrm{stable}}=\frac15\big[r_{\mathrm{velP}}+r_{\mathrm{velQ}}+\lambda_{\mathrm{yaw}}r_{\mathrm{yaw}}+\lambda_{\mathrm{up}}r_{\mathrm{up}}+r_{\mathrm{taut}}\big],
$$
with
$$
r_{\mathrm{velP}}=\exp\Big[-\Big(\frac{\|v^P\|}{c_{\mathrm{swing}}v_{\max}g(d)}\Big)^{c_{\exp}}\Big],\qquad c_{\mathrm{swing}}=0.75,\; c_{\exp}=8,\; v_{\max}=1.5\ \mathrm{m/s},
$$
$$
r_{\mathrm{velQ}}=\frac1Q\sum_i \exp\Big[-\Big(\frac{\|v^i\|}{v_{\max}g(d)}\Big)^{c_{\exp}}\Big],
$$
$$
r_{\mathrm{yaw}}=\frac1Q\sum_i \exp(-s|\omega_z^i|),\qquad \lambda_{\mathrm{yaw}}=10,
$$
$$
r_{\mathrm{up}}=\frac1Q\sum_i \exp[-s\theta^i],\qquad \lambda_{\mathrm{up}}=5,
$$
and
$$
r_{\mathrm{taut}}=\frac1L\Big[\frac1Q\sum_i \|p^i-p^P\|+\frac1Q\sum_i (p_z^i-p_z^P)\Big].
$$

The safety term is
$$
r_{\mathrm{safe}}=\frac15[-r_{\mathrm{coll}}-r_{\mathrm{oob}}-\lambda_s r_{\mathrm{smooth}}-r_{\mathrm{energy}}+r_{\mathrm{dist}}],
$$
where $r_{\mathrm{coll}}=c_{\mathrm{coll}}I_{\mathrm{collision}}$ with $c_{\mathrm{coll}}=10$, $r_{\mathrm{oob}}=c_{\mathrm{oob}}I_{\mathrm{oob}}$ with $c_{\mathrm{oob}}=10$, and
$$
r_{\mathrm{smooth}}=\frac12\Big[
\frac1Q\sum_i \|a_t^i-a_{t-1}^i\|_1 + \frac1Q\sum_i \|a_t^i-\bar a_t^i\mathbf{1}\|_1
\Big],
$$
with
$$
\bar a_t^i=\frac14\sum_j a_{t,j}^i,\qquad \lambda_s=10.
$$
The energy term is
$$
r_{\mathrm{energy}}=\frac1Q\frac14\sum_{i,j}[e^{-c_b|a_{t,j}^i|}+e^{c_b(a_{t,j}^i-1)}],\qquad c_b=50,
$$
and for $Q>1$, $r_{\mathrm{dist}}$ is the average over $i\neq j$ of
$$
\mathrm{clip}\Big(\frac{\|p^i-p^j\|-d_{\min}}{d_{\mathrm{safe}}-d_{\min}},0,1\Big),
$$
with $d_{\min}=0.15$ m and $d_{\mathrm{safe}}=0.18$ m.

Training uses MuJoCo at $250$ Hz with $\Delta t=0.004$ s and MuJoCo tendons modeling hybrid cable slack–taut transitions. Domain randomization includes randomized initial states; actuator model randomization with per-quad base thrust sampled from $U(0.105,0.15)$ N plus $\mathcal N(0,0.008^2)$ N and clipped to $[0.095,0.16]$ N; $\tau^i\sim U(0.004,0.05)$ s; observation noise with $o'=o+\sigma_{\mathrm{obs}}\Lambda\eta$, $\sigma_{\mathrm{obs}}=1.0$; and disturbances comprising random forces and torques on UAVs, random forces on the payload, and occasional RPM jumps. RL hyperparameters include $16{,}384$ environments, rollout length $128$, total environment steps $2\times 10^9$, learning rate $4\times 10^{-4}$, gradient-norm clip $0.5$, entropy coefficient $0.01$, value-loss coefficient $0.5$, discount $\gamma=0.997$, and episode length $3072$ steps, approximately $12.3$ s [2509.14126].

## 6. Empirical results, transfer, and limitations

For the landing application, CrazyMARL was deployed on two Bitcraze Crazyflie 2.1 drones in a Vicon motion-capture arena [2406.04159]. Vicon updates at $100$ Hz provide position and orientation, while velocities are estimated via a first-order filter. Control commands are throttled at $50$ Hz and clipped to $\pm 3$ m/s. Safety thresholds trigger an emergency hover when horizontal speed exceeds $4$ m/s or altitude falls below $0$. Landing pads are described as $\varnothing 0.4$ m cylinders on the UR10 TCP and are instrumented with AprilTags for fail-safe vision fallback.

Quantitatively, the landing system is evaluated across static-platform trials $(N=12)$ and moving-platform trials $(N=8)$. For static platforms, CrazyMARL reports mean Euclidean landing error $\mu=2.26$ cm, standard deviation $\sigma=0.14$ cm, and success rate $91.7\%$, whereas the PID+APF baseline reports $\mu=4.80$ cm, $\sigma=0.30$ cm, and success rate $80.0\%$. For moving platforms, CrazyMARL reports $\mu=3.93$ cm, $\sigma=0.20$ cm, and success rate $75.0\%$, while PID+APF reports $\mu=7.40$ cm, $\sigma=0.45$ cm, and success rate $60.0\%$. A two-sample $t$-test on static errors yields $t(20)=5.12$, $p<0.001$ [2406.04159]. The source text summarizes these results by stating that CrazyMARL halves landing error and raises success rates by approximately $15\%$ over PID+APF while maintaining sub-$4$ cm accuracy on moving targets.

For payload transport, the main baseline comparison is a two-UAV setting with $1{,}000$ trials and harsh initializations [2509.14126]. CrazyMARL reports success $797/1000 = 79.7\%$ and mean payload speed $0.58$ m/s, while a rigid-rod centralized baseline reports success $435/1000 = 43.5\%$ and speed $0.27$ m/s. The source states that the RL policy recovers more reliably and quickly, whereas the baseline often spirals under swing. In generalization sweeps with two UAVs, the paper reports a sharp drop in success for cable lengths $L<0.3$ m, improved success for moderate increases in the range $0.3$–$1$ m, degradation for $L>1$ m, a robust plateau over payload mass variation with moderate degradation for extreme light or heavy payloads, tolerance to observation noise up to approximately $1$, and insensitivity to random seed with all runs within $\pm 2\%$.

Scalability results in the transport setting report $99\%$ success and approximately $2$ s settling for $Q=1$, $81\%$ success for $Q=2$, $60\%$ success for $Q=3$ with an added $20$ g payload and coordinated load-sharing, and near $0\%$ for $Q=6$ with an added $40$ g payload, attributed to breakdown due to peer-ordering sensitivity [2509.14126]. Zero-shot sim-to-real transfer is demonstrated on Crazyflie 2.1 hardware with STM32F405, onboard EKF and motion capture, a TFLite policy running at $250$ Hz, and direct PWM output. Reported trials include single-UAV flight with and without payload, two-UAV cooperative transport, and operation in $3.5$ m/s average wind, with outcomes including stable auto-takeoff, rapid disturbance rejection under pushes and wind, and payload figure-8 tracking under gusts.

The two CrazyMARL systems converge on several common implications. Both emphasize decentralized execution without runtime communication, both rely on shared policy structure, and both use simulation environments intended to preserve transferability to real Crazyflie hardware. At the same time, they expose distinct limitations. In the landing case, the source argues for extensibility from two agents to tens of Crazyflies through parameter sharing and linear inference cost [2406.04159]; this is a stated implication rather than an experimentally demonstrated scaling result. In the transport case, the source explicitly identifies permutation-sensitive peer-state encoding as a current limitation and points to order-invariant or attention-based peer encoders, centralized critics such as MAPPO for larger teams, and integrated obstacle avoidance for cluttered environments as promising extensions [2509.14126].

Source: https://www.emergentmind.com/topics/crazymarl