CrazyMARL: Multi-Agent RL for Aerial Robotics
- CrazyMARL is a family of decentralized multi-agent reinforcement learning systems designed for cooperative aerial robotics using shared policy structures and CTDE strategies.
- One variant employs a PPO-based framework for collision-free, centimeter-accurate landing of drone swarms on static and moving platforms.
- Another variant uses decentralized direct motor-control policies to enable cooperative transport of cable-suspended payloads with improved recovery over traditional baselines.
Searching arXiv for CrazyMARL-related papers to ground the article in current literature. CrazyMARL is the name used for two closely related multi-agent reinforcement learning systems for cooperative aerial robotics, both implemented on Crazyflie platforms but targeting different control problems. In one usage, CrazyMARL denotes a Proximal Policy Optimization-based framework for collision-free, centimeter-accurate landing of drone swarms on static and moving platforms (Aschu et al., 2024). In another, it denotes a decentralized direct motor-control policy for cooperative transport of cable-suspended payloads, explicitly modeling slack–taut cable transitions and emphasizing zero-shot sim-to-real transfer under disturbances (Lorentz et al., 17 Sep 2025). Across both usages, the common thread is decentralized execution with shared policy structure, deployment on resource-constrained Crazyflie hardware, and an attempt to replace or reduce reliance on analytical centralized control.
1. Terminological scope and research context
The term CrazyMARL does not designate a single invariant algorithmic artifact; rather, it identifies two research systems built around Crazyflie quadrotors and multi-agent RL.
| Paper | Task | Core setup |
|---|---|---|
| "MARLander: A Local Path Planning for Drone Swarms using Multiagent Deep Reinforcement Learning" (Aschu et al., 2024) | Precise landing of a drone swarm at relocated target locations | PPO, CTDE, parameter sharing, continuous velocity commands |
| "CrazyMARL: Decentralized Direct Motor Control Policies for Cooperative Aerial Transport of Cable-Suspended Payloads" (Lorentz et al., 17 Sep 2025) | Cooperative transport of cable-suspended payloads | decentralized parameter-shared IPPO under CTDE, direct normalized motor commands |
The first system is presented as a local path-planning and landing method for drone swarms trained in a realistic simulated environment and deployed with Crazyflie drones and a Vicon indoor localization system (Aschu et al., 2024). The second system is formulated as a Dec-POMDP for teams of quadrotors transporting a point-mass payload via massless cables of fixed rest-length , with the dynamics simulated in MuJoCo using tendons (Lorentz et al., 17 Sep 2025).
This suggests that CrazyMARL is best understood as a research lineage or naming convention rather than a single standardized framework. A plausible implication is that the name identifies a family of decentralized or CTDE MARL controllers for aerial multi-robot coordination on Crazyflie-class platforms.
2. Swarm landing formulation on Crazyflie hardware
In the landing setting, CrazyMARL is built atop Proximal Policy Optimization in a centralized-training, decentralized-execution paradigm with parameter sharing (Aschu et al., 2024). A single neural network hosts both an actor head, representing the policy , and a critic head, representing the value function , trained jointly. During training, the centralized critic receives the concatenated observations of all agents, while at execution each agent queries only its own policy. No explicit message passing is required at runtime; coordination is described as emerging through shared rewards and the centralized critic.
The shared actor-critic architecture takes as input a flattened vector of length $4N$ for the experimental case , comprising each agent’s local state. The hidden layers are fully connected with ReLU activations and sizes $512$, $256$, and $128$. The actor head predicts continuous velocity commands, 0, while the critic head emits a scalar value estimate 1 (Aschu et al., 2024).
Each Crazyflie agent 2 at time 3 observes
4
where 5 is the 3D position of drone 6 relative to its assigned landing target, 7 is the unit quaternion orientation, 8 is the linear velocity in the inertial frame, and 9 is the angular velocity. Neighbor information is implicitly encoded through the centralized critic, which sees all 0 for 1; no additional message passing occurs during execution. The action for each agent is a continuous velocity set-point,
2
The simulation environment is a custom multi-agent Gym-style environment in PyBullet. The workspace is a cube 3 m, with agents and platform spawned uniformly at random at each episode start. Maximum commanded velocity is clipped at 4 m/s. No explicit sensor or dynamics noise is injected, but the PyBullet solver is described as replicating aerodynamic drag and ground effects. In moving-platform episodes, the UR10-equipped platform follows a linear trajectory with speed sampled in 5 m/s (Aschu et al., 2024).
3. Reward design, optimization, and coordination in the landing system
The landing system uses a composite per-step reward designed to encourage precise, smooth, and collision-free landings (Aschu et al., 2024):
6
with
7
for each agent, and
8
The terminal bonus 9 is awarded only when all 0 agents are simultaneously within a small radius 1 of their targets. The distance-difference term shapes approach versus retreat, the velocity penalty favors smooth descents, and the hard negative term discourages inter-drone collisions.
Training is reported for 2 M time-steps using PPO with discount factor 3, GAE 4, learning rate 5, clip range 6, value loss coefficient 7, entropy coefficient 8, batch size 9, and $4N$0 epochs per update. The combined PPO loss at update $4N$1 is given as
$4N$2
where
$4N$3
Coordination is attributed to shared network parameters and the centralized critic during training rather than to explicit inter-agent communication. Because agents do not exchange real-time messages, inference is stated to scale in $4N$4. The only quadratic term in computation is within the reward penalty for collision checks, which can be pruned via locality by considering only near neighbors (Aschu et al., 2024). This suggests a particular conception of decentralization: coordination is learned implicitly through training-time access to joint state and a shared reward, while deployment remains communication-free.
4. Cooperative cable-suspended payload transport
In the payload-transport setting, CrazyMARL is formulated for $4N$5 quadrotors jointly transporting a point-mass payload $4N$6 of mass $4N$7 via massless cables of fixed rest-length $4N$8 (Lorentz et al., 17 Sep 2025). The state space at time $4N$9 is
0
Each agent 1 takes action 2 based on its local observation 3, and the agents share a common reward 4 and discount 5.
The cable dynamics are expressed through a hybrid slack–taut law. Cable length is
6
and tension 7 along
8
obeys
9
Thus $512$0 implies a taut cable, while $512$1 implies slack and $512$2. The payload equations of motion are
$512$3
and each UAV satisfies
$512$4
The paper notes that exact kinematics are handled by MuJoCo’s physics engine, but that these expressions capture the slack/taut switch.
The RL framework is described as decentralized, parameter-shared IPPO under a CTDE paradigm, with no runtime communication required. The local observation for agent $512$5 is
$512$6
Actions are mapped from $512$7 to normalized thrusts $512$8, then to commanded thrust $512$9. A first-order motor lag is included per rotor:
$256$0
$256$1
$256$2
The actor is an MLP with layers $256$3 and tanh activations, with Gaussian mean head and learned log-standard deviation. The critic is an MLP $256$4 with tanh and scalar output $256$5. Total trainable parameters are reported as on the order of $256$6–$256$7 k for the actor and about $256$8 k for the critic, with exact counts depending on input dimensions (Lorentz et al., 17 Sep 2025).
5. Learning objective, randomization strategy, and robustness in transport
The transport system uses Independent Proximal Policy Optimization with shared parameters (Lorentz et al., 17 Sep 2025). For each agent,
$256$9
where
$128$0
with $128$1. The value loss is
$128$2
and an entropy bonus $128$3 is included. Updates use $128$4 PPO epochs per batch, $128$5 minibatches, and GAE with $128$6.
The reward is factorized as
$128$7
The tracking term is
$128$8
with
$128$9
and
0
where
1
The stability term is
2
with
3
4
5
6
and
7
The safety term is
8
where 9 with 00, 01 with 02, and
03
with
04
The energy term is
05
and for 06, 07 is the average over 08 of
09
with 10 m and 11 m.
Training uses MuJoCo at 12 Hz with 13 s and MuJoCo tendons modeling hybrid cable slack–taut transitions. Domain randomization includes randomized initial states; actuator model randomization with per-quad base thrust sampled from 14 N plus 15 N and clipped to 16 N; 17 s; observation noise with 18, 19; and disturbances comprising random forces and torques on UAVs, random forces on the payload, and occasional RPM jumps. RL hyperparameters include 20 environments, rollout length 21, total environment steps 22, learning rate 23, gradient-norm clip 24, entropy coefficient 25, value-loss coefficient 26, discount 27, and episode length 28 steps, approximately 29 s (Lorentz et al., 17 Sep 2025).
6. Empirical results, transfer, and limitations
For the landing application, CrazyMARL was deployed on two Bitcraze Crazyflie 2.1 drones in a Vicon motion-capture arena (Aschu et al., 2024). Vicon updates at 30 Hz provide position and orientation, while velocities are estimated via a first-order filter. Control commands are throttled at 31 Hz and clipped to 32 m/s. Safety thresholds trigger an emergency hover when horizontal speed exceeds 33 m/s or altitude falls below 34. Landing pads are described as 35 m cylinders on the UR10 TCP and are instrumented with AprilTags for fail-safe vision fallback.
Quantitatively, the landing system is evaluated across static-platform trials 36 and moving-platform trials 37. For static platforms, CrazyMARL reports mean Euclidean landing error 38 cm, standard deviation 39 cm, and success rate 40, whereas the PID+APF baseline reports 41 cm, 42 cm, and success rate 43. For moving platforms, CrazyMARL reports 44 cm, 45 cm, and success rate 46, while PID+APF reports 47 cm, 48 cm, and success rate 49. A two-sample 50-test on static errors yields 51, 52 (Aschu et al., 2024). The source text summarizes these results by stating that CrazyMARL halves landing error and raises success rates by approximately 53 over PID+APF while maintaining sub-54 cm accuracy on moving targets.
For payload transport, the main baseline comparison is a two-UAV setting with 55 trials and harsh initializations (Lorentz et al., 17 Sep 2025). CrazyMARL reports success 56 and mean payload speed 57 m/s, while a rigid-rod centralized baseline reports success 58 and speed 59 m/s. The source states that the RL policy recovers more reliably and quickly, whereas the baseline often spirals under swing. In generalization sweeps with two UAVs, the paper reports a sharp drop in success for cable lengths 60 m, improved success for moderate increases in the range 61–62 m, degradation for 63 m, a robust plateau over payload mass variation with moderate degradation for extreme light or heavy payloads, tolerance to observation noise up to approximately 64, and insensitivity to random seed with all runs within 65.
Scalability results in the transport setting report 66 success and approximately 67 s settling for 68, 69 success for 70, 71 success for 72 with an added 73 g payload and coordinated load-sharing, and near 74 for 75 with an added 76 g payload, attributed to breakdown due to peer-ordering sensitivity (Lorentz et al., 17 Sep 2025). Zero-shot sim-to-real transfer is demonstrated on Crazyflie 2.1 hardware with STM32F405, onboard EKF and motion capture, a TFLite policy running at 77 Hz, and direct PWM output. Reported trials include single-UAV flight with and without payload, two-UAV cooperative transport, and operation in 78 m/s average wind, with outcomes including stable auto-takeoff, rapid disturbance rejection under pushes and wind, and payload figure-8 tracking under gusts.
The two CrazyMARL systems converge on several common implications. Both emphasize decentralized execution without runtime communication, both rely on shared policy structure, and both use simulation environments intended to preserve transferability to real Crazyflie hardware. At the same time, they expose distinct limitations. In the landing case, the source argues for extensibility from two agents to tens of Crazyflies through parameter sharing and linear inference cost (Aschu et al., 2024); this is a stated implication rather than an experimentally demonstrated scaling result. In the transport case, the source explicitly identifies permutation-sensitive peer-state encoding as a current limitation and points to order-invariant or attention-based peer encoders, centralized critics such as MAPPO for larger teams, and integrated obstacle avoidance for cluttered environments as promising extensions (Lorentz et al., 17 Sep 2025).