Papers
Topics
Authors
Recent
Search
2000 character limit reached

CrazyMARL: Multi-Agent RL for Aerial Robotics

Updated 17 July 2026
  • CrazyMARL is a family of decentralized multi-agent reinforcement learning systems designed for cooperative aerial robotics using shared policy structures and CTDE strategies.
  • One variant employs a PPO-based framework for collision-free, centimeter-accurate landing of drone swarms on static and moving platforms.
  • Another variant uses decentralized direct motor-control policies to enable cooperative transport of cable-suspended payloads with improved recovery over traditional baselines.

Searching arXiv for CrazyMARL-related papers to ground the article in current literature. CrazyMARL is the name used for two closely related multi-agent reinforcement learning systems for cooperative aerial robotics, both implemented on Crazyflie platforms but targeting different control problems. In one usage, CrazyMARL denotes a Proximal Policy Optimization-based framework for collision-free, centimeter-accurate landing of drone swarms on static and moving platforms (Aschu et al., 2024). In another, it denotes a decentralized direct motor-control policy for cooperative transport of cable-suspended payloads, explicitly modeling slack–taut cable transitions and emphasizing zero-shot sim-to-real transfer under disturbances (Lorentz et al., 17 Sep 2025). Across both usages, the common thread is decentralized execution with shared policy structure, deployment on resource-constrained Crazyflie hardware, and an attempt to replace or reduce reliance on analytical centralized control.

1. Terminological scope and research context

The term CrazyMARL does not designate a single invariant algorithmic artifact; rather, it identifies two research systems built around Crazyflie quadrotors and multi-agent RL.

Paper Task Core setup
"MARLander: A Local Path Planning for Drone Swarms using Multiagent Deep Reinforcement Learning" (Aschu et al., 2024) Precise landing of a drone swarm at relocated target locations PPO, CTDE, parameter sharing, continuous velocity commands
"CrazyMARL: Decentralized Direct Motor Control Policies for Cooperative Aerial Transport of Cable-Suspended Payloads" (Lorentz et al., 17 Sep 2025) Cooperative transport of cable-suspended payloads decentralized parameter-shared IPPO under CTDE, direct normalized motor commands

The first system is presented as a local path-planning and landing method for drone swarms trained in a realistic simulated environment and deployed with Crazyflie drones and a Vicon indoor localization system (Aschu et al., 2024). The second system is formulated as a Dec-POMDP for teams of quadrotors transporting a point-mass payload via massless cables of fixed rest-length LL, with the dynamics simulated in MuJoCo using tendons (Lorentz et al., 17 Sep 2025).

This suggests that CrazyMARL is best understood as a research lineage or naming convention rather than a single standardized framework. A plausible implication is that the name identifies a family of decentralized or CTDE MARL controllers for aerial multi-robot coordination on Crazyflie-class platforms.

2. Swarm landing formulation on Crazyflie hardware

In the landing setting, CrazyMARL is built atop Proximal Policy Optimization in a centralized-training, decentralized-execution paradigm with parameter sharing (Aschu et al., 2024). A single neural network hosts both an actor head, representing the policy πθ\pi_\theta, and a critic head, representing the value function VθV_\theta, trained jointly. During training, the centralized critic receives the concatenated observations of all NN agents, while at execution each agent queries only its own policy. No explicit message passing is required at runtime; coordination is described as emerging through shared rewards and the centralized critic.

The shared actor-critic architecture takes as input a flattened vector of length $4N$ for the experimental case N=2N=2, comprising each agent’s local state. The hidden layers are fully connected with ReLU activations and sizes $512$, $256$, and $128$. The actor head predicts N×3N\times 3 continuous velocity commands, πθ\pi_\theta0, while the critic head emits a scalar value estimate πθ\pi_\theta1 (Aschu et al., 2024).

Each Crazyflie agent πθ\pi_\theta2 at time πθ\pi_\theta3 observes

πθ\pi_\theta4

where πθ\pi_\theta5 is the 3D position of drone πθ\pi_\theta6 relative to its assigned landing target, πθ\pi_\theta7 is the unit quaternion orientation, πθ\pi_\theta8 is the linear velocity in the inertial frame, and πθ\pi_\theta9 is the angular velocity. Neighbor information is implicitly encoded through the centralized critic, which sees all VθV_\theta0 for VθV_\theta1; no additional message passing occurs during execution. The action for each agent is a continuous velocity set-point,

VθV_\theta2

The simulation environment is a custom multi-agent Gym-style environment in PyBullet. The workspace is a cube VθV_\theta3 m, with agents and platform spawned uniformly at random at each episode start. Maximum commanded velocity is clipped at VθV_\theta4 m/s. No explicit sensor or dynamics noise is injected, but the PyBullet solver is described as replicating aerodynamic drag and ground effects. In moving-platform episodes, the UR10-equipped platform follows a linear trajectory with speed sampled in VθV_\theta5 m/s (Aschu et al., 2024).

3. Reward design, optimization, and coordination in the landing system

The landing system uses a composite per-step reward designed to encourage precise, smooth, and collision-free landings (Aschu et al., 2024):

VθV_\theta6

with

VθV_\theta7

for each agent, and

VθV_\theta8

The terminal bonus VθV_\theta9 is awarded only when all NN0 agents are simultaneously within a small radius NN1 of their targets. The distance-difference term shapes approach versus retreat, the velocity penalty favors smooth descents, and the hard negative term discourages inter-drone collisions.

Training is reported for NN2 M time-steps using PPO with discount factor NN3, GAE NN4, learning rate NN5, clip range NN6, value loss coefficient NN7, entropy coefficient NN8, batch size NN9, and $4N$0 epochs per update. The combined PPO loss at update $4N$1 is given as

$4N$2

where

$4N$3

Coordination is attributed to shared network parameters and the centralized critic during training rather than to explicit inter-agent communication. Because agents do not exchange real-time messages, inference is stated to scale in $4N$4. The only quadratic term in computation is within the reward penalty for collision checks, which can be pruned via locality by considering only near neighbors (Aschu et al., 2024). This suggests a particular conception of decentralization: coordination is learned implicitly through training-time access to joint state and a shared reward, while deployment remains communication-free.

4. Cooperative cable-suspended payload transport

In the payload-transport setting, CrazyMARL is formulated for $4N$5 quadrotors jointly transporting a point-mass payload $4N$6 of mass $4N$7 via massless cables of fixed rest-length $4N$8 (Lorentz et al., 17 Sep 2025). The state space at time $4N$9 is

N=2N=20

Each agent N=2N=21 takes action N=2N=22 based on its local observation N=2N=23, and the agents share a common reward N=2N=24 and discount N=2N=25.

The cable dynamics are expressed through a hybrid slack–taut law. Cable length is

N=2N=26

and tension N=2N=27 along

N=2N=28

obeys

N=2N=29

Thus $512$0 implies a taut cable, while $512$1 implies slack and $512$2. The payload equations of motion are

$512$3

and each UAV satisfies

$512$4

The paper notes that exact kinematics are handled by MuJoCo’s physics engine, but that these expressions capture the slack/taut switch.

The RL framework is described as decentralized, parameter-shared IPPO under a CTDE paradigm, with no runtime communication required. The local observation for agent $512$5 is

$512$6

Actions are mapped from $512$7 to normalized thrusts $512$8, then to commanded thrust $512$9. A first-order motor lag is included per rotor:

$256$0

$256$1

$256$2

The actor is an MLP with layers $256$3 and tanh activations, with Gaussian mean head and learned log-standard deviation. The critic is an MLP $256$4 with tanh and scalar output $256$5. Total trainable parameters are reported as on the order of $256$6–$256$7 k for the actor and about $256$8 k for the critic, with exact counts depending on input dimensions (Lorentz et al., 17 Sep 2025).

5. Learning objective, randomization strategy, and robustness in transport

The transport system uses Independent Proximal Policy Optimization with shared parameters (Lorentz et al., 17 Sep 2025). For each agent,

$256$9

where

$128$0

with $128$1. The value loss is

$128$2

and an entropy bonus $128$3 is included. Updates use $128$4 PPO epochs per batch, $128$5 minibatches, and GAE with $128$6.

The reward is factorized as

$128$7

The tracking term is

$128$8

with

$128$9

and

N×3N\times 30

where

N×3N\times 31

The stability term is

N×3N\times 32

with

N×3N\times 33

N×3N\times 34

N×3N\times 35

N×3N\times 36

and

N×3N\times 37

The safety term is

N×3N\times 38

where N×3N\times 39 with πθ\pi_\theta00, πθ\pi_\theta01 with πθ\pi_\theta02, and

πθ\pi_\theta03

with

πθ\pi_\theta04

The energy term is

πθ\pi_\theta05

and for πθ\pi_\theta06, πθ\pi_\theta07 is the average over πθ\pi_\theta08 of

πθ\pi_\theta09

with πθ\pi_\theta10 m and πθ\pi_\theta11 m.

Training uses MuJoCo at πθ\pi_\theta12 Hz with πθ\pi_\theta13 s and MuJoCo tendons modeling hybrid cable slack–taut transitions. Domain randomization includes randomized initial states; actuator model randomization with per-quad base thrust sampled from πθ\pi_\theta14 N plus πθ\pi_\theta15 N and clipped to πθ\pi_\theta16 N; πθ\pi_\theta17 s; observation noise with πθ\pi_\theta18, πθ\pi_\theta19; and disturbances comprising random forces and torques on UAVs, random forces on the payload, and occasional RPM jumps. RL hyperparameters include πθ\pi_\theta20 environments, rollout length πθ\pi_\theta21, total environment steps πθ\pi_\theta22, learning rate πθ\pi_\theta23, gradient-norm clip πθ\pi_\theta24, entropy coefficient πθ\pi_\theta25, value-loss coefficient πθ\pi_\theta26, discount πθ\pi_\theta27, and episode length πθ\pi_\theta28 steps, approximately πθ\pi_\theta29 s (Lorentz et al., 17 Sep 2025).

6. Empirical results, transfer, and limitations

For the landing application, CrazyMARL was deployed on two Bitcraze Crazyflie 2.1 drones in a Vicon motion-capture arena (Aschu et al., 2024). Vicon updates at πθ\pi_\theta30 Hz provide position and orientation, while velocities are estimated via a first-order filter. Control commands are throttled at πθ\pi_\theta31 Hz and clipped to πθ\pi_\theta32 m/s. Safety thresholds trigger an emergency hover when horizontal speed exceeds πθ\pi_\theta33 m/s or altitude falls below πθ\pi_\theta34. Landing pads are described as πθ\pi_\theta35 m cylinders on the UR10 TCP and are instrumented with AprilTags for fail-safe vision fallback.

Quantitatively, the landing system is evaluated across static-platform trials πθ\pi_\theta36 and moving-platform trials πθ\pi_\theta37. For static platforms, CrazyMARL reports mean Euclidean landing error πθ\pi_\theta38 cm, standard deviation πθ\pi_\theta39 cm, and success rate πθ\pi_\theta40, whereas the PID+APF baseline reports πθ\pi_\theta41 cm, πθ\pi_\theta42 cm, and success rate πθ\pi_\theta43. For moving platforms, CrazyMARL reports πθ\pi_\theta44 cm, πθ\pi_\theta45 cm, and success rate πθ\pi_\theta46, while PID+APF reports πθ\pi_\theta47 cm, πθ\pi_\theta48 cm, and success rate πθ\pi_\theta49. A two-sample πθ\pi_\theta50-test on static errors yields πθ\pi_\theta51, πθ\pi_\theta52 (Aschu et al., 2024). The source text summarizes these results by stating that CrazyMARL halves landing error and raises success rates by approximately πθ\pi_\theta53 over PID+APF while maintaining sub-πθ\pi_\theta54 cm accuracy on moving targets.

For payload transport, the main baseline comparison is a two-UAV setting with πθ\pi_\theta55 trials and harsh initializations (Lorentz et al., 17 Sep 2025). CrazyMARL reports success πθ\pi_\theta56 and mean payload speed πθ\pi_\theta57 m/s, while a rigid-rod centralized baseline reports success πθ\pi_\theta58 and speed πθ\pi_\theta59 m/s. The source states that the RL policy recovers more reliably and quickly, whereas the baseline often spirals under swing. In generalization sweeps with two UAVs, the paper reports a sharp drop in success for cable lengths πθ\pi_\theta60 m, improved success for moderate increases in the range πθ\pi_\theta61–πθ\pi_\theta62 m, degradation for πθ\pi_\theta63 m, a robust plateau over payload mass variation with moderate degradation for extreme light or heavy payloads, tolerance to observation noise up to approximately πθ\pi_\theta64, and insensitivity to random seed with all runs within πθ\pi_\theta65.

Scalability results in the transport setting report πθ\pi_\theta66 success and approximately πθ\pi_\theta67 s settling for πθ\pi_\theta68, πθ\pi_\theta69 success for πθ\pi_\theta70, πθ\pi_\theta71 success for πθ\pi_\theta72 with an added πθ\pi_\theta73 g payload and coordinated load-sharing, and near πθ\pi_\theta74 for πθ\pi_\theta75 with an added πθ\pi_\theta76 g payload, attributed to breakdown due to peer-ordering sensitivity (Lorentz et al., 17 Sep 2025). Zero-shot sim-to-real transfer is demonstrated on Crazyflie 2.1 hardware with STM32F405, onboard EKF and motion capture, a TFLite policy running at πθ\pi_\theta77 Hz, and direct PWM output. Reported trials include single-UAV flight with and without payload, two-UAV cooperative transport, and operation in πθ\pi_\theta78 m/s average wind, with outcomes including stable auto-takeoff, rapid disturbance rejection under pushes and wind, and payload figure-8 tracking under gusts.

The two CrazyMARL systems converge on several common implications. Both emphasize decentralized execution without runtime communication, both rely on shared policy structure, and both use simulation environments intended to preserve transferability to real Crazyflie hardware. At the same time, they expose distinct limitations. In the landing case, the source argues for extensibility from two agents to tens of Crazyflies through parameter sharing and linear inference cost (Aschu et al., 2024); this is a stated implication rather than an experimentally demonstrated scaling result. In the transport case, the source explicitly identifies permutation-sensitive peer-state encoding as a current limitation and points to order-invariant or attention-based peer encoders, centralized critics such as MAPPO for larger teams, and integrated obstacle avoidance for cluttered environments as promising extensions (Lorentz et al., 17 Sep 2025).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to CrazyMARL.