- The paper introduces H-DQN and H-PPO, combining Hilbert-index observations, curve-guided exploration, and policy-invariant reward shaping for decentralized multi-robot coverage.
- At 16 agents, H-DQN achieves 0.895 coverage versus DQN’s 0.589, while H-PPO reaches 0.927 versus PPO’s 0.683 and reduces redundancy to 1.768 visits per unique cell.
- The paper shows that curve-guided action sampling drives most gains, while a waypoint interface successfully transfers Hilbert paths to a Boston Dynamics Spot robot, despite limited hardware validation.
Overview
This paper proposes a framework that integrates Hilbert space-filling curve (SFC) priors into decentralized multi-robot reinforcement learning for coverage and exploration tasks. The authors introduce H-DQN and H-PPO, which augment DQN (2602.19400) and PPO policies with a normalized Hilbert index appended to the local observation, together with a curve-guided exploration mechanism in which agents probabilistically select actions advancing them to the next cell along the Hilbert ordering. The framework additionally employs potential-based reward shaping along the Hilbert sequence, preserving policy invariance under the original MDP while biasing exploration toward locality-preserving traversal. Beyond simulation, the paper contributes a waypoint interface converting Hilbert orderings into curvature-bounded, time-parameterized SE(2) trajectories executable on a Boston Dynamics Spot legged robot via the Spot SDK.
The central motivation is that standard deep RL exploration strategies (ε-greedy, entropy maximization) perform poorly in large, sparse-reward environments, whereas SFCs provide a systematic traversal with minimal backtracking. Because all agents share the same Hilbert indexing scheme and differ only by initial offset, coordination emerges implicitly without centralized control or inter-agent communication.
Method
Each agent's observation st is augmented as st′=[st,ht], where ht is the normalized Hilbert index of its grid cell. In H-DQN, ε-probability actions advance to the next Hilbert cell rather than being uniformly random; Q-learning updates remain standard TD with target-network replay. H-PPO applies identical state augmentation and biased action sampling during rollouts; the clipped surrogate objective with GAE is unchanged, with ε annealed over training. Reward shaping uses Φ(ht)=αht, adding γΦ(ht+1)−Φ(ht) so that curve progression is encouraged without altering the optimal policy. The authors argue this injects a globally consistent spatial reference into each agent's decision process, reducing policy-gradient variance and providing implicit task decomposition across agents.
For hardware deployment, Hilbert indices are converted to metric waypoints, L-corners are inserted during arc-length resampling to avoid diagonal chords, headings follow path tangents, and time parameterization enforces velocity limits of the form vlim(s)=min{vmax,vcurv(s)} with curvature-dependent caps. Execution over the Spot SDK maintains time sync, lease, and E-Stop keepalive, gates on inflated keep-out zones, and reports command latency within tens of milliseconds buffered by reference-time scheduling.
Simulation Results
Experiments use 32×32 and 64×64 grids with 5×5 egocentric observations, four to sixteen homogeneous agents sharing a convolutional policy, trained for 300k steps and averaged over five seeds. Convergence speed is measured via a windowed stability criterion (rolling average reward within 90% of maximum for at least ten consecutive evaluation windows), a more robust estimator than first-threshold-crossing metrics.
Coverage ratio improves consistently and grows with team size. At 16 agents, DQN drops to 0.589 while H-DQN reaches 0.895; PPO falls to 0.683 while H-PPO reaches 0.927. Notably, baselines degrade at the largest team size—standard DQN "peaks and then collapses" at 16 agents per the authors—whereas both Hilbert-augmented variants improve monotonically:
| Method |
4 Agents |
8 |
12 |
16 |
| DQN |
0.676 |
0.729 |
0.721 |
0.589 |
| H-DQN |
0.831 |
0.882 |
0.890 |
0.895 |
| PPO |
0.712 |
0.760 |
0.772 |
0.683 |
| H-PPO |
0.871 |
0.904 |
0.917 |
0.927 |
Redundancy (visits per unique cell) shows the same pattern: baseline redundancy increases with team size (DQN: 4.17 at 16 agents), while H-PPO decreases it monotonically to 1.768 at 16 agents. This inverse scaling behavior supports the claim that shared spatial priors substitute for explicit communication in coordinating coverage.
Hilbert index computation via integer bit interleaving introduces negligible runtime overhead, supporting embedded deployment—a claim borne out by the real-robot experiments.
Ablations and Hardware Validation
Ablations in the 8-agent, 64×64 environment isolate three components: state augmentation, curve-guided exploration, and reward shaping. A salient finding is that the No State Augmentation variant performs comparably to full H-PPO, indicating that curve-guided sampling—not the augmented input—is the dominant driver of the benefit, particularly in early learning and spatial efficiency. State augmentation primarily improves late-stage stability and variance reduction. Removing exploration bias while retaining augmentation causes significant degradation. This attribution is a useful corrective against interpreting the gains as a representation-learning effect alone.
Hardware validation runs on a Spot quadruped covering a 10 m × 10 m area discretized into a 5×5 grid, with five trials per policy from a fixed start pose, logging odometry at 50 Hz. Waypoints are translated into linear steps (Δs=0.25 m at ~0.8 m/s) and discrete ±30° turns. Qualitatively, PPO covers all cells but exhibits additional revisits and longer traversals relative to H-PPO, whose planned waypoints align closely with executed odometry. Parameter sensitivity tests over step size and turn resolution show expected trade-offs between traversal speed and path fidelity. Reported failure modes include visual odometry drift and tight-turn misalignment, mitigated by heading snaps and keep-out inflation.
Limitations and Open Questions
The paper concedes several constraints directly. The framework assumes uniform grids and precomputable spatial priors, limiting generalization to irregular topologies, continuous domains, and dynamic obstacles—the very conditions common in real deployments. The hardware validation is limited to a single robot executing precomputed trajectories in a small, structured indoor area; multi-robot hardware trials remain future work, so the decentralized-coordination claim is validated only in simulation. The ablation finding that state augmentation contributes little beyond curve-guided sampling raises an open question about whether the learned policies actually exploit the Hilbert index or merely inherit structure from exploration. Finally, robustness to localization drift and perception noise, extension to SE(3) motion, and obstacle-aware or learned SFC variants that reindex around occlusions are left unaddressed.
Conclusion
This paper demonstrates that a simple geometric prior—the Hilbert curve—embedded as observation features, exploration bias, and potential-based shaping yields consistent gains in coverage ratio, redundancy, and scalability across team sizes in sparse-reward multi-agent RL, at negligible compute cost and without explicit communication, and translates to executable SE(2) trajectories on a legged platform. The strongest empirical result is monotone improvement with swarm size where baselines collapse; the most instructive caveat is the ablation showing curve-guided exploration, not state augmentation, drives most of the improvement.