HEHA for Multi-Robot Exploration
- The paper introduces a HEHA framework that assigns dedicated scout and task roles, using MI-UCB for joint optimization of exploration and exploitation.
- It employs hierarchical and decentralized planning with global mHPP allocation and local TSP decomposition to efficiently coordinate heterogeneous robots.
- Empirical results demonstrate up to 134% performance improvement over conventional methods, with reduced redundancy and communication overhead.
HEHA (Heterogeneous Exploration–Exploitation Architecture) is a hierarchical planning paradigm for multi-robot exploration characterized by explicit division of roles, tight coupling of information gain with task objectives, and scalable decentralized coordination. In the context of multi-robot systems, especially those comprising agents with diverse functional and mobility capabilities, HEHA frameworks enable superior exploration efficiency, minimize redundancy, and outperform conventional monolithic or heuristic baselines across a wide spectrum of environments and robot platforms.
1. Conceptual Foundations and Role Decomposition
HEHA formalizes the division of a robot team into separate exploration and exploitation agents, often referred to respectively as “scouts” and “tasks” (Lee et al., 2021). For an -robot team operating in an unknown environment , two subsets are defined:
- : Scout robots, tasked exclusively with gathering information and reducing environmental uncertainty.
- : Task robots, responsible for mission completion using gathered information (e.g., high-resolution mapping, target confirmation). Scouts and tasks may partially overlap (dually-equipped robots). The salient innovation of HEHA is to sidestep the classical scalar “explore-exploit” trade-off by concurrently executing both through dedicated subsystems, optimized jointly over a unified surrogate metric.
2. MI-UCB Surrogate Objective for Joint Policy Optimization
Classical approaches to multi-robot exploration often decouple information gathering and objective fulfillment, or combine them via heuristic utility weighting. HEHA instead defines a mutual-information upper-confidence-bound (MI-UCB) as a single acquisition function for joint optimization (Lee et al., 2021): where:
- : mutual information between scout measurements and latent environment, quantifying prospective uncertainty reduction,
- : reward for task trajectories given the (unknown) environment,
- : optimism-tuning parameter,
- The first term guides scouts toward informative regions; the second enables task robots to maximize expected concrete returns.
This acquisition function admits tractable evaluation using Bayesian grid-based beliefs, and its structure is such that information-seeking and task-reward maximization are simultaneously, yet independently, incentivized.
3. Hierarchical and Decentralized Planning Architectures
HEHA frameworks employ hierarchical architectures to address scalability and computational complexity, especially as the number of robots and environmental volume increases.
3.1 Global Planning
At the global level, task allocation is handled by algorithms that account for agent heterogeneity and capability constraints (Yang et al., 5 Oct 2025). Problem instances are typically formulated as constrained multiple Hamiltonian Path Problems (mHPP), with frontier clusters and traversability models: where is the subset of agents capable of traversing frontier 0.
The Partial Anytime Focal search (PEAF) algorithm introduced in (Yang et al., 5 Oct 2025) reduces the branching factor of global planning using partial expansion and label dominance, returning solutions with bounded sub-optimality in makespan.
3.2 Local Planning
Local planners further decompose the assigned clusters into tractable TSPs and utilize cluster-level heuristics to prioritize “hetero-frontiers” (key frontiers only visitable by a subset of robots). Local cost adjustment and assignment policies (based on robot priority 1) prevent duplicated exploration among heterogeneous teams.
3.3 Decentralized Coordination
Decentralized Monte Carlo Tree Search (Dec-MCTS) is adopted for coordination, with each robot maintaining its own MCTS tree over action sequences and sampling other agents’ intentions via exchanged distributions. The MI-UCB reward is employed at every rollout, ensuring that agents, even without perfect global synchronization, select mutually informative and complementary behaviors (Lee et al., 2021).
4. Deep Learning Integration and Communication-Efficient Representations
Frontier-based clustering, sparse map encoding, and DRL-based assignment—together with graph neural network (GNN) allocation functions—are integrated into modern HEHA frameworks (Cai et al., 2024). The three-tier stack comprises:
- Frontier identification and sparsification: Removes map noise and clusters frontiers to O(2) candidate centers, reducing communication overhead.
- Multi-graph neural network (mGNN) planner: Learns robot-to-cluster preferences with Proximal Policy Optimization (PPO), optimized for coverage gain and time efficiency.
- Local routing tier: Utilizes a fast subsequence-reversal search to optimize local tours, circumventing the cost of full TSP enumeration.
Adaptive compression and graph abstraction reduce cross-robot bandwidth by over 30% compared to grid-based transmit paradigms (Cai et al., 2024).
5. Geometric Priors and Hilbert-Augmented Approaches
A distinct instantiation of HEHA is the introduction of Hilbert space-filling curves to guide decentralized RL agents (Gurunathan et al., 23 Feb 2026). Each robot augments its observation space with a normalized Hilbert index 3 representing its current spatial position along a globally consistent 1D sweep. This positional prior supports:
- Structure-preserving exploration (robots naturally spread, reducing overlap and redundancy).
- Rapid convergence in sparse reward coverage MDPs for both DQN and PPO agents.
- Efficient online trajectory conversion to 4 waypoints with curvature/velocity constraints, ensuring physical feasibility for both swarm and legged robots.
The entirety of the control loop, including Hilbert index computation, path smoothing, and time-parameterization, remains computationally lightweight for real-time execution on constrained platforms.
6. Empirical Performance and Deployment
The following table summarizes representative empirical results:
| Method | Exploration Speed-up | Redundancy Reduction | Platform/Scenario |
|---|---|---|---|
| HEHA-PEAF (Yang et al., 5 Oct 2025) | Up to 30% (vs. NBVP) | 14–32% | Heterogeneous (wheeled, legged, aerial) robots, large mixed-terrain fields |
| MI-UCB (Lee et al., 2021) | 50–134% more task reward | (N/A) | Multi-drone surveillance, targets confirmation |
| Hier.-Region-based (Meng et al., 17 Mar 2025) | 20% vs. flat global plan | 15–25% | ROS2, simulated/real-world quadrupeds, 1–4 robots |
| DRL+SparseMap (Cai et al., 2024) | 13–20% fewer steps | >30% comm reduction | TurtleBots, simulation, various indoor scales |
| Hilbert-Augm. RL (Gurunathan et al., 23 Feb 2026) | 15–25% higher coverage ratio | ~30% | Grid swarms, Boston Dynamics Spot |
Multi-environment benchmarks (including 2D/3D fields with obstacle-rich, multi-level, and indoor-outdoor topologies) and real-world deployments confirm robustness of HEHA to sensor/model error, communication lag, and heterogeneous payloads. For example, (Yang et al., 5 Oct 2025) realized coordinated exploration of a 35×35 m environment by assigning stairs to a legged robot and flat regions to a wheeled robot within 140 s, with the solution robustly respecting traversability constraints.
7. Limitations, Extensions, and Future Work
Key limitations noted in current HEHA implementations include:
- Scalability: Centralized or partially centralized planners may incur polynomial complexity as 5, 6 grow. Distributed, multi-level task decompositions remain underexplored (Yang et al., 5 Oct 2025).
- Dynamic Environments: Most frameworks assume quasi-static scenes; explicit online adaptation to moving obstacles and time-varying constraints is not standard.
- Communication Robustness: Full connectivity is often assumed; intermittent or lossy communication and explicit decentralized belief fusion are active areas for development.
- Richer Heterogeneity: Extending the assignment/utility maps to encapsulate energy constraints, sensor heterogeneity, or dynamic task priorities remains an open direction. A plausible implication is that combining HEHA’s exploration–exploitation coupling with learning-based frontier selection and geometric priors may yield further gains in high-dimensional, time-varying real-world domains.
References
- "An Upper Confidence Bound for Simultaneous Exploration and Exploitation in Heterogeneous Multi-Robot Systems" (Lee et al., 2021)
- "HEHA: Hierarchical Planning for Heterogeneous Multi-Robot Exploration of Unknown Environments" (Yang et al., 5 Oct 2025)
- "A Hierarchical Region-Based Approach for Efficient Multi-Robot Exploration" (Meng et al., 17 Mar 2025)
- "An Enhanced Hierarchical Planning Framework for Multi-Robot Autonomous Exploration" (Cai et al., 2024)
- "Hilbert-Augmented Reinforcement Learning for Scalable Multi-Robot Coverage and Exploration" (Gurunathan et al., 23 Feb 2026)